Podvek
Console

Quantization on a DGX Spark: FP4, FP8 and GGUF

Quantization is how large models fit. The format you pick decides how much memory you save and which software you can use.

Updated October 7, 2026

Bytes per parameter

A model's weights take roughly the number of parameters multiplied by the bytes used per parameter. Lower precision means fewer bytes.

Approximate weight memory by precision
PrecisionBytes per parameter70B model120B model
FP16 or BF162140 GB240 GB
FP8170 GB120 GB
4-bit (FP4, GGUF Q4)about 0.5about 35 GBabout 60 GB

These are weights only. The KV cache, activations and the operating system need room too, and some 4-bit formats add a small overhead for scaling factors. A 70B model at FP16 does not fit in 128GB; at FP8 it does, with limited headroom; at 4-bit it fits comfortably.

Convert a 16-bit GGUF to 4-bit with llama.cppExample output
llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M

Formats you will meet

  • NVFP4 and MXFP4: 4-bit floating-point formats. The DGX Spark's GPU has hardware support for FP4, and NVIDIA's January 2026 post says NVFP4 reduces memory use by about 40 percent.
  • FP8: half the memory of FP16, typically at a small quality cost.
  • GGUF: the file format used by llama.cpp, with quantisation levels such as Q4 and Q8 that you can pick per model.

Quantization and speed

Because decode speed is bound by memory bandwidth, reading fewer bytes per token usually means faster generation. A 4-bit model reads about a quarter of the data of the same model at FP16.

Check quality on your own task

Lower precision can reduce accuracy, and the effect varies by model and task. Run your evaluation set on each variant before choosing one.

Frequently asked questions

Will a 70B model fit at FP8?

The weights take about 70GB, so it fits in 128GB with room for a moderate KV cache. Long contexts or large batches use more memory.

Launch when you are ready

Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.