Bytes per parameter
A model's weights take roughly the number of parameters multiplied by the bytes used per parameter. Lower precision means fewer bytes.
| Precision | Bytes per parameter | 70B model | 120B model |
|---|---|---|---|
| FP16 or BF16 | 2 | 140 GB | 240 GB |
| FP8 | 1 | 70 GB | 120 GB |
| 4-bit (FP4, GGUF Q4) | about 0.5 | about 35 GB | about 60 GB |
These are weights only. The KV cache, activations and the operating system need room too, and some 4-bit formats add a small overhead for scaling factors. A 70B model at FP16 does not fit in 128GB; at FP8 it does, with limited headroom; at 4-bit it fits comfortably.
llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_MFormats you will meet
- NVFP4 and MXFP4: 4-bit floating-point formats. The DGX Spark's GPU has hardware support for FP4, and NVIDIA's January 2026 post says NVFP4 reduces memory use by about 40 percent.
- FP8: half the memory of FP16, typically at a small quality cost.
- GGUF: the file format used by llama.cpp, with quantisation levels such as Q4 and Q8 that you can pick per model.
Quantization and speed
Because decode speed is bound by memory bandwidth, reading fewer bytes per token usually means faster generation. A 4-bit model reads about a quarter of the data of the same model at FP16.
Check quality on your own task
Lower precision can reduce accuracy, and the effect varies by model and task. Run your evaluation set on each variant before choosing one.
Frequently asked questions
Will a 70B model fit at FP8?
The weights take about 70GB, so it fits in 128GB with room for a moderate KV cache. Long contexts or large batches use more memory.