The rule
A model's weights take about parameters times bytes per parameter. Billions of parameters times bytes per parameter gives gigabytes directly: a 70 billion parameter model at 1 byte per parameter is about 70GB.
| Parameters | FP16 (2 bytes) | FP8 (1 byte) | 4-bit (about 0.5 byte) |
|---|---|---|---|
| 8 billion | 16 GB | 8 GB | about 4 GB |
| 32 billion | 64 GB | 32 GB | about 16 GB |
| 70 billion | 140 GB | 70 GB | about 35 GB |
| 120 billion | 240 GB | 120 GB | about 60 GB |
| 200 billion | 400 GB | 200 GB | about 100 GB |
A DGX Spark has 128GB of memory shared between the CPU and GPU. That is the total for weights, cache and everything else, so you never get all of it for the model.
Leave room for the rest
- KV cache: grows with context length and the number of requests served at once. Long contexts and big batches can use tens of gigabytes.
- Activations, and for training, gradients and optimizer state, which can be several times the size of the weights for full fine-tuning.
- The operating system and your own software.
Reading the table
- 70B at FP16 (140GB) does not fit. At FP8 (70GB) it fits with moderate headroom. At 4-bit (about 35GB) it fits comfortably.
- 120B fits at 4-bit and, tightly, at FP8, which leaves little for cache.
- 200B fits only at 4-bit, at about 100GB. This agrees with NVIDIA's statement that a single DGX Spark can run inference on models up to 200 billion parameters at FP4.
- Two nodes pool 256GB across two machines, which fits larger models or lets you use higher precision.
For fine-tuning
Methods like QLoRA keep the base model in 4-bit and train small adapters, so the memory cost is close to the 4-bit weights plus a modest overhead. That is why a 70B model is practical to fine-tune this way while full fine-tuning of the same model is not.
These are estimates from arithmetic, not measurements. Check your model's real memory use with a short test before you plan a long run.
Frequently asked questions
Does unified memory mean I get all 128GB for the model?
No. The CPU, GPU, operating system and your software all share it, so plan for noticeably less than 128GB of usable room for weights and cache.