Podvek
Console

What fits in 128GB? A sizing guide for model weights

Before you rent anything, do this arithmetic. It tells you which models fit, and at what precision.

Updated October 7, 2026

The rule

A model's weights take about parameters times bytes per parameter. Billions of parameters times bytes per parameter gives gigabytes directly: a 70 billion parameter model at 1 byte per parameter is about 70GB.

Approximate weight memory for common model sizes
ParametersFP16 (2 bytes)FP8 (1 byte)4-bit (about 0.5 byte)
8 billion16 GB8 GBabout 4 GB
32 billion64 GB32 GBabout 16 GB
70 billion140 GB70 GBabout 35 GB
120 billion240 GB120 GBabout 60 GB
200 billion400 GB200 GBabout 100 GB

A DGX Spark has 128GB of memory shared between the CPU and GPU. That is the total for weights, cache and everything else, so you never get all of it for the model.

Leave room for the rest

  • KV cache: grows with context length and the number of requests served at once. Long contexts and big batches can use tens of gigabytes.
  • Activations, and for training, gradients and optimizer state, which can be several times the size of the weights for full fine-tuning.
  • The operating system and your own software.

Reading the table

  • 70B at FP16 (140GB) does not fit. At FP8 (70GB) it fits with moderate headroom. At 4-bit (about 35GB) it fits comfortably.
  • 120B fits at 4-bit and, tightly, at FP8, which leaves little for cache.
  • 200B fits only at 4-bit, at about 100GB. This agrees with NVIDIA's statement that a single DGX Spark can run inference on models up to 200 billion parameters at FP4.
  • Two nodes pool 256GB across two machines, which fits larger models or lets you use higher precision.

For fine-tuning

Methods like QLoRA keep the base model in 4-bit and train small adapters, so the memory cost is close to the 4-bit weights plus a modest overhead. That is why a 70B model is practical to fine-tune this way while full fine-tuning of the same model is not.

These are estimates from arithmetic, not measurements. Check your model's real memory use with a short test before you plan a long run.

Frequently asked questions

Does unified memory mean I get all 128GB for the model?

No. The CPU, GPU, operating system and your software all share it, so plan for noticeably less than 128GB of usable room for weights and cache.

Launch when you are ready

Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.