Podvek
Console

Why token speed depends on memory bandwidth, and what it means for a DGX Spark

A quick calculation explains why a 70B model feels slow on this machine while a larger mixture-of-experts model feels quick.

Updated October 7, 2026

The bound

To produce one token, a dense model reads essentially all of its weights from memory once. So the best possible speed is memory bandwidth divided by the model's size in memory. It is an upper bound: it ignores the KV cache and every other overhead.

Upper bound on decode speed on a DGX Spark at 273 GB/s
Dense modelWeights in memoryUpper bound, tokens per second
70B at FP870 GBabout 3.9
70B at 4-bitabout 35 GBabout 7.8

LMSYS measured Llama 3.1 70B at FP8 on SGLang at 2.7 tokens per second, below the 3.9 bound and in the same range, as expected.

Why mixture-of-experts models are different

A mixture-of-experts model has many parameters in total but activates only some of them for each token, so it reads far less than its full size per token. NVIDIA reports gpt-oss-120b at MXFP4 on llama.cpp decoding at about 55 tokens per second on a DGX Spark, much faster than the dense 70B figure, even though the model is larger overall. The two numbers come from different publishers and software, so compare the pattern, not the exact ratio.

How the machine compares

Memory bandwidth, NVIDIA published figures, and ratio to the DGX Spark
DeviceMemory bandwidthRatio to DGX Spark
DGX Spark273 GB/s1x
H100 PCIe2,000 GB/sabout 7.3x
RTX 50901,792 GB/sabout 6.6x
H100 SXM3.35 TB/sabout 12x
H2004.8 TB/sabout 18x

What to do about it

  • Quantise. Fewer bytes per weight means less to read per token and faster generation.
  • Prefer mixture-of-experts models when you need a large model that responds quickly.
  • Batch requests. Reading the weights once can serve several requests, so total throughput rises with batch size even when single-stream speed does not.
  • Use the machine for what its capacity is good at: large models, long contexts, experiments, fine-tuning.

Frequently asked questions

Is the bound the same as measured speed?

No, it is an upper limit. Real speed is lower once the KV cache, software overhead and other work are included.

Launch when you are ready

Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.