The bound
To produce one token, a dense model reads essentially all of its weights from memory once. So the best possible speed is memory bandwidth divided by the model's size in memory. It is an upper bound: it ignores the KV cache and every other overhead.
| Dense model | Weights in memory | Upper bound, tokens per second |
|---|---|---|
| 70B at FP8 | 70 GB | about 3.9 |
| 70B at 4-bit | about 35 GB | about 7.8 |
LMSYS measured Llama 3.1 70B at FP8 on SGLang at 2.7 tokens per second, below the 3.9 bound and in the same range, as expected.
Why mixture-of-experts models are different
A mixture-of-experts model has many parameters in total but activates only some of them for each token, so it reads far less than its full size per token. NVIDIA reports gpt-oss-120b at MXFP4 on llama.cpp decoding at about 55 tokens per second on a DGX Spark, much faster than the dense 70B figure, even though the model is larger overall. The two numbers come from different publishers and software, so compare the pattern, not the exact ratio.
How the machine compares
| Device | Memory bandwidth | Ratio to DGX Spark |
|---|---|---|
| DGX Spark | 273 GB/s | 1x |
| H100 PCIe | 2,000 GB/s | about 7.3x |
| RTX 5090 | 1,792 GB/s | about 6.6x |
| H100 SXM | 3.35 TB/s | about 12x |
| H200 | 4.8 TB/s | about 18x |
What to do about it
- Quantise. Fewer bytes per weight means less to read per token and faster generation.
- Prefer mixture-of-experts models when you need a large model that responds quickly.
- Batch requests. Reading the weights once can serve several requests, so total throughput rises with batch size even when single-stream speed does not.
- Use the machine for what its capacity is good at: large models, long contexts, experiments, fine-tuning.
Frequently asked questions
Is the bound the same as measured speed?
No, it is an upper limit. Real speed is lower once the KV cache, software overhead and other work are included.