Published specifications
| DGX Spark | GeForce RTX 5090 | |
|---|---|---|
| Memory | 128 GB LPDDR5x, unified with the CPU | 32 GB GDDR7 |
| Memory bandwidth | up to 273 GB/s | 1,792 GB/s |
| What it is | A complete AI computer with Arm CPU | A graphics card for a PC |
What the numbers mean
The RTX 5090 has about 6.6 times the memory bandwidth of a DGX Spark (1,792 divided by 273, our arithmetic). For a model that fits in 32GB, it will generate text much faster. In LMSYS's review, gpt-oss-20b at MXFP4 with Ollama and batch size 1 decoded at 205 tokens per second on the RTX 5090 against 49.7 on the DGX Spark. Both were measured the same way, and the roughly four-fold gap is in the same direction as the bandwidth gap.
But 32GB is a hard limit. A 70B model needs about 35GB for its weights even at 4-bit, which does not fit. The DGX Spark can hold it with room to spare.
A simple way to choose
| Your model needs | Better fit |
|---|---|
| Under about 27 GB including cache | RTX 5090, for speed |
| 27 to 108 GB including cache | DGX Spark, because it fits |
| More than 108 GB | Dual or quad DGX Spark nodes (planned) |
Our rule of thumb: use 85 percent of each machine's memory (27 of 32 GB, 108 of 128 GB) and keep the rest for the operating system and runtime overhead. It is arithmetic, not a measurement, so check your own model before relying on it.
def weights_gb(params_billion, bytes_per_param):
return params_billion * bytes_per_param
print(weights_gb(70, 0.5)) # 35.0 GB at 4-bit: does not fit in 27 GB, fits in 108 GB
print(weights_gb(32, 0.5)) # 16.0 GB at 4-bit: fits on an RTX 5090Rent instead of buy either one
If you only need the large-memory machine for a few experiments, renting the DGX Spark by the minute avoids buying hardware that will mostly be idle.
Frequently asked questions
Which is faster for AI?
For models that fit in 32GB, the RTX 5090 is faster because its memory bandwidth is much higher. For anything larger, only the DGX Spark can hold the model.