Podvek
Console

Run large models for inference on a DGX Spark

Capacity is the strength of this machine. Decode speed depends heavily on the model type and the software.

Updated October 7, 2026

What fits

NVIDIA states a single DGX Spark can run inference on models up to 200 billion parameters using FP4 precision. At about half a byte per parameter, 200 billion parameters is roughly 100GB of weights, which leaves headroom in 128GB for the KV cache and the operating system.

What speed to expect

Generating each new token requires reading the model's active weights from memory, so decode speed is limited by memory bandwidth, which is 273 GB/s on this machine. Two things follow. Dense models decode slowly as they grow. Mixture-of-experts models, which read only part of their weights per token, decode much faster than dense models of similar total size.

Published decode speeds on one DGX Spark
ModelPrecisionSoftwareBatch sizeDecode tokens per secondSource
gpt-oss-20bMXFP4llama.cpp182.74NVIDIA
gpt-oss-120bMXFP4llama.cpp155.37NVIDIA
gpt-oss-20bMXFP4Ollama149.7LMSYS
Llama 3.1 8BFP8SGLang120.5LMSYS
Llama 3.1 70BFP8SGLangnot stated2.7LMSYS

NVIDIA ran its inference tests with 2048 input tokens, 128 output tokens and batch size 1. These numbers come from different publishers, software and dates, and they change as the software improves. The two gpt-oss-20b rows show how much the software alone matters: 82.74 with llama.cpp against 49.7 with Ollama, on the same hardware class.

Same model, three machines

LMSYS measured gpt-oss-20b at MXFP4 with Ollama and batch size 1 on three machines in the same review. This is the one place where speeds can be set side by side, because one author measured them the same way.

gpt-oss-20b, MXFP4, Ollama, batch size 1, measured by LMSYS
MachinePrefill tokens per secondDecode tokens per second
DGX Spark2,05349.7
GeForce RTX 50908,519205
RTX Pro 6000 Blackwell10,108215

The DGX Spark is about four times slower here, which fits its lower memory bandwidth. What it buys you is room: the RTX 5090 stops at 32GB. See DGX Spark vs RTX 5090.

Try it yourself

OllamaExample output
ollama run gpt-oss:20b

To measure decode and prefill speed for any GGUF model, use llama-bench. This is the command from the llama.cpp project's DGX Spark benchmark thread, followed by the results its maintainers published:

llama-bench: published result, llama.cpp build 11fb327bf (7941), NVIDIA-SMI 580.95.05, CUDA 13.0Published result
llama-bench -m gpt-oss-120b-mxfp4.gguf -ngl 99 -fa 1 -ub 2048 -p 2048 -n 32

model                      test      t/s
gpt-oss 120B MXFP4 MoE     pp2048    2443.91 ± 7.47
gpt-oss 120B MXFP4 MoE     tg32      58.72 ± 0.20
gpt-oss 20B MXFP4 MoE      pp2048    4505.82 ± 12.90
gpt-oss 20B MXFP4 MoE      tg32      83.43 ± 0.59
qwen3moe 30B.A3B Q8_0      pp2048    2986.97 ± 18.87
qwen3moe 30B.A3B Q8_0      tg32      61.06 ± 0.23

pp2048 is prefill of 2048 tokens and tg32 is generating 32 tokens. The rows are condensed from the maintainers' table, and the command line is shortened to the flags that matter.

Batching changes the picture

LMSYS measured Llama 3.1 8B at FP8 on SGLang at 20.5 tokens per second for one request and 368 tokens per second in total at batch size 32. If you serve several users or run many prompts at once, total throughput is much higher than the single-stream number.

Good fits

  • Trying a large open model before committing to bigger hardware.
  • Offline batch jobs: summarising, classifying or labelling a dataset.
  • Long-context work where the model and its KV cache do not fit on a smaller card.
  • Agent prototypes that call a local model many times.

Frequently asked questions

Why is a 70B model so much slower than gpt-oss-120b?

A dense 70B model reads all of its weights for every token. gpt-oss-120b is a mixture-of-experts model that reads only part of its weights per token, so it decodes faster despite being larger overall.

Launch when you are ready

Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.