Podvek
Console

Why we do not write "X times faster than an H100"

Every number on this site follows the same rules. Here they are, with three details about this machine that are easy to get wrong.

Updated October 7, 2026

Rule 1: a number comes with its context

A speed figure means little without the model, the precision, the software and who measured it. When one of those is missing from the source, we write "not stated" rather than fill it in. The software alone can move the result a lot: for gpt-oss-20b, NVIDIA reports 82.74 decode tokens per second with llama.cpp and LMSYS reports 49.7 with Ollama. See what to expect from inference.

Rule 2: only compare comparable things

We put the DGX Spark next to an H100 on memory capacity, memory bandwidth, power and form factor, because NVIDIA publishes those for both. We do not set the Spark's FLOPS next to an H100's, and we only write a multiple when it is plain division of published numbers, labelled as our arithmetic. See DGX Spark vs H100.

Three details that are easy to get wrong

1. The "1 petaFLOP" figure

NVIDIA's datasheet footnote says the up to 1 petaFLOP of AI performance is theoretical FP4 throughput using the sparsity feature. Datacenter GPUs are usually quoted at dense FP8 or FP16 numbers, so the two are not comparable and neither tells you how fast your job will run.

2. Memory in nvidia-smi

The CPU and GPU share one pool of memory, so NVIDIA's known-issues page says nvidia-smi reports memory usage as "Not Supported" on this machine. Per-process GPU memory is still listed. If a screenshot of a DGX Spark shows a tidy used/total memory bar from nvidia-smi, it did not come from this chip.

3. The compute capability

According to an NVIDIA employee's reply on the NVIDIA developer forum, the GB10 reports compute capability 12.1 (sm_121). If a snippet or a rental page tells you the DGX Spark is SM 10.0, it is not describing this chip. You can ask the machine directly:

Ask the GPU for its name and compute capabilityExample output
nvidia-smi --query-gpu=name,compute_cap --format=csv,noheader

We have not run this command on a DGX Spark ourselves yet, so we do not print its output here. When we have measured output, it will go on this site labelled as measured, with the date and software versions.

Frequently asked questions

Is the DGX Spark a 1 petaFLOP machine?

NVIDIA quotes up to 1 petaFLOP of AI performance, and the datasheet footnote says that is theoretical FP4 throughput using sparsity. It is not comparable with dense FP8 or FP16 figures.

Why does nvidia-smi not show my memory usage?

The CPU and GPU share one memory pool. NVIDIA's known-issues page says nvidia-smi shows memory usage as Not Supported on this machine, while per-process GPU memory is still listed.

Launch when you are ready

Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.