Specifications
| Component | Specification |
|---|---|
| Superchip | NVIDIA GB10 Grace Blackwell, CPU and GPU linked by NVLink-C2C |
| CPU | 20-core Arm: 10 Cortex-X925 and 10 Cortex-A725 |
| GPU | NVIDIA Blackwell architecture with 5th-generation Tensor Cores |
| Memory | 128 GB LPDDR5x, one coherent pool shared by CPU and GPU, 256-bit |
| Memory bandwidth | up to 273 GB/s |
| AI performance | up to 1 petaFLOP at FP4 with sparsity (theoretical) |
| Storage | 4 TB NVMe M.2 with self-encryption (the version we offer) |
| Networking | ConnectX-7 at 200 Gbps, 10 GbE RJ-45, Wi-Fi 7 |
| Power | 240 W external power supply |
| Size and weight | 150 x 150 x 50.5 mm, 1.2 kg |
| Operating system | NVIDIA DGX OS, based on Ubuntu 24.04 |
Reading the headline numbers
The 1 petaFLOP figure is a theoretical FP4 number that relies on sparsity. It is not comparable to the dense FP8 or FP16 figures usually quoted for datacenter GPUs, so do not use it to predict how fast your job will run.
The number that most shapes day-to-day experience is memory bandwidth. At 273 GB/s it is far below datacenter GPUs, which limits how quickly large dense models can generate text. See the bandwidth post for how to estimate it.
Because memory is unified, nvidia-smi reports memory usage as "Not Supported" on this machine, although per-process GPU memory is still listed. Check usage with system tools or from inside your framework. Why this trips people up.
What NVIDIA says it is for
- Inference on models up to 200 billion parameters, using FP4 precision.
- Fine-tuning models up to 70 billion parameters.
- Two linked systems (256GB): models up to 400 billion parameters at FP4. Four systems need an external switch: up to 700 billion.
- Prototyping and development, with work then moved to DGX Cloud or a datacenter.
What to expect when you log in
These are the commands we would run first on a fresh instance. The output shows the shape of what you should see: the names and values follow NVIDIA's specifications, but we have not measured them on a DGX Spark ourselves yet.
$ nvidia-smi -L
GPU 0: NVIDIA GB10 (UUID: GPU-…)
$ nvidia-smi --query-gpu=name,compute_cap --format=csv,noheader
NVIDIA GB10, 12.1
$ nproc
20NVIDIA lists 20 Arm cores and reports the GB10's compute capability as 12.1 on its developer forum. The compute_cap query field may not exist on every driver version, so treat that line as the one most likely to differ.
Sizes we rent
| Size | Machines | Unified memory | Status |
|---|---|---|---|
| Single node | 1 | 128 GB | Available first |
| Dual node | 2, linked with a 200G cable | 256 GB | Planned |
| Quad node | 4, through a switch | 512 GB | Planned |
Where it is not the right tool
It is not built to compete with datacenter GPUs on throughput. If you need to serve many users at high speed or train a model from scratch, a datacenter GPU is the better choice. See the comparisons for numbers.
Frequently asked questions
How much memory does a DGX Spark have?
128 GB of LPDDR5x unified memory shared between the CPU and GPU, with up to 273 GB/s of bandwidth.
Can it run a 70 billion parameter model?
Yes. NVIDIA states it can fine-tune models up to 70 billion parameters and run inference on models up to 200 billion parameters at FP4. A 70B model needs about 35GB of weights at 4-bit and about 70GB at FP8.
Is the 1 petaFLOP number comparable to an H100's?
No. It is a theoretical FP4 figure with sparsity, while datacenter GPUs are usually quoted at dense FP8 or FP16.