GB10 · 128GB unified · 273 GB/s · 1 tenant/machine
A dedicated NVIDIA DGX Spark,billed by the minute.
128GB of unified memory, root over SSH and your own encrypted volume. For fine-tuning, local inference and evals that do not fit on a 24GB or 32GB card.
What you are renting
One NVIDIA DGX Spark, all to yourself. The numbers below come from NVIDIA's datasheet.
- Superchip
- GB10
- 20 Arm cores and a Blackwell GPU, linked by NVLink-C2C
- Memory
- 128 GB
- LPDDR5x shared by CPU and GPU, up to 273 GB/s
- AI compute
- 1 PFLOP
- Theoretical FP4 with sparsity. Not comparable with FP8 or FP16 figures.
- Power and size
- 240 W
- Whole system, 1.2 kg
Three steps to start
01
Sign in and add an SSH key
Sign in with your email, then save the public key you will log in with.
02
Top up and launch
Add funds to your wallet, pick a size and launch. If every machine is busy you join the queue.
03
Log in
SSH in with the command shown in the console and get to work. Delete the instance when you are done and billing stops.
What you get
One machine, one tenant
No sharing of GPU or memory. All 128GB of unified memory is yours. Two- and four-node sizes, up to 512GB, are planned.
Billed per minute
Billing starts when your container is ready and stops when you delete the instance. Queueing, provisioning and wiping are free.
Free queue when machines are busy
When your turn comes we start the instance and email you, or hold it until you confirm.
Delete means destroyed
Your data lives on its own encrypted volume with a fresh key for every rental. Deleting the instance destroys the key.
What people run on a DGX Spark
Fine-tuning, local inference, retrieval pipelines and more. Each page explains what fits in 128GB of unified memory and how to get started.
Fine-tune large language models on a rented DGX Spark
128GB of unified memory is enough to fine-tune models that will not fit on a typical 24GB or 32GB graphics card.
Run large models for inference on a DGX Spark
Capacity is the strength of this machine. Decode speed depends heavily on the model type and the software.
Build retrieval-augmented generation pipelines on one machine
A RAG system needs several models and a database. On one machine with a shared memory pool, they can live side by side.
Quantization on a DGX Spark: FP4, FP8 and GGUF
Quantization is how large models fit. The format you pick decides how much memory you save and which software you can use.
Run model evaluations and benchmarks without owning the hardware
An evaluation run is a few hours of work you may repeat a dozen times. You do not need a machine for the days in between.
Priced by the hour, billed by the minute
On demand, with no minimum spend. Prices are in US dollars.
Single node
×1$0.75per hour
$0.0125 per minute
- 1 DGX Spark
- 128GB unified memory
- Reservation (planned): $16.80/day
Two- and four-node sizes are planned. Prices and details are on the pricing page. All sizes and billing details
Launch when you are ready
Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.