Blog
Sizing, costs and practical notes for running AI workloads on a DGX Spark.
What fits in 128GB? A sizing guide for model weights
Work out in a minute whether a model will fit on a DGX Spark: parameters times bytes per parameter, plus room for the KV cache. With a table for common sizes.
October 7, 2026
Why token speed depends on memory bandwidth, and what it means for a DGX Spark
Text generation reads the model's weights for every token, so speed is capped by memory bandwidth. A back-of-envelope bound for the DGX Spark's 273 GB/s.
October 7, 2026
What does a 70B QLoRA fine-tune cost on a rented DGX Spark?
Using NVIDIA's published QLoRA throughput for Llama 3.3 70B on a DGX Spark (759.79 tokens/s), estimate hours and cost for 10M, 100M and 1B training tokens.
October 7, 2026
Why we do not write "X times faster than an H100"
Our rules for DGX Spark numbers, and three details often quoted wrongly: the 1 petaFLOP figure, the nvidia-smi memory readout and the compute capability.
October 7, 2026
Launch when you are ready
Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.