What people run on a DGX Spark
Fine-tuning, local inference, retrieval pipelines and more. Each page explains what fits in 128GB of unified memory and how to get started.
Fine-tune large language models on a rented DGX Spark
NVIDIA says a single DGX Spark can fine-tune models up to 70 billion parameters. What fits in 128GB, which methods to use, and how to estimate a run.
Run large models for inference on a DGX Spark
Serve open-weight models up to about 200 billion parameters at FP4 on one machine. What to expect from decode speed, and which model types run best.
Build retrieval-augmented generation pipelines on one machine
Run the embedding model, vector index and generation model together on a single DGX Spark with 128GB of unified memory, and prototype a RAG system end to end.
Quantization on a DGX Spark: FP4, FP8 and GGUF
How lower-precision formats change memory use and speed on a DGX Spark, and how to estimate whether a quantised model will fit in 128GB.
Run model evaluations and benchmarks without owning the hardware
Evaluation jobs are short, repeatable and bursty. Why per-minute rental on a DGX Spark suits them, and how to run a clean, reproducible comparison.
Launch when you are ready
Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.