What NVIDIA says it can do
NVIDIA's product page says a single DGX Spark can fine-tune AI models up to 70 billion parameters and run inference on models up to 200 billion parameters. The datasheet specifies FP4 precision for the 200 billion figure. NVIDIA's January 2026 software post also describes distributed fine-tuning of models up to 70B across two linked systems, using FSDP and LoRA.
Why memory is the point
Fine-tuning needs room for the model weights, the optimizer state, gradients and activations. A DGX Spark's CPU and GPU share one 128GB pool, so a model that would have to be split or offloaded on a smaller card can stay in memory. The trade-off is memory bandwidth: 273 GB/s is far lower than datacenter GPUs, so runs take longer than on an H100.
Methods that suit the machine
- LoRA and QLoRA: train small adapter matrices instead of every weight. QLoRA keeps the base model in 4-bit, which is how a 70B model fits.
- Full fine-tuning of smaller models: NVIDIA's benchmark includes full fine-tuning of Llama 3.2 3B.
- Distributed fine-tuning across two nodes with FSDP and LoRA, on the dual-node size.
Published throughput
NVIDIA's performance post reports "peak tokens/sec" for three fine-tuning runs on one DGX Spark. The post defines it as batch size × steps × sequence length ÷ total training time, over a run of only 64 steps. Treat each number as one data point for that configuration, not a promise: real runs add data loading, evaluation and checkpointing, and longer sequences change the result.
| Model | Method | Batch size | Peak tokens per second |
|---|---|---|---|
| Llama 3.2 3B | Full fine-tune | 8 | 13,519.54 |
| Llama 3.1 8B | LoRA | 4 | 6,969.59 |
| Llama 3.3 70B | QLoRA | 8 | 759.79 |
At about 760 tokens per second, one billion training tokens on a 70B QLoRA run is roughly 366 hours of compute. The cost estimate shows the arithmetic for other sizes.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model
name = "meta-llama/Llama-3.3-70B-Instruct"
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(name, quantization_config=bnb, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(name)
lora = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"], task_type="CAUSAL_LM")
model = get_peft_model(model, lora)
model.print_trainable_parameters()This is the standard transformers + PEFT recipe, shown to illustrate the shape of a QLoRA run. We have not run it on a DGX Spark. Whether bitsandbytes ships working builds for aarch64 with CUDA 13 is something to check before you rely on it.
Why rent instead of buy for this
Fine-tuning is bursty: a few long runs, then days of data preparation and evaluation. Per-minute billing means you pay for the run, not for a machine sitting idle between experiments.
Frequently asked questions
Can I fine-tune a 70B model on one DGX Spark?
NVIDIA states a single DGX Spark can fine-tune models up to 70 billion parameters. NVIDIA's own 70B benchmark uses QLoRA, where the base model is held in 4-bit.
Is it as fast as an H100?
No. Memory bandwidth is much lower, so expect longer runs. The advantage is capacity: 128GB of unified memory.