Podvek
Console

Fine-tune large language models on a rented DGX Spark

128GB of unified memory is enough to fine-tune models that will not fit on a typical 24GB or 32GB graphics card.

Updated October 7, 2026

What NVIDIA says it can do

NVIDIA's product page says a single DGX Spark can fine-tune AI models up to 70 billion parameters and run inference on models up to 200 billion parameters. The datasheet specifies FP4 precision for the 200 billion figure. NVIDIA's January 2026 software post also describes distributed fine-tuning of models up to 70B across two linked systems, using FSDP and LoRA.

Why memory is the point

Fine-tuning needs room for the model weights, the optimizer state, gradients and activations. A DGX Spark's CPU and GPU share one 128GB pool, so a model that would have to be split or offloaded on a smaller card can stay in memory. The trade-off is memory bandwidth: 273 GB/s is far lower than datacenter GPUs, so runs take longer than on an H100.

Methods that suit the machine

  • LoRA and QLoRA: train small adapter matrices instead of every weight. QLoRA keeps the base model in 4-bit, which is how a 70B model fits.
  • Full fine-tuning of smaller models: NVIDIA's benchmark includes full fine-tuning of Llama 3.2 3B.
  • Distributed fine-tuning across two nodes with FSDP and LoRA, on the dual-node size.

Published throughput

NVIDIA's performance post reports "peak tokens/sec" for three fine-tuning runs on one DGX Spark. The post defines it as batch size × steps × sequence length ÷ total training time, over a run of only 64 steps. Treat each number as one data point for that configuration, not a promise: real runs add data loading, evaluation and checkpointing, and longer sequences change the result.

NVIDIA-reported fine-tuning throughput on one DGX Spark (PyTorch, sequence length 2048, 64 steps)
ModelMethodBatch sizePeak tokens per second
Llama 3.2 3BFull fine-tune813,519.54
Llama 3.1 8BLoRA46,969.59
Llama 3.3 70BQLoRA8759.79

At about 760 tokens per second, one billion training tokens on a 70B QLoRA run is roughly 366 hours of compute. The cost estimate shows the arithmetic for other sizes.

qlora.py: minimal QLoRA setup (not tested on DGX Spark)Example output
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model

name = "meta-llama/Llama-3.3-70B-Instruct"
bnb = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(name, quantization_config=bnb, device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(name)

lora = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"], task_type="CAUSAL_LM")
model = get_peft_model(model, lora)
model.print_trainable_parameters()

This is the standard transformers + PEFT recipe, shown to illustrate the shape of a QLoRA run. We have not run it on a DGX Spark. Whether bitsandbytes ships working builds for aarch64 with CUDA 13 is something to check before you rely on it.

Why rent instead of buy for this

Fine-tuning is bursty: a few long runs, then days of data preparation and evaluation. Per-minute billing means you pay for the run, not for a machine sitting idle between experiments.

Frequently asked questions

Can I fine-tune a 70B model on one DGX Spark?

NVIDIA states a single DGX Spark can fine-tune models up to 70 billion parameters. NVIDIA's own 70B benchmark uses QLoRA, where the base model is held in 4-bit.

Is it as fast as an H100?

No. Memory bandwidth is much lower, so expect longer runs. The advantage is capacity: 128GB of unified memory.

Launch when you are ready

Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.