Podvek
Console

What does a 70B QLoRA fine-tune cost on a rented DGX Spark?

A quick estimate from a published throughput figure, and why it is a starting point for a budget, not a quote.

Updated October 7, 2026

The input

NVIDIA's DGX Spark performance post reports 759.79 "peak tokens/sec" for Llama 3.3 70B QLoRA fine-tuning with PyTorch, at sequence length 2048 and batch size 8, over 64 steps. The post defines the figure as batch size × steps × sequence length ÷ total training time. Time is tokens divided by throughput, and cost is time multiplied by the hourly rate. We use the single-node on-demand rate at the time of writing, $0.75 per hour.

The arithmetic
tokens = 100_000_000
tokens_per_second = 759.79   # NVIDIA: Llama 3.3 70B QLoRA, seq 2048, batch 8, 64 steps
hours = tokens / tokens_per_second / 3600
print(round(hours, 1), round(hours * 0.75, 2))   # 36.6 hours, $27.42

The estimate

Estimated time and cost for Llama 3.3 70B QLoRA at 759.79 tokens per second and $0.75 per hour
Training tokensEstimated timeEstimated cost
10 millionabout 3.7 hoursabout $2.74
100 millionabout 36.6 hoursabout $27.42
1 billionabout 366 hoursabout $274

The same arithmetic for the two smaller runs

NVIDIA published two more fine-tuning runs in the same post. Here is what one billion training tokens would take at each published throughput.

One billion training tokens at NVIDIA's published throughput and $0.75 per hour
RunPublished tokens per secondHoursCost
Llama 3.2 3B, full fine-tune13,519.54about 20.5about $15.41
Llama 3.1 8B, LoRA6,969.59about 39.9about $29.89
Llama 3.3 70B, QLoRA759.79about 366about $274

Why a real run can differ

  • The figure comes from a 64-step run in one configuration. It is one data point, not a guarantee.
  • The estimate counts training only. Loading the model, evaluating, saving checkpoints and trying settings add time.
  • Your sequence lengths and batch size may differ, and throughput changes with them.
  • Software versions change results. Published figures date from the time they were measured.

Use the table to size a budget. Then run a short test on a few thousand tokens and extrapolate from your own measured throughput.

Tokens in practice

A training token is not a word. English text often averages somewhat more than one token per word, and other languages usually need more tokens per word. To count your own dataset, run it through the tokenizer for the model you plan to train.

Frequently asked questions

How much does it cost to fine-tune a 70B model?

At NVIDIA's published 759.79 tokens per second and $0.75 per hour, 100 million training tokens is about 36.6 hours and about $27. Real runs can differ, so measure a short run first.

Launch when you are ready

Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.