Podvek
Console

Run model evaluations and benchmarks without owning the hardware

An evaluation run is a few hours of work you may repeat a dozen times. You do not need a machine for the days in between.

Updated October 7, 2026

Why evaluation suits renting

Benchmarks, regression suites and A/B comparisons run in bursts: start a job, wait for it to finish, read the numbers. Paying by the minute means a two-hour run costs two hours, with no idle machine to justify.

A clean setup

  1. Launch a fresh instance, so there is nothing left over from earlier runs.
  2. Install exact versions of your serving software and evaluation harness, and write the versions down.
  3. Fix random seeds, sampling settings and prompts.
  4. Run every model variant on the same instance in one session, so hardware conditions match.
  5. Copy the raw outputs and logs out, then delete the instance.

What to record

  • Model name, precision and the software and its version.
  • Batch size, context length and sampling settings.
  • Both prefill speed and decode speed, not just one of them.
  • The quality metric for your task alongside the speed numbers.

Published speed figures for this hardware vary with framework and version, so compare your own runs against each other rather than against numbers from elsewhere.

When one machine is not enough

Models too large for 128GB can use the dual-node size, which pools 256GB across two machines.

Frequently asked questions

Are results on a rented machine comparable between runs?

Each rental is a dedicated machine with nobody else on it, which helps consistency. Still record software versions and settings, since those change results more than the machine does.

Launch when you are ready

Top up, then pick a size. If every machine is busy, waiting in the queue costs nothing.