The pieces
- An embedding model that turns documents and queries into vectors.
- A vector index or database that finds the closest documents.
- Optionally a reranker that reorders the top results.
- A generation model that writes the answer from the retrieved text.
Why one machine helps
On a small graphics card you often have to choose which of these fits in memory. With 128GB shared between CPU and GPU, the embedding model, the reranker and a quantised generation model can all stay loaded while the vector index sits in system memory. Add up the weights of each model before you start; the sizing guide shows how.
A workflow that suits per-minute billing
- Launch an instance and copy your documents over.
- Chunk and embed the documents once, and save the index to your own storage.
- Start the generation model and iterate on prompts, chunk size and retrieval settings.
- Run an evaluation set against the pipeline and record the results.
- Copy the index and results out, then delete the instance.
There is no persistent storage between rentals, so the index is the thing to copy out. Next time, restore it instead of embedding again.
Serving the result
Each instance has ten service ports. You can expose your API on one of them for a demo or an internal test, with authentication in front of it.
Frequently asked questions
Can I keep my vector database between rentals?
Not on the machine. Copy the index out before you delete the instance and restore it on your next rental.