Part of a series on building cz-agents → under the hood

The thousands-of-dollars question

The client use case was unambiguous: RAG over confidential documents — bids in a tender — with a hard constraint that none of them may leave the local network. No third-party API, no cloud model. The data stays home, and so does the model. This is an increasingly common brief for this class of task, and it is technically satisfiable: Open WebUI as the interface, bge-m3 for embeddings running on an ARM server, and the generative model via ollama on a work laptop with a Ryzen 6800H, where I dedicated six cores.

The question was not whether it could be built. The question was whether it was good enough — and if not, whether a GPU was needed, and which one. Behind that question sits a decision worth thousands: buy a card, then discover it does not have the VRAM for your model, and you bought the wrong thing. I wanted that decision anchored in numbers, not in forum vibes.

What a CPU can do (and what it can't)

First I measured what I already had. llama3.1:8b on six cores of the Ryzen 6800H gave roughly 20 tokens per second on prompt reading and 6.5 tokens per second on generation. A real RAG query that pulls around seven thousand tokens of retrieved documents into context therefore took 4 minutes 22 seconds. The answer quality was usable. The latency was not — nobody works interactively at that per-query cost, and an analyst who has to go through dozens of bids will not sit through it.

So I tried a bigger model. qwen3:14b on the same CPU produced an answer in about 15 minutes. Qualitatively it was visibly better — better structure, more reliable table reading, fewer missed details. But 15 minutes per query is a dead number. I knew I wanted the quality of the 14B model at sub-minute latency. A CPU cannot do that, and no env var will change it. The question narrowed: which is the cheapest GPU that delivers that combination.

A rental instead of a purchase

This is where most people reach for the manufacturer's spec sheet and guess. Instead, I rented the GPU for one evening. vast.ai is a marketplace where you rent a specific GPU by the hour from individuals and datacenters; an RTX 3090 with 24 GB of VRAM came to $0.126 per hour, around $0.19 per hour with storage. I uploaded the same stack, ran the identical prompt, and measured.

The result: 1,889 tok/s on prompt reading, 59 tok/s on generation, an answer in 14.7 seconds. Eighteen times the CPU. And crucially — the model correctly found both prices I had deliberately planted in the test documents. That evening I added an RTX 5090 with 32 GB at $0.416 per hour. There qwen3:32b fit entirely into VRAM and answered in 35 seconds, while qwen3:14b ran at 1,986 tok/s / 108 tok/s / 19.6 s and was good enough in quality on my test set. The total bill for both benchmarks, storage included: a few dollars.

The decision stopped being an act of faith. The 14B model plus a 24 GB card handles the task. The 32B model did not meaningfully raise quality — it just asked for a more expensive card. The GPU purchase is deferred until a real, measured need pays for it — and when that need arrives, I know exactly what to reach for.

So much for the story. More useful than the numbers, though, were the things that broke along the way. Half of them had nothing to do with the card's performance.

Flash attention: one env var, a fivefold difference

OLLAMA_FLASH_ATTENTION=1 is off by default. On the RTX 5090 it was the difference between 296 and 1,539 tok/s on prompt reading — more than fivefold, from a single environment variable. For RAG, where the overwhelming majority of work is precisely reading a long context, that is not a detail; it is the headline number.

The other side of that coin is more dangerous. Had I run the benchmark with the default, prompt eval would have looked terrible, the conclusion would have read "the 3090 is slow, buy a stronger card," and that decision would have been factually wrong. A benchmark that does not run with the right configuration does not lie about one thing — it lies about everything. Before I compare hardware, I have to be sure I am measuring the software in its best shape.

24 GB is not 24 GB

A 32B model in Q4 quantization with an 8k context does not fit in 24 GB. Ollama does not throw an error — instead it quietly offloads part of the layers to the CPU (partial offload), and performance collapses from tens of tokens per second to 9.5 tok/s. An order-of-magnitude degradation, without a single warning; only the log reveals that part of the model is not on the GPU.

The lesson is banal, and yet I keep repeating it even to myself: VRAM has to hold not just the model, but the model plus the KV cache, which grows with context length. The longer the context (and RAG is by definition a long context), the larger the cache. Anyone sizing a card by the weight-file size is sizing it wrong, and finds out in production, where it is the worst place to find out.

num_ctx 4096 silently truncates RAG

The default context length in ollama is 4096 tokens. When you push seven thousand tokens of retrieved documents into the prompt, the model simply discards the first half — the log shows a laconic truncating input prompt and nothing more. The answer arrives, sounds confident, and was in fact produced over half the source material. Nobody notices, because the output looks complete.

This is the most insidious bug in the whole benchmark, because it has no symptom. You can spot a slow model. A model that cannot see half the documents and confidently says nothing about it, you spot only if you have a test set with a known answer. Without one, RAG on the default quietly degrades quality while pretending everything is fine.

When streaming goes silent

The ollama:latest image had streaming that did not get along with Open WebUI — the interface received empty responses, while the same query via the non-stream API returned correct text. I spent half an hour hunting for a bug in the prompt and the RAG pipeline before it dawned on me that the problem was a version incompatibility at the ends of the stream.

The fix was boring: pin specific versions on both sides rather than trusting latest. But it is a reminder that with a local stack you carry the integration risk yourself. When two open-source components stop talking to each other, there is no vendor behind you — it is you, the log, and two versions that have to meet.

Slow network = SIGBUS

Cheap GPU hosts on the marketplace have one quiet downside: a slow connection. When the download of runtime libraries is cut off halfway, you get SIGBUS at runtime — a crash that looks like a hardware fault but is in fact an incompletely downloaded file. I lost one instance that way before I traced the cause.

So inet_down>1000 belongs in the instance filter just as naturally as the card type and VRAM size. Network speed is not a luxury; it is the precondition for the stack to even come up reliably. The cheapest offer on a slow link ends up costing more than an instance a few cents dearer that comes up on the first try.

On a rented GPU, only fictional data

And now the most important thing, which has nothing to do with performance. A rented GPU runs on someone else's machine, to which the host has root. For a benchmark of confidential data, those two facts are mutually exclusive — so only fictional data went to the rented GPU. Real tender bids never touched the vast.ai landscape.

But governance does not end there; it begins there. I built the test set as generated fictional documents into which I deliberately inserted planted hooks: two similar items differing in a single detail — say, two variants of the same component with a different specification and a different price. That makes it objectively measurable whether RAG found that detail, or whether the two items merged into a single confident falsehood. It is precisely on that hook that the difference shows between a card that answered fast and a card that answered fast and correctly.

Eval without ground truth is just vibes. "Looks good" is not a metric. Without a known correct answer you cannot tell a model that solved the task from a model that convincingly faked it — and with RAG over sensitive documents, that difference is the whole product. A test set with planted hooks is cheaper than one wrong purchase order, and it answers a question that a speed benchmark on its own cannot.

What I take from it

A GPU worth thousands can be test-driven before you buy it for the price of lunch. A few dollars on a GPU marketplace, your own test set with known answers, and an evening of measuring — and the purchase decision stops being faith in a datasheet. That alone is reason enough to rent first and buy later.

But the main lesson is elsewhere. Of the six things that decided the outcome, only one was genuinely about hardware. The rest were two env vars, one silent default, one version incompatibility, and one filter on network speed. Had I measured naively, I would have gotten precise numbers about a misconfigured system and, with a clear conscience, bought a more expensive card than I needed. The costliest mistake in sizing is not a slow GPU. It is a confident measurement that nobody validated.

Facing a similar decision in your company? → AI consulting