Ask what an “AI employee” costs and you get a per-seat price. It is the one line item that arrives on an invoice, which is why everybody quotes it. It is also the least informative number in the whole calculation, and this piece is about what belongs there instead.
The best anchor I have for the Czech market is a survey by Iva Brejlová, published on Lupa on 10 June 2026: how much AI costs technology firms in the Czech Republic. The figures: at one firm a developer cost around €15 per month in 2024 and around €120 in 2026 (a frontend developer with Figma about €130) — an eight- to ninefold rise, most of the difference coming from a single item, a €105 subscription. Another firm reports about $150 per developer and about $20 per non-developer. A third spends roughly 3,000 koruna per head monthly, with an internal ceiling of 20,000 koruna per developer and its heaviest users reaching $1,000 a month. One firm now puts 30% of its annual software budget into AI tooling. Token bills in that survey range from hundreds of koruna to 50 thousand a month, and a bespoke agent costs 50,000 to 800,000 koruna as a one-off. Globally, from the same article: Stripe spends about $100,000 a day on tokens, and Uber burned its annual AI budget in four months and then capped spending at $1,500 per employee per month.
Notice what all those numbers have in common. Every one of them is per-seat or subscription-shaped. And the range — hundreds of koruna to fifty thousand — is the actual finding. Not any average you could compute from it.
What I measured on myself
Rather than estimate, I went into the local logs of agentic work on a single workstation. The sample: 120 session files, records from 7 September to 5 October 2026, 713 billable model calls. The billable records fall on five days only — on the other days the machine was off or the work happened elsewhere. So this is a lower bound, not total consumption.
Total tokens: 88,010,250. The breakdown that matters:
| Category | Tokens | Share |
|---|---|---|
| Cache read (re-reading context already paid for) | 83,791,432 | 95.21% |
| Cache write (1h TTL 2,651,114 + 5m TTL 1,051,471) | 3,702,585 | 4.21% |
| Output — what the models actually wrote | 514,781 | 0.58% |
| Genuinely new input text over the whole period | 1,452 | 0.0016% |
That last figure is not a typo. Across nearly a month of work, 1,452 tokens of text entered the models that had never been there before. Everything else was context handed over again and again. The metered equivalent of this consumption at Claude API prices is $70.28: Opus 5 $56.38 (80.2%, 462 calls), Sonnet 5 $11.72 (16.7%, 237 calls), Opus 4.7 $2.18 (3.1%, 14 calls).
The unit is the turn, not the token
Here is the accounting point. Price a single model call and Opus 5 gives a median of 142,909 tokens and $0.1006, a 90th percentile of 199,940 tokens and $0.19, and a maximum of 234,602 tokens and $0.5148. Averaged across all models it is $0.0986 per call. One agent turn is therefore a fairly stable line item of about ten cents — containing over a hundred and forty thousand tokens, almost all of them a re-read of what the agent already knows.
In my sample, 171 tokens are read for every token produced. That is not waste you can switch off; it is how agentic work operates. Every turn the model sees the task, the history, the files and the tool results again. You are paying for reading your own context, not for writing.
The consequence for anyone building a spreadsheet is direct. My effective price per million delivered output tokens came to $137. The list price for Opus 5 output is $25 per million. So estimating cost as “how much text do I want, times the list output price” is wrong by roughly 5.5× — and wrong downward, which is the unpleasant direction. For completeness, the price list I computed against (verified 5 October 2026): Opus 5 $5 input / $25 output per million, Sonnet 5 $2 / $10, Haiku 4.5 $1 / $5; cache reads at a tenth of the input price, cache writes at 1.25× (5-minute TTL) or 2× (1-hour TTL).
A subscription is an anaesthetic
Now the admission that matters most: I did not pay that $70.28. The work came out of a flat-rate subscription. It is a shadow price — the metered equivalent of work billed as a flat fee. And that is precisely why subscriptions are treacherous in accounting: they hide everything happening underneath.
They hide this, for instance. The same work without prompt caching would have cost $383.30 at metered prices instead of $70.28 — 5.45× more. That entire difference is context hygiene: how long sessions live, how they get compacted, how often the cache is warm. It is the single largest lever in the whole calculation, and your subscription invoice does not show it.
They also hide the distribution over time. The most expensive day in the sample (5 October) cost $50.70 — 72% of the entire period — on 66,208,484 tokens. The most expensive 10% of calls accounted for 25.4% of the cost. With consumption shaped like that, a monthly average is a worthless figure. What matters is the peak and the tail, because that is what hits the cap.
I have a direct comparison for scale. On the genuinely metered API — the one that powers production features of my own applications — I hold a ceiling of $30 a month, with $40 as an absolute maximum. That is less than a single day of my agentic work would cost at metered prices. As long as a flat fee covers it, nobody sees it in the budget.
The line items with no invoice
The cost of a cap. Twice — on 31 May and 25 June 2026 — I ran out of credit or hit the monthly ceiling on a metered API. A model-based extraction step in one of my pipelines silently stopped for about five days before anyone noticed. Monitoring watched data ingestion, which kept running; nothing watched the extraction, which was failing. No data was lost, but for five days there was no output. The accounting entry is therefore not “saved X”, it is “five days not delivered”. A spending cap is not a savings instrument, it is a silent off switch.
The cost of fixing and reviewing. I have a case where an agent reported finished work that was not in the code — a commit announcing a job it had not done. And a case where three finished articles were live in production but missing from git; that one surfaced by accident. Tests do not catch this class of thing; human review does, and that review is a line item whose cost rises with the agent’s output volume. The more productive the agent, the more expensive the review — not less.
Human review time. And here my data runs out. How many hours review costs me I cannot compute from the logs. I do know what would have to be measured: timestamps from the moment output is presented to approval or rejection, and the number of rounds before it passes. I do not measure that yet, so the item belongs in the budget as explicitly unquantified — not as zero. The most common accounting lie in this field is not a wrong number. It is a line booked as zero because no invoice arrived for it.
The comparison trap: why 343× is not a price change
The same author wrote on Lupa on 24 September 2026 about a Czech e-shop that replaced generative agents with a decision model, reporting 343× lower costs, processing time down from 32.4 s to 1.7 s, and accuracy 30 to 40% higher. That swap may well have been exactly right, and I have no reason to dispute it.
Methodologically, though, it is worth naming what the number is. 343× is not a change in price, it is a change in task. Categorising products with a fixed, short output and doing agentic work over a long context are not the same job — in one you pay for a decision, in the other for repeatedly reading state. Cost per task may only be compared at equal task and equal acceptance bar. Otherwise a factor of two or of three hundred can be manufactured almost at will, in either direction.
A five-line method
Copy this into your own spreadsheet:
- Subscriptions and licences — the only item with an invoice. Record it, but do not lean on it.
- Metered tokens — real consumption by production features, kept separate from agentic work.
- Context mechanics — cache read, cache write, compaction. For me this was a factor of 5.45×, the biggest lever in the calculation.
- Failed turns and rework — calls that produced no accepted output, plus the work to repair them.
- Human review time — keep the row even when you cannot fill in the number.
And three traps people fall into with that table: the average instead of the peak (72% of my cost landed on one day), the list output price as the basis of an estimate (off by about 5.5×), and the line booked as zero because no document exists for it.
What this does not establish
Three limits, all of them material. This is one workstation, not a fleet, and the billable records land on five days, so it is a lower bound. It is a metered equivalent of work actually paid for by subscription — the numbers describe the price that work would carry, not the price I paid. And review time was not measured, which means the probably-largest item in the whole calculation is missing from it.
I was also measuring over a sample that included the very session doing the counting. During the hour I spent on it, the figures moved by a few percent. This is a snapshot, not a set of closed accounts.
One claim that can be re-measured in a month and that could prove me wrong: I expect the cache-read share to stay above ninety percent and the cost distribution to stay similarly lopsided, with the most expensive day carrying most of the period. If a month from now consumption turns out to be spread evenly, then my argument against averages rested on a small sample, and I will say so.