It started with ordinary frustration. For over twenty years I worked in tech around finance — payment systems, banks, R&D teams that had to take seriously where data came from and how much it could be trusted. When I later wanted to hand a language model Czech public data, I hit an old familiar problem from a new angle: the source registers look tidy until you try to read them by machine. Formats vary by record type, the important fields end up in PDFs, and one field is sometimes full and sometimes empty with no warning. I decided to turn that into something small, readable and auditable — and to run it for six months so I'd learn what running it actually costs.

This page is therefore a permanent notebook, not a launch. The product lives elsewhere; here I wanted to record the decisions and what they cost.

Why I got into it

I spent long enough in an environment where "missing value" and "zero" are not the same thing, and where guessing in place of a missing fact is an operational risk, not a detail. That instinct travels badly into a world where a model gets a prompt and happily fills in what it lacks. I wanted the opposite default: an interface that would rather return a partial result with a reason than a smooth sentence about a company that may not exist.

Czech registers are a rich source — they just weren't built for programmatic reads with consistent latency, predictable errors and stable contracts. They were built for a human with a browser. The whole project is really a long answer to one question: how do you turn data meant for the human eye into a layer a machine can trust? I dug into that in the essay on the economics of dirty data; here it's enough to say this is the motive that holds everything else together.

Betting on a protocol

The first big decision was about format. I could have built a plain REST API and been done — I know that world and it's predictable. Instead I bet on the Model Context Protocol, at the time a fairly young open standard for connecting tools and data sources to models. Betting on a protocol in its early days always has two sides. The downside is obvious: the spec moves, libraries mature as you go, the docs run ahead of reality, and every so often a change you didn't expect catches up with you.

The upside is subtler, but for this kind of product it matters. A REST endpoint describes how to get the data — the path, the parameters, the shape of the response. A protocol like MCP additionally carries what the tool is: named tools, their schemas, and the boundaries a model can read without me spelling them out in every prompt. For data where the structure and its limits matter, the difference between "download this JSON" and "here is a tool and these are its rules" is surprisingly large. And because an ecosystem of clients grew around MCP, one well-described server suddenly worked in several places without me writing an integration for each.

In hindsight it was the right bet, paid for in small change. When you move inside an early protocol, you write code against a moving target, and part of the work is making peace with rewriting something because the world around you shifted. But in exchange I got a boundary between agent and system that I don't have to keep re-explaining.

What running it revealed

Most of what I learned came not from building but from running. Nine servers have been in production for half a year, and a few things honestly surprised me.

Who actually calls

Naively I expected mostly humans, or their agents, on the other end. Reality was more varied. A large share of traffic is catalogue crawlers indexing available MCP servers, and monitors periodically checking the service is alive. Then come agents that actually compute something — and only behind them the occasional human. The most important lesson from this is banal and yet easy to miss: a request is not the same as usage. A server can take thousands of inbound connections a day and only a handful of them mean a real query. When I first looked at raw request counts, I nearly read an entirely different story than the one that was actually happening. Useful metrics only started once I separated handshakes and health checks from real tool calls.

Rate limits are a protocol question, not just a shield

At first I treated rate limiting as a defence against overload. Running it convinced me it's more a part of the contract. An agent behaves differently from a human: on a single task it can stack dozens of queries in seconds, because none of them cost it anything. So a limit isn't a wall against an attacker but a way to tell the model clearly that behind the data sits a finite resource. The key was that the refusal be readable — that the response make plain this is a temporary ceiling, not a data error. A model given a clear "not now, try later" behaves sensibly. A model given emptiness fills in the blank.

Open-core as an architectural boundary

At first I assumed the open servers would be enough — and technically they are. Over time I hit a boundary that doesn't run between "free" and "paid" but between two kinds of value. Code you can read, run and modify is one thing. Operational responsibility — that the underlying datasets are current, that the parser didn't break on a new combination of documents, and that someone handles the incident when a public source behaves differently — is another. I wanted that boundary in the architecture, not just in a price list. The open part therefore lives as standalone packages with no hidden runtime dependencies; the layer on top is separated on purpose, so the self-hosted variant never depends on something you can't see. That decision was primarily technical; the business consequences followed from it, not the other way around.

What I took away

After half a year the main lesson fits in one sentence: the protocol is the easy part, the data is the hard one. MCP gave me a clean boundary between model and system, but the overwhelming majority of the work still sits in understanding the sources — their rate limits, failure behaviour, the meaning of fields, and the places where the honest answer is "I don't know". The second lesson is about measurement: without separating noise from signal, your telemetry tells you a story that didn't happen. And the third is about humility toward young standards — betting on an early protocol pays off when you're willing to pay continuously in small rewrites rather than once in a big failure.

What's next: keep the servers boring and predictable, improve the readability of refusals and partial results, and slowly widen coverage where it makes sense by the data, not by marketing. The code is open on GitHub, and I'm happy to be shown where I'm wrong.

Deeper dives

This page is a signpost. Each decision above has its own essay where I go into depth: