For years I ran engineering. Depending on the stage of the company that meant a team of eight or a team of fifty, but the job never really changed: split the goal, hand each piece to someone specific, notice early when it is going wrong, and turn down nine good ideas out of ten. I do more or less the same thing today. The team is just one person and a fleet of AI agents.

The org chart looks familiar. There is a coordinator who holds a release together. There is operations, the only role allowed to touch the live system. There is a development lead who takes longer topics and breaks them into pieces for its own juniors. There is a tester and an analyst. The one genuinely different thing is the unit of capacity. With people it was headcount and a calendar. Here it is a budget: each role runs on a separate subscription with its own weekly ceiling, and that ceiling can be spent in a single afternoon.

I expected the hard part to be output quality. Quality turned out to be the easy half. What broke were the things I never had to manage with people, because people do them on their own.

An agent will not tell you it is out of capacity

When a person runs out of capacity, they say so. Sometimes late, usually reluctantly, but they say it. An agent cannot. When it hits the ceiling of its subscription, reporting the problem is itself work, and work is exactly what it is no longer allowed to do. It simply stops mid-sentence.

The worse part is that stopping looks identical to succeeding. An agent that finished its turn and is waiting for the next instruction is, from the outside, indistinguishable from an agent that ran out of budget, and from an agent stuck on a question it never asked anyone. Three very different states, one identical signal: silence.

With people, a stand-up solves this. You meet, someone looks uncomfortable, you ask. In a fleet you cannot learn the state from the agent; you have to learn it about the agent, from outside. So I had to build a watchdog that checks, on a fixed rhythm, whether anything has moved in the last few dozen minutes, and separates working silence from dead silence. It is the first piece of equipment a human team never needs.

Unspoken intent does not exist

A human team runs largely on things nobody says out loud. I mention once that I want to be told before anyone touches a particular part of the system, and it holds for a year. Nobody re-raises it. Nobody asks whether it still applies.

Nothing holds that way with agents. The rules they read at the start fade after a few hours of work — not through a fault, but through crowding: step twenty of the task is simply nearer than a sentence from the preamble. I once handed over a task made of several dependent steps. The agent did the first step, reported it, and waited. That was not a failure; it did exactly what was written. What failed was the intent. "Do this" meant "and then keep going" to me. It did not to the agent. And from the outside, that waiting was indistinguishable from being stuck.

So now two sentences go into every single brief, to the point of tedium: run it to completion and a finding is not a task. Whatever has to be true always has to be in every brief, not in the onboarding document. That is an uncomfortable mirror for a manager. It turns out that what I had called a "clear brief" for years was half a brief and half shared context that the other person quietly supplied for me.

Diligence is expensive

In one day my agents burned half the weekly ceiling of the most expensive subscription. Not on a large task. On small things nobody had asked for: tidying, renaming, readability, two fixes that "presented themselves" along the way. Every individual change was defensible. Together they were a full day of work and zero approved results.

Human teams have a safeguard against this that nobody puts on a slide, because it sounds unflattering: fatigue and reluctance. A person who is supposed to close one ticket only starts cleaning the room next door if it genuinely bothers them. They have other priorities, a Friday afternoon, two meetings. An agent has none of that. Give it access and it stays diligent until the budget is gone.

The rule that came out of this is blunt and it works. A plan with a cost estimate first, work second. And a finding an agent trips over on the way does not become work; it goes on a list, with just enough analysis to decide whether it deserves any at all. "I noticed that" is not a mandate.

Bureaucracy grows after every incident

This is the lesson I am least proud of, because I did it to myself. Every time something failed, I added a check. A deploy broke, so I added a gate. A mistake slipped through, so I added a second reader. Each individual step was sensible and, at the time, cheap. A few months later, changing a single sentence of copy on a website passed through four to six reviews by different models.

I only saw it when I added up what each part of the fleet had cost me over a week. The most expensive line item was not any of the things the fleet exists to do. It was the coordinator relaying messages — roughly four and a half times more than all code review put together. I was mostly paying for my agents to pass notes to each other.

This is where the contrast with people is sharpest. Human bureaucracy brakes itself, because people quietly route around it; when a rule is stupid it stops being followed, and sooner or later you find out. Agents execute it faithfully, to the last item. No pressure ever comes from below to delete it. Whoever introduced it has to delete it.

What helped was sorting changes into three classes by consequence rather than by size. Copy ships without a blessing. Anything touching payments never ships without one. In between sits a band where one independent read is enough. Plus one change of topology: technical roles talk to each other directly, and the coordinator receives decisions, not traffic.

"Verified" is not verified

Hundreds of green tests, and signing in to the application still did not work. The test had substituted its own stand-in for the authentication layer and then diligently tested the stand-in. The green was honest: the stand-in worked perfectly. Real sign-in did not. One click in a browser found it.

The second time was more expensive. An agent reported that a password change had gone through and had been verified. It had verified it by comparing a file on disk. But the running services had read that password at startup and had been using the old one ever since. They spent hours with no database behind them while the state on disk looked completely fine.

Both have the same core. An agent does not verify reality; it verifies its model of reality — and it does so calmly, because it lacks the uneasy feeling a person gets looking at their own finished work. An experienced engineer occasionally thinks "this went suspiciously well". An agent has no such instinct, and its report sounds exactly as confident whether it is true or not.

The resulting rule sounds harsh but is mostly practical: the result is not verified by whoever produced it. And it is not verified by me either — I have neither the time nor the distance. It is verified by a role with a single job: to try, from the outside and by the path a user would take, whether the thing does what it claims.

Agents build spaceships

The last lesson is the oldest management disease in new packaging. A simple request came in: correct three specific records. What came back was a proposal for a new tool that would handle such cases in general, with its own interface and its own review cycle. Again, none of it was stupid. Every step followed logically from the one before. The only thing missing was someone to say that three records do not justify new code.

On a human team, experience and — again — laziness do this job. A senior engineer will not start building a tool for three records, because they know what it costs and who will be maintaining it a year from now. An agent has no such memory. So cutting scope is work that nobody in the fleet can do except the human. It is the same managerial job it always was; it just has to happen far more often, and out loud.

What did not change

Set against the years with people, the fundamentals hold. A clear brief is still a clear brief, except you now find out immediately, because there is nobody to fill in the rest. One owner per thing still applies; two agents on one task argue exactly like two people, only faster. A decision that is not written where the next person or agent will see it is gone within a month — with people it took a little longer, because they remembered.

And the hardest part of the job is unchanged: saying "we are not doing this". With people it cost me popularity. With agents it costs me money, because every unspoken "no" turns into finished work within minutes — work somebody then has to read and throw away.

It does not replace management. It multiplies it

If I had to compress it into one sentence that not everyone will agree with: an AI team does not replace leadership, it multiplies it. Mistakes included. A vague brief does not produce confusion in three people by the end of the week; it produces it in five agents within ten minutes, and on the meter. A missing "no" does not stay a single idea, it comes back as a finished thing. A check added in a moment of fright does not fade away; it keeps running obediently until I remove it.

The multiplication runs both ways, which is why I keep doing this. A well-written brief propagates to five places at once. A decision recorded once holds for the whole fleet. None of it happens by itself, though. Compute turned out cheaper than I expected. The overhead of managing it turned out dearer.