The Kitchen Brigade: How I Run a Team of AI Agents Without Burning Money
My AI agents processed 28 billion tokens in the last six months. Almost none of it was new.
In short: I don't use one big AI model for everything. One model works like a head chef: it plans, tastes and signs off. Cheaper models work like cooks: they cook from clear orders and report back. And almost everything the models read comes from a prepared station — the cache — so it's quick and cheap. That's how a number as scary as 28 billion stays reasonable.
Why a kitchen?
For six months, most of my software work — my camera e-commerce, internal tools, bots, scrapers, this website — has been done together with AI coding agents inside OpenCode.
I started the obvious way: one assistant, one long conversation, everything in it. It works, but it's like asking the head chef to also peel every potato. Expensive, slow, and the chef is too busy peeling to taste the soup.
So I reorganised it like a professional kitchen — a brigade. AI models come in sizes; from Anthropic, think Haiku (small and cheap), Sonnet (mid-size) and Opus (the top tier). Each size gets a job that fits it:
| Kitchen role | AI role | What it does |
|---|---|---|
| Head chef | The director — my main conversation, a top model | Plans, splits the work, tastes the result, ships it |
| Line cook | Implementer — a mid-size model | Cooks most of the code from a clear order |
| Senior cook | Senior implementer — a top model | Delicate dishes, or when the line cook burned it |
| Junior cook | Small implementer — a small model | Follows an exact recipe card, nothing more |
| Runner | Reader — the cheapest model | Goes through the pantry, reads every label, comes back with a short list |
| Visiting chef | Reviewer — a model from another company | Tastes and comments. Never touches the stove |
The numbers behind it
I keep a small dashboard that reads OpenCode's local activity database (a side project I call katas). From April 6 to October 3, 2026:
- 1,337 conversations with the head chef.
- 2,631 helper sessions — cooks called in for one specific job.
- 28.2 billion tokens in total. A token is roughly three-quarters of a word.
- 93.8% of those tokens were the cache being re-read. 5.4% was text stored in the cache for the first time. Brand-new input was just 0.43%, and the models' own writing 0.36%.
- The helpers used about a quarter of all tokens, but wrote about half of everything the system produced.
So the huge total is mostly the same text being read again and again — cheaply. Keep that in mind; it's the heart of this post.
Rule 1: the head chef doesn't peel potatoes
Calling a cook isn't free. Before they start, someone has to explain the order: what we're making, where things are, how it should look. When I measured it, that briefing came to 28,000–49,000 tokens per helper (median) before it did any work — and the helper re-reads it at every one of its steps, about 21 on average.
Analogy: You don't call a cook over from another station to boil one egg. By the time you've explained where the pots are, you could have boiled it yourself.
I learned this the expensive way. In July my head chef was calling cooks for everything — on the busiest days, 100 to 270 helpers, with peak days re-reading almost a billion tokens of cache. On July 29 I measured the briefing cost, trimmed the instructions every helper has to read, and wrote one rule: the head chef only hands work out when it buys something real — two dishes cooking at the same time, or a big job kept off its own counter. Small things it simply does itself.
The effect was immediate. In the four weeks before the rule, 144 conversations called 1,525 helpers — about 11 each. In the four weeks after, 150 conversations called 620 — about 4 each. About the same number of conversations — less than half the helpers.
Two small habits keep this cheap:
- Give the address, not the groceries. The head chef tells a cook which files to open, never pastes their contents into the order. Pasting means paying for the same text twice.
- Send orders out together. Independent jobs leave in one round, so three cooks work side by side instead of waiting in line.
Rule 2: the prepared station is the secret (caching)
Here's what most people don't realise. A language model has no memory between steps. Each time it takes a step — opens a file, runs a command, writes a line — it reads the entire conversation again from the top: the instructions, the files, everything said so far.
That sounds hopelessly wasteful. It would be, without caching.
Analogy: Mise en place. Before service, a good cook chops the onions, measures the spices and lines everything up. During service they don't chop again — they grab from the prepared station. The cache is the AI's prepared station: what it has already read is kept ready, and grabbing it costs a small fraction of reading it fresh.
Overall, about 94% of everything the models read came from the cache. For the helpers, with their short and focused orders, it was almost 96%. The kitchen rules that protect it:
- Don't rearrange the station mid-service. The cache only works on the beginning of the conversation. Change the instructions halfway and everything after them must be prepped again.
- Keep the station clean. A long, wandering conversation is a station buried in leftovers; every step means digging through it again. One clear job per helper keeps it tidy.
- Prepped onions don't wait forever. The cache expires after a few idle minutes, so steady work is cheaper than stop-and-go.
Rule 3: escalate once, never loop
Cheap cooks fail sometimes. The rule is strict: one step up, once. If that fails too, the head chef stops and rethinks the plan instead of trying harder.
The point isn't saving a few tokens on one dish. It's avoiding the classic disaster: an agent retrying the same broken thing twenty times, re-reading the whole conversation on every try.
It happened while I was writing this post. The cheapest reader wouldn't start because of a configuration problem on my machine. Nobody retried it. The job went one step up — four mid-size readers went through my projects' documentation in parallel — and the configuration problem went on the to-fix list. A slightly more expensive afternoon, but no loop.
Rule 4: a visiting chef tastes, but never cooks
Only one family of models is allowed to write code. Models from a different company may only read: review a plan, check a deploy, point out risks. And I call them in myself — the head chef can't invite them.
Analogy: A good kitchen invites a chef from another restaurant to taste the new menu. A fresh palate catches what everyone in the kitchen has stopped noticing — and because the guest never touches the stove, there's never any doubt about who changed the recipe.
What I'd tell someone starting out
- Give each model a role, and write the roles down. Mine live in one short file the agents read at the start of every session.
- Hand out work only for two jobs at once, or to keep a big job off your counter. Every helper pays a briefing cost.
- Protect the cache. Stable instructions at the top, focused sessions, file names instead of file contents.
- Escalate once. Loops are where budgets go to die.
- Measure it. I only learned what a helper really costs — and that almost all of my tokens were the cache being re-read — because I built a dashboard. Before that, I was guessing.
The real win of running it like a kitchen isn't even the savings. It's that the head chef — the expensive, careful model — spends its time on what it does best: deciding what to cook, and tasting every plate before it leaves the kitchen. That's how 28 billion tokens turn into finished work instead of a very expensive mess.
More in AI Agents
This is the only post here so far — browse the section or all posts.