AI Agents

The Kitchen Brigade: How I Run a Team of AI Agents Without Burning Money

My AI agents processed 28 billion tokens in the last six months. Almost none of it was new.

In short: I don't use one big AI model for everything. One model works like a head chef: it plans, tastes and signs off. Cheaper models work like cooks: they cook from clear orders and report back. And almost everything the models read comes from a prepared station — the cache — so it's quick and cheap. That's how a number as scary as 28 billion stays reasonable.

Why a kitchen?

For six months, most of my software work — my camera e-commerce, internal tools, bots, scrapers, this website — has been done together with AI coding agents inside OpenCode.

I started the obvious way: one assistant, one long conversation, everything in it. It works, but it's like asking the head chef to also peel every potato. Expensive, slow, and the chef is too busy peeling to taste the soup.

So I reorganised it like a professional kitchen — a brigade. AI models come in sizes; from Anthropic, think Haiku (small and cheap), Sonnet (mid-size) and Opus (the top tier). Each size gets a job that fits it:

Kitchen role AI role What it does
Head chef The director — my main conversation, a top model Plans, splits the work, tastes the result, ships it
Line cook Implementer — a mid-size model Cooks most of the code from a clear order
Senior cook Senior implementer — a top model Delicate dishes, or when the line cook burned it
Junior cook Small implementer — a small model Follows an exact recipe card, nothing more
Runner Reader — the cheapest model Goes through the pantry, reads every label, comes back with a short list
Visiting chef Reviewer — a model from another company Tastes and comments. Never touches the stove

The numbers behind it

I keep a small dashboard that reads OpenCode's local activity database (a side project I call katas). From April 6 to October 3, 2026:

  • 1,337 conversations with the head chef.
  • 2,631 helper sessions — cooks called in for one specific job.
  • 28.2 billion tokens in total. A token is roughly three-quarters of a word.
  • 93.8% of those tokens were the cache being re-read. 5.4% was text stored in the cache for the first time. Brand-new input was just 0.43%, and the models' own writing 0.36%.
  • The helpers used about a quarter of all tokens, but wrote about half of everything the system produced.

So the huge total is mostly the same text being read again and again — cheaply. Keep that in mind; it's the heart of this post.

Rule 1: the head chef doesn't peel potatoes

Calling a cook isn't free. Before they start, someone has to explain the order: what we're making, where things are, how it should look. When I measured it, that briefing came to 28,000–49,000 tokens per helper (median) before it did any work — and the helper re-reads it at every one of its steps, about 21 on average.

Analogy: You don't call a cook over from another station to boil one egg. By the time you've explained where the pots are, you could have boiled it yourself.

I learned this the expensive way. In July my head chef was calling cooks for everything — on the busiest days, 100 to 270 helpers, with peak days re-reading almost a billion tokens of cache. On July 29 I measured the briefing cost, trimmed the instructions every helper has to read, and wrote one rule: the head chef only hands work out when it buys something real — two dishes cooking at the same time, or a big job kept off its own counter. Small things it simply does itself.

The effect was immediate. In the four weeks before the rule, 144 conversations called 1,525 helpers — about 11 each. In the four weeks after, 150 conversations called 620 — about 4 each. About the same number of conversations — less than half the helpers.

Yes

No

No

Yes

Read a pile of files

Exact recipe card

Needs judgment

Delicate or risky

New task

Small and obvious?

Head chef does it

Worth splitting?
two at once, or a big job

What kind of work?

Runner
cheapest model

Junior cook
small model

Line cook
mid model

Senior cook
top model

Head chef tastes:
does it run? does it look right?

Serve it

How the head chef decides who does a job

Two small habits keep this cheap:

  • Give the address, not the groceries. The head chef tells a cook which files to open, never pastes their contents into the order. Pasting means paying for the same text twice.
  • Send orders out together. Independent jobs leave in one round, so three cooks work side by side instead of waiting in line.

Rule 2: the prepared station is the secret (caching)

Here's what most people don't realise. A language model has no memory between steps. Each time it takes a step — opens a file, runs a command, writes a line — it reads the entire conversation again from the top: the instructions, the files, everything said so far.

That sounds hopelessly wasteful. It would be, without caching.

Analogy: Mise en place. Before service, a good cook chops the onions, measures the spices and lines everything up. During service they don't chop again — they grab from the prepared station. The cache is the AI's prepared station: what it has already read is kept ready, and grabbing it costs a small fraction of reading it fresh.

Yes

No

Next step

Has this beginning
been read before?

Grab it from the cache
fast and cheap

Read it fresh
full price

Store it in the cache
small one-time extra

Read only what's new:
the latest message
or the file just opened

Write the answer

Why most of the "reading" is nearly free

Overall, about 94% of everything the models read came from the cache. For the helpers, with their short and focused orders, it was almost 96%. The kitchen rules that protect it:

  • Don't rearrange the station mid-service. The cache only works on the beginning of the conversation. Change the instructions halfway and everything after them must be prepped again.
  • Keep the station clean. A long, wandering conversation is a station buried in leftovers; every step means digging through it again. One clear job per helper keeps it tidy.
  • Prepped onions don't wait forever. The cache expires after a few idle minutes, so steady work is cheaper than stop-and-go.

Rule 3: escalate once, never loop

Cheap cooks fail sometimes. The rule is strict: one step up, once. If that fails too, the head chef stops and rethinks the plan instead of trying harder.

works

fails

works

fails

Cheaper cook tries

Done

One step up
tries once

Head chef
rethinks the plan

One escalation, then the head chef steps in

The point isn't saving a few tokens on one dish. It's avoiding the classic disaster: an agent retrying the same broken thing twenty times, re-reading the whole conversation on every try.

It happened while I was writing this post. The cheapest reader wouldn't start because of a configuration problem on my machine. Nobody retried it. The job went one step up — four mid-size readers went through my projects' documentation in parallel — and the configuration problem went on the to-fix list. A slightly more expensive afternoon, but no loop.

Rule 4: a visiting chef tastes, but never cooks

Only one family of models is allowed to write code. Models from a different company may only read: review a plan, check a deploy, point out risks. And I call them in myself — the head chef can't invite them.

Analogy: A good kitchen invites a chef from another restaurant to taste the new menu. A fresh palate catches what everyone in the kitchen has stopped noticing — and because the guest never touches the stove, there's never any doubt about who changed the recipe.

What I'd tell someone starting out

  1. Give each model a role, and write the roles down. Mine live in one short file the agents read at the start of every session.
  2. Hand out work only for two jobs at once, or to keep a big job off your counter. Every helper pays a briefing cost.
  3. Protect the cache. Stable instructions at the top, focused sessions, file names instead of file contents.
  4. Escalate once. Loops are where budgets go to die.
  5. Measure it. I only learned what a helper really costs — and that almost all of my tokens were the cache being re-read — because I built a dashboard. Before that, I was guessing.

The real win of running it like a kitchen isn't even the savings. It's that the head chef — the expensive, careful model — spends its time on what it does best: deciding what to cook, and tasting every plate before it leaves the kitchen. That's how 28 billion tokens turn into finished work instead of a very expensive mess.

More in AI Agents

This is the only post here so far — browse the section or all posts.