How I use several AI models to code without the bill blowing up
In short: One expensive model plans and reviews; cheaper models do a good share of the work. And almost everything they read comes from a cache, which costs a fraction of reading it from scratch.
Between April and October my AI agents processed about 28 billion tokens. That sounds like a fortune, but 94% of it was text the models had already read before. What keeps the bill under control is four rules about who does what.
I build almost all of LENZ's software with AI agents in OpenCode: the online store, the repair shop's internal tools, a few bots and this website. I organize them the way I organize the store:
| In the store | In the AI setup | What it does |
|---|---|---|
| Store manager | Top model (Opus) | Plans, hands out the work and reviews before delivery |
| Experienced salesperson | Mid model (Sonnet) | Does the work from a clear brief |
| Senior tech | Top model | The delicate jobs, or what the salesperson couldn't solve |
| Intern | Small model | Reads stacks of documents and brings back a summary |
| Outside auditor | Model from another company | Reviews and gives an opinion. Never touches the code |
Where 28 billion tokens come from
From April 6 to October 3, 2026, my dashboard (the panel where I track usage) counted 1,337 conversations with the manager and 2,631 helper sessions. A token is about three quarters of a word.
94% of all that was the cache read again, and only 0.43% was new text. The helpers used about a quarter of the tokens, but they wrote close to half of the code.
Rule 1: don't send the senior tech to change a battery
Every helper starts with an explanation of the job: the instructions, the tools and the task. I measured it: that's between 28,000 and 49,000 tokens before it does anything, and it reads them again at each of its 21 or so steps.
So the manager only hands work out when it buys real time (two tasks at once) or when one big task would fill up his own conversation. The small stuff he does himself. The delicate work, or whatever the helper couldn't finish, goes to the senior tech.
In the four weeks before July 29, each conversation called about 11 helpers. That day I wrote this rule down and trimmed the opening explanation. The next four weeks dropped to around 8, and the four after that to around 2.
Rule 2: keep what's already been read within reach
An AI model has no memory between one step and the next: at every step it reads the whole conversation again from the top. The cache is what makes that cheap. What it already read stays stored, like the catalogue a salesperson keeps in his head, and using it again costs a fraction.
The cache only works if the start of the conversation doesn't change, so I keep the instructions at the top fixed and give each helper a single job. It also expires after a few minutes without use, so working straight through costs less than starting and stopping.
Rule 3: you escalate once
If a cheap model fails, the task moves up one level and gets one more attempt. If it fails there too, the manager changes the plan. Nobody repeats the same broken step over and over, because every attempt reads everything again.
Rule 4: the audit comes from outside
Only one family of models writes code. The ones from other companies read and comment, and I'm the one who calls them in. Whoever didn't write the code has nothing to defend.
What I'd do if you're starting now
- Write down in a short file what each model does
- Hand out work only to run things in parallel or to isolate big tasks
- Don't change the start of the conversation, so the cache keeps working
- Allow one escalation, never a loop
- Measure your usage: I only saw these numbers once I built a dashboard
More in AI Agents
This is the only post here so far. Browse the section or all posts.