AI Agents

How I use several AI models to code without the bill blowing up

Nicolás Biondi

In short: One expensive model plans and reviews; cheaper models do a good share of the work. And almost everything they read comes from a cache, which costs a fraction of reading it from scratch.

Between April and October my AI agents processed about 28 billion tokens. That sounds like a fortune, but 94% of it was text the models had already read before. What keeps the bill under control is four rules about who does what.

I build almost all of LENZ's software with AI agents in OpenCode: the online store, the repair shop's internal tools, a few bots and this website. I organize them the way I organize the store:

In the store In the AI setup What it does
Store manager Top model (Opus) Plans, hands out the work and reviews before delivery
Experienced salesperson Mid model (Sonnet) Does the work from a clear brief
Senior tech Top model The delicate jobs, or what the salesperson couldn't solve
Intern Small model Reads stacks of documents and brings back a summary
Outside auditor Model from another company Reviews and gives an opinion. Never touches the code

Where 28 billion tokens come from

From April 6 to October 3, 2026, my dashboard (the panel where I track usage) counted 1,337 conversations with the manager and 2,631 helper sessions. A token is about three quarters of a word.

94% of all that was the cache read again, and only 0.43% was new text. The helpers used about a quarter of the tokens, but they wrote close to half of the code.

Rule 1: don't send the senior tech to change a battery

Every helper starts with an explanation of the job: the instructions, the tools and the task. I measured it: that's between 28,000 and 49,000 tokens before it does anything, and it reads them again at each of its 21 or so steps.

So the manager only hands work out when it buys real time (two tasks at once) or when one big task would fill up his own conversation. The small stuff he does himself. The delicate work, or whatever the helper couldn't finish, goes to the senior tech.

No

Yes

Fails

New task

Hand it out?

Manager
does it

Intern reads or
salesperson works

Senior tech

Manager
reviews

Who takes each task

In the four weeks before July 29, each conversation called about 11 helpers. That day I wrote this rule down and trimmed the opening explanation. The next four weeks dropped to around 8, and the four after that to around 2.

Rule 2: keep what's already been read within reach

An AI model has no memory between one step and the next: at every step it reads the whole conversation again from the top. The cache is what makes that cheap. What it already read stays stored, like the catalogue a salesperson keeps in his head, and using it again costs a fraction.

Yes

No

Next step

Read this
opening before?

Takes it from cache

Reads at full price
and stores it

Reads only new text

Answers

Why re-reading comes cheap

The cache only works if the start of the conversation doesn't change, so I keep the instructions at the top fixed and give each helper a single job. It also expires after a few minutes without use, so working straight through costs less than starting and stopping.

Rule 3: you escalate once

If a cheap model fails, the task moves up one level and gets one more attempt. If it fails there too, the manager changes the plan. Nobody repeats the same broken step over and over, because every attempt reads everything again.

Rule 4: the audit comes from outside

Only one family of models writes code. The ones from other companies read and comment, and I'm the one who calls them in. Whoever didn't write the code has nothing to defend.

What I'd do if you're starting now

  1. Write down in a short file what each model does
  2. Hand out work only to run things in parallel or to isolate big tasks
  3. Don't change the start of the conversation, so the cache keeps working
  4. Allow one escalation, never a loop
  5. Measure your usage: I only saw these numbers once I built a dashboard

More in AI Agents

This is the only post here so far. Browse the section or all posts.