Reviewing a single WhatsApp conversation with AI cost us around 33,000 tokens and 20 to 30 seconds. The client had 7,527 conversations from about two months, and we worked with an AI account that has a usage cap, not an API key billed per token. Putting all 7,527 through the model was like asking one technician at our repair desk to inspect every camera in Lima: he can start, he never finishes.
What the client wanted to know was concrete: how long they took to answer, how many people were left waiting, how many wrote in the middle of the night. My first reflex was to look for a cheaper model that could take the volume. I did the math and it didn't add up either. The project got unstuck when I accepted that almost all of those questions are answered with arithmetic.
Paying tokens for 7,527 chats
One SQL query and two hours of patience
A ground floor that spends no tokens
I pulled a read-only copy of the live database with VACUUM INTO, a SQLite instruction that writes a clean file to the side: 0.7 s, without taking the database lock away from the process that was answering chats at that moment.
On that copy we cut the history into conversations every time there were more than 6 hours of silence, and we labeled the 89,371 messages as customer, bot, human or system with local rules. An outgoing message counts as bot only if its text matches an automatic reply template from the business exactly, or if it's a photo attached less than 15 seconds after one. The rest of the outgoing traffic stays as human, and none of this calls a model.
With that we calculated first response time, waits, abandonment and messages outside business hours over the 7,527 conversations, in plain, repeatable SQL. We call that ground floor Tier 1.
43.6% ended up waiting for a human
75.7% of the conversations got a human reply at some point and 24.3% were handled by the bot alone. 57.9% started outside business hours. The median first human response was 28 seconds, which sounds fine until you cross it with the 43.6% that closed with the customer still waiting for a human.
Triage decides who gets the specialist
At the fire station we use the word triage to decide who gets treated first. Here it works like a hospital vital signs monitor: you put it on everybody because measuring a pulse costs almost nothing, and you call the expensive specialist only for the charts the monitor flagged.
The monitor is a set of signals that cost nothing: negative sentiment words, conversations with no human reply at all, waits over 30 minutes and purchase intent. We added a random sample as a control, because a list built only from bad cases confirms the suspicion you walked in with.
Triage sorts the candidates by severity and cuts the list at a limit: in this run, eight conversations from the last 7 days went up to Tier 2, the expensive floor. An AI CLI, the model running from the command line, takes the transcript labeled by role and returns strict JSON with a summary, resolved yes or no, customer sentiment, four scores from 1 to 5, a lost sale flag, findings with severity and an escalation flag. The model doesn't answer any customer and doesn't touch the chat: it hands in a chart and leaves.
What SQL couldn't see
On average, the eight scored tone 2.6 out of 5, clarity 2.1 out of 5 and resolution 2.4 out of 5. 37.5% came back flagged as a lost sale and half as a case to escalate.
Those eight turned up three things no SQL query would have found: one case of mistreating third parties, a sensitive personal detail written out in full inside the chat, and deliveries arranged by the bot alone, with no human confirming anything.
The 4 second median I don't believe
Attribution by templates is a heuristic and it's written down as one in the project README. In the shortest window we validated, the median human response came out at 4 seconds, and to me that means some replies from people are being counted as bot. We left it noted as a Tier 1 limitation, in the same report where the number appears.
Deciding how to mask personal data in the transcripts is still pending, and we have to solve it before Tier 2 goes to production.
Before looking for a cheaper model, I now work out how much of the answer comes from arithmetic, and every cheap heuristic travels with its margin of error in the same report that uses it. There's still one technician, but I no longer send him every camera in Lima: I send him eight.
