I found the problem while looking for something else, which by now is my favourite diagnostic method: ten days of repair intakes without a single alert in the internal WhatsApp group, and nobody on the team had noticed. Nobody misses a channel that only informs, so those ten silent days felt like ten quiet days at the service shop.
The system is supposed to post an alert to the group for every unit that comes in for repair, with the case summary, the photos and the intake PDF. That way the technician knows what's coming before the camera reaches his bench, and sales can answer when the client calls to ask. Without the alert, the data sits in the system and nobody looks at it.
The new token went out and the shop kept the old one
We rotated the token on the messaging side, the credential that authorises a system to send messages through the WhatsApp API. The shop system kept presenting the previous one, so every attempt got a 401, the response an API uses to say that credential is no good. That was 68 failed sends, 9 of them intakes that were never announced.
The system wrote every rejection into a log nobody opens, except when they're already hunting a problem. It was a doorbell with a cut cable: nobody complains about the silence and you assume nobody came.
A log with no reader is not monitoring
Change my mind
The deploy was the obvious culprit
I thought the sending code had broken in a deploy, because we'd touched that part not long before. The code was unchanged.
We confirmed it with a read-only endpoint, a query that returns 200 if the credential works and 401 if it doesn't, without posting any message. It answered 401. Testing it with a real send would have put a duplicate alert in the group.
Resending the 9 alerts without inventing the format
With the token updated the new alerts went out again, but the 9 intakes from those ten days were still unannounced. We wrote a resend script that reuses the production functions instead of rebuilding the messages, because an alert in a different format earns nobody's trust.
We gave it three handbrakes. It runs in dry-run by default, which prints what it would send without sending anything, and it refuses to start if the system isn't running in its usual mode. It also skips units already delivered, because announcing the intake of a camera the client picked up a week ago only confuses people. We re-announced the 9 intakes.
Two weeks later, the worker vanished
The messaging platform migrated and the worker dedicated to the group, the process that only sent those internal alerts, got left behind without leaving a note. This time the gap was caught before a single intake went unannounced, because I already knew the channel could go silent without warning: I compared the intakes in the database against the chat history and they all matched, 0 lost events against 9 the first time.
Two doorbells for the same door, one with no owner
Underneath there were two paths doing the same thing. Messages to clients went out through the main API and the group alerts through a shortcut of their own, with its own credential and its own separate process. Two routes are two places where a token can go stale, and the second one had no owner.
We removed the shortcut. Internal alerts now go out through the same API the team uses to write to a client, and the log records the successes too so we can count them. Renewing the token closed this incident; deleting the second route should prevent the ones lined up behind it.
Two things are still open: there's no alert when the group goes a business day with no messages, and the cross-check against the chat history is still manual. Every internal channel now has an owner with a name even if it only informs, and before we touch a credential we test it with the query that posts nothing. The scoreboard reads 9 unannounced intakes the first time and 0 the second.
