BlogAI Agents

One sign per agent

Four failures of the shared switch ended in per-agent permits with an expiry

Nicolás Biondi

A dim doorway with a blank hanger on the handle and three glowing duplicates fading beside it.

Ten WhatsApp messages went out to the internal group in two minutes, all of them announcing repairs we had delivered weeks earlier. The system forgets nothing, and it picks the worst moment to remember. It was the fourth time.

There was no new bug: the usual rule, "turn notifications off before you test and set them back to normal when you're done", breaks on its own as soon as two agents work at the same time.

Tests run against real customer orders

The after-sales module at LENZ Service runs against the production database, the one with real customer orders. When an AI agent opens the browser to verify a status change, the system does what it always does and fires the WhatsApp message, the email and the push to the customer's phone. That's why the mute switch existed: without it, one careless test turns into an awkward phone call the next day.

With a single agent the switch works, and I almost never work with a single agent.

The sign anyone can take down

The notification mode was one field with two values: normal or off. Agent A turns it off and starts testing. Agent B, on another branch, turns it off too, finishes fifteen minutes earlier and sets it back to normal, exactly as the instruction asks. A is still running and the system has no way to know.

I blamed the agents first, sure that one of them had ignored the rule. They followed it to the letter, all four times. The field didn't store who asked for the silence or until when, so "restore" meant "cancel everyone's silence", including the silence of whoever was still inside. It was the "do not disturb" sign on a meeting room door: there was only one, and the last person out took it down even with people still working inside.

Meme Batman Slapping Robin: The agents ignored the instruction / They followed it to the letter The agents ignored the instruction They followed it to the letter

A turns notifications off

A runs tests

B turns notifications off

B finishes and restores

Normal mode, A continues

Real messages to the group

Agent B restores the shared value while A is still testing

Every agent hangs its own sign

We swapped the single value for a table of temporary permits. Each agent takes its own row with its name, the reason and an expiry time: 180 minutes by default, hard ceiling of 12 hours. It releases only its own row and can't touch anyone else's.

One function answers what the current mode is, and all three channels use it. If at least one permit is live, it returns off; if none are left, it reads the team setting. During tests we no longer touch that setting, so the restore step disappeared, which was the exact step that kept failing.

The expiry covers the agent that hangs halfway through: an orphan permit with no deadline would mute real customers forever. The read also fails closed, and if the query to the table errors out the function returns off. I'd rather have one customer wait for their "your gear is ready" notice than another batch of test notices going out to the group, or worse, to a customer.

The test database we didn't build

We could have written a sterner paragraph in the repo policy, or set up a rotation or a queue, but instructions had already lost us four times.

We didn't stand up a separate test database either, which is the textbook answer. Duplicating production with its nine services, each with its own database and its own cron jobs (tasks that run by themselves at a fixed time), is a project of weeks. One table took care of the risk.

The admin header says who hung the sign

Since the change, no test message has gone out to the internal group again. The admin header shows live who holds a permit and why, so when a notification doesn't arrive I look at that screen instead of asking in the chat. The test suite takes its own permit when it starts and releases it at the end, and the repo policy requires taking one before any browser check that can write.

The 12 hour ceiling is still open, which is a lot of silence for one forgotten permit, and nothing warns us yet when a permit is about to expire with tests running. We left it there because the bad case went from eternal silence to half a day of silence, and half a day we notice.

Each agent takes its own permit and nobody can release anyone else's. The silence lifts when the last sign comes down, and those repairs from weeks ago have nothing left to announce.

postmortem notifications concurrency workflow