Agent failures rarely look like failures
Agent failures rarely look like failures
Agent monitoring usually answers a narrower question than the one owners think they are asking. It asks whether each step ran, not whether the result was true. The model writes "I've booked your appointment." The code wrapping the booking tool logs that the request returned a success code. The dashboard counts the conversation as resolved. Each layer reports faithfully on its own process, and none of them goes back to check the calendar. Stack enough of these reports together and you get a screen full of green ticks that measures activity, not correctness.
This is the single mechanism behind everything that follows, and it takes four recurring forms:
- The confident wrong answer. The agent replies fluently, on time, in the right tone, with a fact that is false. The log records "answered."
- The phantom action. The agent tells the customer something was done, but the tool call failed, timed out, or never fired. The transcript says "sent." The inbox says otherwise.
- The wrong-target success. The action really happened, just not the right one: a booking on the wrong date, a confirmation to an old email address, a refund against the wrong order. Every system involved reports success because, from its point of view, it succeeded.
- Drift. An agent that was correct at launch slowly falls out of step with prices, policies, and stock that changed underneath it. Nothing about its behaviour looks different from day one.
Ordinary monitoring misses all four for the same reason. It is built to catch exceptions: crashes, error codes, latency spikes, rising failure rates. None of these failures throws an exception. A wrong answer delivered in normal time is, to an alerting system, indistinguishable from a right one. Nothing crashes, so nothing alerts, and the first signal is often a customer who acted on what they were told.
What happens when "nothing found" becomes "all clear"?
The pattern usually starts with one defensive line of code. A lookup is wrapped so that if it times out or the integration returns an error, the function hands back a safe-looking default instead of stopping: an empty list, a balance of zero, a status of "available", the first slot of the day. The intent is reasonable, since one slow API should not freeze a whole conversation. But the default travels downstream looking exactly like real data, and every later step treats it as real.
In staffing and dispatch, the availability check for a candidate fails, the wrapper returns "no conflicts", and the agent offers that person for a Tuesday shift they are already working. At an optometry practice, the calendar sync never loads, the slot list defaults to the first morning opening, and the agent confirms an exam into a time already taken. In e-commerce, an order history call errors out, comes back as an empty list, and the agent tells the customer there is no record of their purchase. Each answer is wrong in a specific way, and none carries a trace of the error behind it.
The fault is that three distinct states have been collapsed into one. Failed means the system could not ask. Empty means it asked and the true answer was nothing. Unknown means it asked and got something it cannot interpret. A convention that names the fix is "unknown is not failure." Each state needs its own visible path. Empty can be stated plainly. Failed should halt the action and say so, or retry. Unknown should prompt a question or a handoff, never a guess. The cost is a few extra branches in the code. The payoff is that "nothing found" can no longer pass for "all clear".
Why do some guards check nothing?
A safeguard can be present in the code, listed in the runbook, and still have no power to stop anything. Three patterns account for most inert guards.
The first checks form instead of truth. A validator confirms that the booking reply contains a date field, that the date parses, that the email matches an address pattern, that the response is non-empty. All of that passes when the date is wrong, the address is stale, or the reply is fluent nonsense. Schema validation answers "is this shaped like an answer?" and nothing more.
The second runs too late. An overnight quality score on yesterday's transcripts can flag a bad conversation, but the customer has already driven to the appointment. A post-hoc check is an audit. It is useful for finding patterns, but it cannot be counted as a guard on the action it reviews.
The third is wired so it cannot fail. The check sits inside an exception handler that logs a warning and carries on. Its threshold was loosened during a noisy launch week and never tightened. Its failure branch returns the same value as its success branch. It runs on every request and blocks none.
Put each guard through three questions. First, has it ever fired in production? A guard with a spotless record over months of live traffic is more likely broken than flawless. Second, what happens when you feed it a known-bad input on purpose, such as a booking for a slot already taken? If it passes, it is decoration. Third, does the agent being checked also do the checking? The principle here is called verification independence. A model grading its own output shares the blind spots that produced the error. The check has to come from somewhere else: a direct read of the source system, a deterministic rule, or a separate model.
Why does an agent answer confidently from stale knowledge?
An agent answers from a snapshot: the documents it was given, the price list loaded at setup, last quarter's filing deadlines, a property listing exported on onboarding day. None of those records carries a line saying when it stopped being true. A rent figure from a listing that was reduced last week sits in the knowledge base looking exactly as authoritative as the day it was added, and the agent has no way to tell the difference.
Confidence comes from somewhere else entirely. Tone, grammar, and fluency are produced by the language model, and they are identical whether the underlying fact is an hour old or a year old. A retrieved passage that closely matches the question produces a smooth, specific answer. Nothing in that process asks how recent the passage is, so a wrong answer and a right one arrive in the same voice.
The fix starts by sorting facts into two kinds. Slow facts, such as how to prepare for an eye exam, what a service includes, or how returns work, can safely live in a knowledge base. Fast facts, such as price, stock, open slots, whether a unit is still available, or this month's deadline, must be fetched from the source system at the moment of answering. The rule is query before you act: if a fact can change between setup and the conversation, read it live.
When the live read is impossible, the agent needs a permitted exit. An agent built to always produce an answer will fill the gap from the snapshot, because that is the only material it has. One built to say "I don't know" or to hand off to a person trades a slower reply for not misleading anyone.
Finally, every stored fact needs a named owner and a review date. Without both, nobody notices when it expires.
How can work be generated, logged, and never delivered?
The word "done" usually gets written at the wrong moment. The agent drafts a follow-up email, hands it to the sending service, and the log line fires as soon as the hand-off returns. A CRM update is marked complete when the request leaves. An invoice is placed in a queue and reported as issued. Each event is real, but each records generation, not arrival. Between the two sit a mail server that can bounce, a CRM that can reject a field, and a queue worker that can stall overnight. By then the log has already moved on.
The gap gets worse when a failure is ambiguous. A write times out. Did it land? The agent cannot tell, so the simplest code path retries. If the first attempt actually succeeded, the customer now has two follow-ups, the CRM has two contact records, and the client has two invoices. This is retry-as-write: a retry meant to recover from failure performs the action a second time. Timeouts are where it bites hardest, because a timeout means the reply never came back, not that the work never happened.
Closing the gap takes three changes. First, verify at the destination. After acting, read the target system: check the provider's delivery status, fetch the CRM record and compare the field, confirm the invoice exists with the right amount. This is backend state verification, and only its result should write "done". Second, change what you count. Track confirmed deliveries next to generations and watch the difference between them. A widening gap is the earliest warning you will get. Third, make every retry check state first. Before resending, look for the record. Attach a unique key to each action so the destination can refuse a duplicate even when the check is skipped.