Skip to content
Security & Trust Talking point

Agents can report a task done when it silently failed, so someone must read the work

Laurie Voss of Arize AI says agent failures can look like success unless someone reads the inputs, outputs and traces.

W
WebPulse Newsroom
AI-assisted · 2 min read
Share on X LinkedIn
Agents can report a task done when it silently failed, so someone must read the work
In brief
  • Laurie Voss described an agent taking correct steps on the wrong subject, and showed another saying a report was saved when the save failed.
  • Voss said the wrong-subject error appears when someone reads inputs and outputs, and the silent write only shows in the trace. Evals can then check for both.

Laurie Voss of Arize AI argued on the AI Engineer show that AI agents can fail in ways that look like success. An agent can take sensible steps on the wrong subject. It can also say a task is finished when it silently failed. Catching this takes someone reading what went into and came out of the agent. For the silent failure, it takes reading the agent's step-by-step record, known as a trace.

What was said

Voss gave two examples, one hypothetical and one real. The hypothetical: a user asks about Tesla, the car company. The agent searches the web and finds material on an 18th-century inventor. It then writes a polished report about the inventor and hands it to the boss. "Nothing that it did there was wrong," Voss said. Each step followed the instructions. Voss called this worse than an obvious failure. He said no unit test would have caught it. The error shows only if someone reads the inputs and outputs.

The real one came from Voss's own demo of a financial research agent. The agent tried to save its report as a file, but the notebook it ran in had no disk. Three of the 13 reports came back as short summaries, with the full body sent nowhere. Voss put it this way: "The write silently failed, and it told us the report was saved." Voss said that pattern only shows up if you read the trace. Voss added that teams need evals, meaning automated checks, to detect an agent doing something helpful but incorrect. Another run looped through endless searches on Microsoft. It ended with a report that pointed to a file that did not exist.

Why it matters

Our reading: an agent makes many decisions, and Voss said each step raises the chance of a mistake. A team that never reads what agents take in and put out can ship confident errors. The report looks finished, so nobody looks closer.

For buyers and managers, the practical question is whether anyone reads those inputs, outputs and traces. Voss's advice was blunt: "Read your traces before you write evals." Automated checks come second. Voss also showed a cheap code check that the right stock ticker appears in the report. Voss's early demos failed that check a number of times.

The other side

The disk failure came from the demo setup, a notebook with no file system. Voss called it unexpected. It shows the pattern, not how often it happens in production.

Reading has limits. Voss said human review catches subtle problems but cannot scale to every trace. Voss admitted that in the demo the traces were not all coded by hand. Automated judges can be wrong too. A correctness judge rejected all 13 reports because it graded without the research context. Giving a faithfulness check the collected sources produced a more useful split: six faithful reports and seven unfaithful. Separately, Voss urged layering checks, from code to an AI judge to human review, since no single one catches everything. How much reading is enough was left open.

Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.

The conversation this talking point comes from

Share this insight