- Abhi Arya of Reducto argued that agents rarely crash when they are wrong. They return confident results that nobody can easily check.
- His test for agent products is who finds out about a wrong answer and how fast. In auto mode, he said, the builder pays when errors reach customers.
On the AI Engineer show, Abhi Arya of Reducto's product team argued that agents are most dangerous when they fail without noise. A wrong answer can look fine. His test for any agent product is simple: when the agent is confidently wrong, who finds out, and how quickly?
What was said
Arya said building software for agents starts with deciding who bears the cost of the agent's mistakes. In what he called auto mode, the agent runs free on a prompt. Its errors go unseen, he said, and the builder pays through its customers and users.
He first built an MCP server, a standard way to let an AI agent call a company's tools, with one tool for every endpoint in Reducto's product. In a curated demo it worked. He said that demo told him nothing.
The real test came when the whole team used it. People fed it half-formed notes from customer calls. The agent built whole workflows and confidently said all was well. By his account, nobody, human or agent, could say what the finished workflow really did.
He then scrapped that code in a single commit. The replacement had fewer tools, and each one encoded how the work is really done. He also added a snapshot tool that returns the live state of a pipeline, including specific error codes.
Then came the runtime problem. A document pipeline can run to the end without any error, yet still miss something. In Arya's words, "And with agents, it goes wrong quietly." His fix was to make uncertainty visible. Reducto's extraction outputs confidence scores, and its evals penalize the agent for shaky guesses and reward it for asking. He closed by asking: "when your agent is really confidently wrong, who finds out and how fast do they find out?"
Why it matters
Our reading: this gives buyers a question that a demo cannot fake. Ask a vendor what happens after the agent gets something wrong. Does the product show its doubt, explain what it did, and leave a human able to check?
For builders, the lesson is where to spend effort. Arya said the fix was architecture, not a longer prompt. We add one inference that Arya did not state: the longer an error stays hidden, the more people it may reach before anyone sees it.
The other side
This is one company's account, given in a conference talk by someone promoting his own product. Arya gave no figures on how much faster errors surface after the rebuild, or what human review costs.
His method also fits his field. Document extraction can produce confidence scores and bounding boxes, which show where on a page a value came from. Many agent tasks offer no such signal. Arya conceded that fixing the build stage did not finish the job, because the problem returned at runtime. Making doubt visible still leaves a person to review the work, and he left open how often people will actually do it.
Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.
The conversation this talking point comes from
- AI Engineer: Why We Deleted Our MCP Server and Rebuilt It — Abhi Arya, Reducto (2026-10-08)





