Skip to content
The AI-First Web

Fewer respondents at agent-deploying firms allow or plan unreviewed pushes

VentureBeat survey: 56% of respondents at agent-deploying firms allow or plan it, down from 75% in July

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Fewer respondents at agent-deploying firms allow or plan unreviewed pushes
In brief
  • Among respondents at firms deploying AI agents, a VentureBeat survey found 56% allow unreviewed production pushes or are building toward it, down from 75% in July.
  • Among all respondents who run pre-release evaluations, 61% reported that in the past 12 months an agent or feature passed its tests and then failed customers.
  • Leaders should ask which agents ship without a person, and whether monitoring flags wrong answers or only speed, errors and cost.

A passing test is not a second opinion

An automated test can only vouch for cases someone thought to write. A human reviewer may notice the case nobody thought of. That difference is the backdrop to a shift in a new survey. The survey does not say why views moved.

VentureBeat Intelligence polled people at companies with at least 100 employees in August and compared the answers with July. This story covers respondents whose companies deploy autonomous AI agents. That group had 118 people in August and 96 in July.

In August, 32% said their company already lets agents push specific low-risk changes to production without human review. Another 24% said they were building pipelines to allow it within 12 months. Together that is 56%, down from 75% in July. VentureBeat calls the drop a real change.

The opposite camp grew as well. In July, 20% of these respondents expected a person to stay in the approval loop for production changes. By August that was 42%, more than double. VentureBeat also calls that a real change.

56%
Allow unreviewed agent pushes or are building toward it (32% allow now, 24% building)
Source: VentureBeat Intelligence, VB Pulse survey (October 5, 2026); 75% in July
42%
Expect people to keep reviewing production changes
Source: VentureBeat Intelligence, VB Pulse survey (October 5, 2026); 20% in July

How a passing test still fails customers

Here, an evaluation means a scored test run on an agent's work before release. If the score passes, software can ship the change with no person involved. The limit is coverage. A test set holds only the cases someone wrote down. Real customers ask for other things.

The survey shows the gap in practice. Set aside the 5% of respondents with no pre-release testing. Of the rest, 61% reported that in the past 12 months an AI agent or language-model feature cleared their tests, went live and then failed customers. This base covers all respondents who run evaluations, not only agent deployers. About 22% said this happened more than once.

VentureBeat adds a caution. The question asks whether a company had any such failure, not how many per agent. So the figure counts affected companies. A firm with a large fleet of agents has more exposure than a firm with a handful.

61%
Ran evaluations, then shipped a test-passing feature that failed customers (past 12 months)
Source: VentureBeat Intelligence, VB Pulse survey (October 5, 2026)

Monitoring that watches the machine, not the answer

Catching a wrong answer after release needs a different kind of monitoring. Health metrics tell a team the agent was running, how quickly it replied and what it cost. Quality checks ask whether the answer was right.

Among agent-deploying respondents, only 29% said live, automated quality checks are the core of their production monitoring. A larger group, 53%, mainly use trace logs or gateway tracking. Those options were worded around activity, token counts, latency, errors and cost, not correctness. Each person could pick one approach, so some of the 53% may also run quality checks.

A dashboard can show all green while customers get wrong answers. The health numbers do not read the answer.

29%
Monitoring built around real-time quality checks (respondents at firms deploying autonomous agents)
Source: VentureBeat Intelligence, VB Pulse survey (October 5, 2026)

Fewer expect to ship unreviewed, but the survey does not say why

It is tempting to link the shift to failures. The survey does not support that link. Among agent-deploying respondents, those who reported a test-passing failure and those who reported none gave similar answers on unreviewed pushes. VentureBeat calls the gap "too close to call."

In short, the survey found no clear link between failures and willingness to ship unreviewed. The samples are modest. That is not proof that no link exists.

Final purchasing decision-makers moved most. Their combined figure fell from 88% in July to 61% in August. Among other respondents it fell from 63% to 48%, a drop VentureBeat calls "too close to call" on its own. VentureBeat did not test whether the two drops differ.

Trust in automated checks is low. Asked what most limits their trust, only 9% chose "we trust automated evaluation today." Another 27% named weak alignment with real-world outcomes.

Spending plans are split. Asked which reliability spending will grow most next year, 30% named human review workflows and 26% production observability tooling. VentureBeat calls that gap "too close to call." Another 21% named automated evaluation pipelines.

Questions to put to your team

First, which agents may ship changes with no person looking, and who decided they were low-risk? Second, when did a passing test last miss a real customer failure, and did that case become a new test?

Third, does your monitoring alert when an answer is wrong, or only when speed, errors or cost move? Fourth, what happens to reviewer workload as agents multiply? Human review catches problems that tests miss. VentureBeat points to the trade-off: each added agent and each added change means more review work to pay for.

A passing score shows only that the test was passed. Whether the customer was served is a separate question, and someone still has to ask it.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: VentureBeat.

Share this insight