- In VentureBeat's August survey, 56% of agent-deploying respondents allow or are building toward unreviewed production changes, down from 75% in July. The survey does not show why.
- Only 29% of agent-deploying respondents mainly use real-time quality checks on live outputs, so wrong answers may go unseen by infrastructure monitoring.
- Before removing human approval, ask what checks output quality after release and how a wrong answer reaches a person.
A passing test says a change was ready to ship. It does not say the change stays right once customers use it. A VentureBeat survey suggests that fewer respondents whose organizations use AI agents want tests alone to approve production changes. The survey does not say why. It does suggest a gap that matters either way: many respondents' primary monitoring is built to show that an agent ran, not that it was right.
What the survey found
VentureBeat's VB Intelligence group surveyed 140 respondents in August. Some of their organizations deploy autonomous AI agents. Among those, 56% said they already let automated test results approve some production changes, or are building toward that within a year. In July the figure was 75%.
The mix of respondents changed between waves, which could skew the result. Final AI purchasing decision-makers made up 53% of August respondents, against 44% in July. Yet the drop also appeared inside that group, from 88% to 61%.
The limits are clear. VentureBeat says the two surveys drew on separate, self-selected groups of its readers and panel members. The results do not necessarily show a change across the whole enterprise market. They also do not show what caused the shift.
Using the tests, doubting the verdict
The clearest change is a split. Among agent-deploying respondents, the share expecting to keep human approval for the foreseeable future rose from 20% to 42%. VentureBeat calls that rise statistically significant.
Over the same period, OpenAI's developer platform appeared in 59% of stacks, up from 31%. VentureBeat calls this the strongest month-over-month vendor finding. Use of that evaluation tooling grew, yet more respondents want a person to sign off. VentureBeat reads this as a difference between using tests to assess agents and trusting them as the only approval.
The numbers offer one possible reason for the hesitation, though the survey does not show the cause. Among respondents who run pre-deployment evaluations, 61% reported at least one case in the past 12 months. In each case an AI feature passed internal testing and then failed a customer. That includes 22% who reported more than one such failure. Another 39% reported exactly one.
The top stated objection to automated evaluation, at 27%, was that it does not match real-world outcomes.
Two cautions apply. The August figure was not statistically different from July's comparable 53%. And the survey counts organizations with at least one failure, not how often any one agent fails. A firm running hundreds of AI features has more chances to hit one.
The mechanism: a clean log, a wrong answer
Pre-deployment testing is a gate at the door. Production monitoring is what you see once the work is live. The survey suggests that, as a primary method, the second is less common than the first.
Just 29% of the 118 agent-deploying respondents named automated checks on live answers, run in real time, as their main monitoring method. Another 36% mainly rely on transaction trace logging. That keeps a record of each request, including token counts, for debugging later. It does not necessarily judge whether the answer was right.
That matters because of how software reports success. VentureBeat points out that an agent can answer fast, sound sure and be wrong. The log still looks clean and the status code reads 200, which means only that the request finished.
Nothing crashes, so an infrastructure alert may never fire. The survey does not say how these failures are found. But without live quality checks, the person who finds the error may be a customer.
The gap is similar where agents already ship unreviewed changes. Only 10 of the 38 respondents who allow such changes named live quality checks as their main approach, about 26%. The question allowed one answer. Those organizations may run other checks the survey did not capture.
Experience did not predict the plan
One finding cuts against the obvious story. Among agent-deploying August respondents who had seen an evaluation-passing failure, 59% allow unreviewed changes or are building toward them. Among those who had not, it was 55%. VentureBeat calls that gap too small to show a relationship.
In July the pattern ran the other way. Firms with such failures were more likely to pursue zero-human deployment. That link did not hold up clearly in August.
So the survey does not show what drives these decisions. Incident history did not clearly predict autonomy plans. One hypothesis, ours and not the survey's, is that ambition and cost pressure count for more than past failures. The data cannot test that.
What to ask your team
Start with the changes an agent can ship today without a person. Ask who defined 'low risk', and when that definition was last checked against a real incident.
Next, ask what happens after release. Does anything confirm that an agent's output is correct, or only that it ran? If the answer is trace logs and uptime dashboards, a polished wrong answer can pass through unseen.
Then ask how a wrong output reaches a human. If the honest answer is 'a customer emails us', that is your monitoring plan. Fund live quality checks before you remove the approval step, not after.
A test can open the gate. Only monitoring can tell you what walked through it.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: VentureBeat.





