Skip to content
Security & Trust Talking point

Cleric's CTO says AI agents are trained to sound sure, and humans must still check them

Willem Pienaar of Cleric says models trained to answer fast often blame the first error log they find.

W
WebPulse Newsroom
AI-assisted · 2 min read
Share on X LinkedIn
Cleric's CTO says AI agents are trained to sound sure, and humans must still check them
In brief
  • Willem Pienaar of Cleric says AI models are trained to be decisive, which is the opposite of what debugging production needs.
  • He says agents often blame the first error log they find, so a human must check. Cleric says grounding in real outcomes keeps confidence honest.

Willem Pienaar, co-founder and CTO of Cleric, said that AI agents pointed at a production alert often name the first error log they find as the root cause. A person then has to check it. He spoke on the AI Engineer show.

What was said

Pienaar contrasted code with production. In code, tests and linters push back quickly. Production has no such focused signal. It is spread across teams and clusters, it keeps changing, and some failures cannot be reproduced.

He then pointed to how the models are built. In his words, "large language models are trained to almost be the opposite for what you need for a production debugger." Debugging needs doubt and many competing theories. The models, he said, are "trained to be decisive, confident, and to give you an answer as quickly as possible."

He described four common failures. The agent mistakes a symptom for the cause. It does not know what is normal. Sub-agents pass up lossy summaries. It anchors on a past incident. His example: after a spike in 500 errors, the agent confidently blamed a memory leak from a new deploy. The human had to go and check. Then it blamed a readiness probe, and the events showed otherwise.

His fix is grounding: check the agent's claim against what actually happened, such as whether the service returned to normal. Without it, he said, "if the agent is just using vibes effectively, it'll always give you a very confident answer."

Why it matters

Our reading: a confident agent multiplies risk, because the cost of a wrong answer does not disappear when an agent produces it. It lands on the on-call engineer, who must second-guess a fluent story during an incident. Speed of answer is not speed of resolution.

For leaders buying or building these tools, the question is whether the agent's confidence has been tested against outcomes. Pienaar said that with grounding, a stated 90% confidence becomes a signal you can trust. Without it, a confident tone tells you little.

The other side

This is one vendor's view, and Cleric sells the grounding approach. The excerpts show a chart of confidence against accuracy, but no figures we can state, and no independent test.

Pienaar also said this has been Cleric's focus for three years, and that the work is still moving toward catching failures before deployment. The excerpts do not show how well grounding works across very different production setups. He also described the overconfidence as partly a product of training to match human preferences, so it is something to counteract, not a bug that is simply fixed.

Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.

The conversation this talking point comes from

Share this insight