Skip to content
The AI-First Web Talking point

Cleric's CTO argues ops agents' confidence needs checking against outcomes

Cleric's CTO says an agent's confidence means little until it is tested against whether fixes worked.

W
WebPulse Newsroom
AI-assisted · 2 min read
Share on X LinkedIn
Cleric's CTO argues ops agents' confidence needs checking against outcomes
In brief
  • Willem Pienaar of Cleric argued that an operations agent's stated confidence only means something once it is checked against whether fixes worked and services recovered.
  • Running several theories at once burns far more tokens than a single coding-agent session, Pienaar said, but keeps engineers out of the loop longer. He gave no figure for how that trades against engineer time.

Willem Pienaar, co-founder and CTO of Cleric, said on the AI Engineer show that people should judge an operations agent by what happens after its advice, not by its manner. Check whether engineers applied its fix and whether the service recovered. Then, he argued, the confidence the agent states becomes a signal you can rely on.

What was said

Pienaar's starting point is that language models are built to answer fast and sound certain. He said they are "trained to be decisive, confident, and to give you an answer as quickly as possible." Debugging production needs the opposite: doubt.

Most teams use agents only to diagnose a problem. Cleric also checks what happens next. It watches whether the engineer merges code that matches the diagnosis. Then it looks for metrics, logs and other state to return to their usual baseline.

One signal is not enough. Muting an alert also makes things look normal, so Cleric compares several signals against what it expects.

Pienaar showed a chart of stated confidence against actual accuracy. A well-calibrated agent sits on the diagonal. Without grounding, he said, "if the agent is just using vibes effectively, it'll always give you a very confident answer." With grounding, an agent that says it is 90% sure can be believed. The chart is Cleric's own, and he showed it without figures.

He was open about the price. Cleric's runs cost far more than a single coding-agent session. Testing several theories at once burns more tokens, but it keeps the human out of the loop longer.

Why it matters

Our reading: the token bill is easy to see, while the cost of a confident wrong answer is hidden. That cost lands on the on-call engineer, who chases a false cause, and on the users waiting for recovery.

For buyers, the practical test is evidence. Ask a vendor whether it records if fixes were applied and services recovered, and whether its confidence figures match those results. A smooth demo shows tone, not calibration.

For teams building their own agents, the lesson is to keep an outcome record. It is the only way to learn whether "90% sure" has ever been true.

The other side

The calibration result comes from the vendor itself. The excerpts give no numbers, sample size or independent check, so it is one company's account.

Pienaar did not put a figure on the trade between token cost and human time. The case that grounding pays for itself is plausible, but the excerpts do not prove it.

He also conceded that recovery is imperfect proof. A muted alert looks like a fix, which is why Cleric needs several signals. The method also depends on seeing whether an engineer actually applied the suggested change, and on a trustworthy baseline for what normal looks like.

Those conditions may not hold in every environment, and the talk did not say how often they fail.

Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.

The conversation this talking point comes from

Share this insight