Skip to content
Security & Trust Talking point

A passing AI test proves little until the judge is checked against people

Tejas Kumar showed a keyword test and an AI judge both passing bad customer-support answers.

W
WebPulse Newsroom
AI-assisted · 2 min read
Share on X LinkedIn
A passing AI test proves little until the judge is checked against people
In brief
  • Tejas Kumar showed a keyword test and an AI judge both passing wrong customer-support answers, because each rewarded the look of a good answer.
  • His remedy was to give the judge the real policy and score it against about 30 human verdicts before trusting its green results.

A passing test does not show that an AI system gives right answers. Tejas Kumar of IBM made this point on the AI Engineer show, in an episode about evals, the tests teams write to check AI behaviour. He argued that keyword checks and AI judges can both approve wrong answers. A judge earns trust only after it is checked against human verdicts.

What was said

Kumar built a support-bot test live. A customer wants to return headphones bought 20 days ago, but the store allows 14 days. His first check looked for the word cannot in the answer. An AI writes the answers, so a reply about not paying for shipping can contain that word too. Kumar said of that result: "And now that absolute nonsense is green."

He then switched to an AI judge, which brings its own biases. He said judges tend to pick the friendliest answer, because they are set up as helpful assistants. Asked about the headphones, a judge will lean toward telling the customer they can return them, even though policy says no. Judges also favour text from their own model family. In his demo, the judge kept picking a wrong answer written by a GPT model.

His fixes were practical. Put the real return policy in the judge's prompt. Use a judge from a different model family than the one writing answers. Then build a set of about 30 scenarios with human pass or fail verdicts, hidden from the judge, and aim for roughly 80% agreement. His summary: "Just because it's green doesn't mean it works."

Why it matters

Our reading: a green result is a claim about the test, not about the product. If the check passes a bot that approves out-of-policy returns, the store absorbs the refunds. If it passes a garbled reply, the customer receives it. Nobody sees the cost on a dashboard that stays green.

For buyers, the useful question to a vendor is how often the judge agreed with human verdicts, and on how many examples. For builders, the judge is a component that needs its own tests. Kumar also called the work of checking disagreements and supplying context "literally the work of building your judge."

The other side

Kumar's demo was small and run live. It used about 30 scenarios with human verdicts he wrote by hand, and a cheap model (GPT 3.5 Turbo) as the judge. He noted that position bias is especially common in cheaper models. Even with the policy added, six cases failed, so agreement was still under his 80% bar.

The bar itself is his choice. On 30 cases, it tolerates a few wrong verdicts. He also warned that 100% agreement is a red flag, a sign of overfitting. The excerpts do not show what to do when human reviewers disagree with each other. Nor do they say how often a judge must be rechecked once real customers arrive.

Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.

The conversation this talking point comes from

Share this insight