Smallest failing receipt in the article's example: 2 lines (Source: InfoWorld, Isaac Sacolick (September 30, 2026))
You cannot grade an exam without an answer key. That is the problem facing anyone who tests AI. The system can give different answers to the same question, and often nobody can say which answer is right. One way around this is to stop asking whether the answer is correct. Ask instead whether the answer holds steady when the meaning stays the same.
That idea is called property-based testing, or PBT. Isaac Sacolick describes it in a September 30, 2026 InfoWorld analysis, drawing on comments from vendors and consultants. The article is one practitioner's survey, not a controlled study.
The technique has a long history. Sacolick traces the name to research published in 1997 and the method to a paper from 2000. His article's focus is applying it to AI systems whose outputs vary from run to run. He calls it “a promising testing approach” and “only one validation step.”
Why fixed answers break down
Traditional software testing is simple. You write an input, state the expected output and check for a match. Sacolick notes that AI models and agents are stochastic, meaning identical prompts can produce different results. A single expected answer then tells you little.
His example is a language model that sorts expense receipt lines into categories. With few categories and similar receipts, he says traditional methods can work. With hundreds of categories and countless receipt formats, they do not.
Some lines also have no clear right answer. A line reading “Marketplace $97.63” on a hotel bill could be hotel, meals or grocery. Testing researchers call this the oracle problem: there is no practical way to know the correct classification.
How the test works
PBT sidesteps the problem. You state a property, such as “each line item is assigned to one category,” and list the allowed categories. A generator is the part that produces the test inputs.
In the article's example, the team builds hundreds of receipts and runs each several times to gauge how consistent the model is. Next, each receipt is rewritten many times in ways that leave its meaning unchanged. Lines are reordered, amounts are formatted differently, or the guest name changes.
If the same line lands in two categories across those variants, at least one result is wrong. You have found a defect without ever learning the right answer.
The tool then shrinks the failing case. It drops line items, reduces amounts and shortens descriptions. It keeps only the simplifications that still fail. The result is the smallest input that still breaks.
In the illustration, a folio might shrink to “Room Charge $289.00” and “Marketplace $97.63.” Swapping their order flips Marketplace from hotel to grocery. A model showing this pattern would be one whose category for a line depends on what precedes it. A team can act on a finding that specific.
From receipts to agents
The same logic applies to AI agents that act, not just classify. Sanmi Koyejo of Virtue AI gives two examples of properties. Paraphrasing a prompt should not flip a safety verdict. An agent should not cross a permission boundary, however a request is worded.
Harshil Shah of R Systems, RSI, says teams getting real value from PBT on agents make four moves. They use an LLM or domain model to generate realistic variations of intent. They write rules about the full run of steps an agent takes, rather than judging only what it says at the end. They assert rules such as a write tool not being called before a validation tool. And they run the suite in continuous integration, repeating each case enough times to be statistically meaningful, and fail the build when pass rates fall.
What it cannot do
The article is frank about limits. Koyejo says PBT “only checks properties you thought to write.” Failure modes nobody anticipated slip through. He suggests pairing it with adversarial red teaming, where testers actively try to break the system.
Cost is another limit. LLM-based generators cost money. Teams must balance enough variants to trust the result against the time and spend to run them. Bo Li, also of Virtue AI, adds that real validation means watching agents in realistic settings with live tools and changing permissions.
Scale is a third concern. Joseph Hurley of Digital.ai says proving a property across a few functions is different from proving quality standards across an entire AI-generated system.
Sacolick also says teams should set release-ready criteria for agents, with further steps for security, business value and AI cost. He closes by expecting PBT, alongside other testing methods, to matter for organizations building reliable agents. That is his opinion, not a measured finding.
The decision this forces
This shows where one kind of risk sits. The defect in the receipt example was not a wrong answer. It was an unstable one. A system that looks right in a demo but shifts under rewording can pass a scripted test and still disappoint a customer or an auditor.
It also moves authorship. Engineers can build the generators, but the properties are business rules, such as who may approve what and what must reconcile with what. EY's Sridhar V. Dasaratha calls it a best practice to write properties in business terms, so they can be checked independently and tracked over time.
Leaders can ask their teams four things:
1. Which properties have we written for each AI system, and who in the business signed them off? 2. Do our tests check consistency under meaning-preserving changes, or only a fixed set of example answers? 3. For agents, do we test the whole sequence of actions, including permission boundaries? 4. What do we pair this with to catch the failures nobody thought to write down?
A test that needs no answer key still depends on the rules someone thought to write.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: InfoWorld.





