In the authors' medical-agent test (361 claims, 40 held-out answers): experts' should-fail claims that were blocked: 138 of 139 (Source: Multiverse Computing, Hugging Face blog (September 29, 2026))
A fact can be true and still be misattributed. For an organisation that acts on what an AI agent says, the second error can do as much harm as the first. That is the argument behind a new paper from Multiverse Computing, and it changes what leaders should ask of any agent that touches sensitive data.
The error: right fact, wrong source
Multiverse Computing's researchers call the failure "cross-source conflation". A claim is true somewhere in the evidence, but the answer credits a different source.
Their first example is a customer support agent. It says, "According to the account record, this plan includes a 30-day refund window." The refund window is real, but it sits in a policy document, not the account record. A checker that pools all the evidence sees support and passes the answer.
Their second example is clinical. Imagine a medication detail that comes from a patient-history tool. The answer then presents it as a finding from the medical literature. The detail is real, but the reader is misled about where it came from and what it means.
The authors' point is that in sensitive settings, a wrong source can hurt as much as a wrong fact.
How the check works
ProvenanceGuard is a verification layer that runs after an agent has answered. It does not retrain the agent. It reads the recorded trace of the agent's tool calls, including each output and its source ID.
Then it works in five steps. It splits the answer into individual claims. It finds the most relevant source for each claim. It checks whether that source supports the claim. It asks whether that source matches the one the answer names or implies. Finally, it gives a verdict per claim and an allow-or-block decision for the whole answer.
The key design choice is that source identity is never pooled into one anonymous block of evidence. It is also strict about literal values. A number, date or identifier missing from the source cannot pass because the sentence sounds plausible.
In the paper's tests, local models did the work, so the recorded traces could be processed offline in a controlled setup. MiniLM found the relevant source. A DeBERTa language-inference model checked support. A local language model split each answer into claims. The researchers say these models are what they evaluated, not a requirement. A team using hosted models would need its own testing and calibration.
What the tests showed
The main test used answers from a medical agent that drew on patient records, research articles and other tools. Human experts checked 361 claims from 40 held-out answers. Experts said 139 claims should not pass. ProvenanceGuard blocked 138 and let one through.
The setting was cautious. The system also held 67 claims that experts considered supported and sent them for review or repair. Where a claim had an identifiable source, the system chose the correct one roughly 86% of the time. Four other support checkers were run on the same claims. None of them said which tool output supported each claim.
In a controlled test, the researchers changed the named source in 50 cases and left the evidence intact. ProvenanceGuard caught all 50.
Where it is weaker
A harder test used several similar sources. ProvenanceGuard scored 0.846 F1 on deciding which claims to block. F1 balances catching bad claims against blocking good ones. But it identified the exact source in only 50.3% of claims.
The authors say telling similar sources apart remains an area for improvement. This matters for companies whose systems hold many near-duplicate documents, such as several policy versions.
Two more limits apply. The main test covers one medical agent, and the results come from the authors' own paper. Buyers should treat them as a promising first result, not independent validation.
What happens to a blocked answer
A block is only useful if something follows it. The researchers wired in a repair step in the style of RARR, a method that revises answers against evidence. In the full-trace run, all 173 blocked answers were resolved. But 144 ended as fallback text, not a real rewrite. A separate repair run on reconstructed multi-source traces ended in fallback for only 2 of 59 blocked answers.
The authors describe this as the system declining to produce an answer it cannot verify. The overhead is roughly half a second per answer on their local setup.
There is one outside data point. NVIDIA's NVFlow project added an optional verification stage to its finance agent. The stage borrows ProvenanceGuard's source-aware approach. It tests completed answers against the SEC excerpts the agent had retrieved. One integration is not a trend.
What leaders should ask
The lesson here is that an agent's answer is a claim about evidence, not only about the world. Accuracy scores alone cannot tell you whether the claim points to the right record.
Put three questions to your team. Does our agent log every tool output with a source ID? Do we test attribution, not just accuracy? And what should the agent do when it cannot verify an answer: guess, escalate, or fall back?
A correct answer with the wrong citation is like a receipt from the wrong vendor. The amount may be right, but it still fails an audit.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Hugging Face.





