Negation questions answered correctly: 12 of 12 (Source: paultendo.github.io (September 29, 2026), every model on both altered contracts)
The test that failed at deception and registered on the invoice
Security teams tend to ask whether an AI model can be tricked into a wrong answer. A test published on September 29 by researcher Paul Tendo points to a second question: what does it cost the model just to read the input? The lesson here is that the risk moved from the answer to the meter. The model's judgment held. The billing did not.
Tendo wrote a new consultancy agreement, eight clauses and 89 lines, and produced two altered copies. In the first, 18 words that reverse a clause if misread, such as "not", "without" and "waives", were spelt with Unicode confusables, characters that resemble ordinary letters but are different symbols. In the second, 60% of the lowercase letters were replaced with confusables. The exercise covered seven recent GPT and Claude models. Each reviewed both copies and answered a dozen questions whose correct answers hinge on a negation, then was asked whether anything looked odd. That came to 91 calls.
Nobody was fooled, and everybody was charged
Accuracy was uniform: no model missed a negation question on either version. Each also interpreted clause 5.1, "shall not be limited", as uncapped liability. Even Haiku 4.5, the model that faltered in the February round, answered without error.
The cost side looks different. Tendo's February run used GPT-5.2 and Claude Sonnet 4.6, and reading the heavily substituted contract consumed 5.2 times the tokens. He labels this Denial of Spend. Reading is not the whole bill, though. A question also produces an answer, the flood does not lengthen that answer, and output tokens are priced above input tokens, so the total rises by less than the token count. Pricing the full exchange at list rates, where Claude charges five times as much for output as for input and GPT eight times, gives a rise of up to 3.9x by his calculation. On tokens alone, the newer Claude models land at 4.0x because they already spend more tokens on plain English.
Noticing is not the same as saving
Some models remarked on the strange characters. Unprompted, three Claude models, Fable 5.1, Opus 5.5 and Sonnet 5, called them out. GPT-6 Astra and Haiku 4.5 never did. Spotting the anomaly saves nothing, because the tokens were consumed before any comment was made. Sonnet 5 went a step further and declined, in all three tries, to say whether the flooded contract seemed unusual. Even those refusals drew roughly 5,800 input tokens each.
For a budget-holder, the picture is easy to describe. To a reviewer reading the document, the text looks normal. The change shows up only as a line item on the invoice.
What the evidence does and does not show
The limits matter. Tendo tested one contract with one or two runs per model, through the Codex CLI and Claude Code rather than the APIs. The bill figures are list-price ratios without caching, while the token counts were measured. He also says he cannot find a single reported case of this technique being used in practice.
He points to two nearby facts as reasons to prepare. Microsoft reported this September on a phishing run that embedded invisible characters within words to slip past spam filters, at volumes up to 2.37 million emails a day. Its recommendation was to clean up text before pattern matching and before an AI sees it. Separately, the 2026 OWASP Top 10 for LLM applications now ranks Unbounded Consumption sixth, up from tenth. The mitigations it names centre on token and size caps, and Tendo's point is that a cap counted in characters would miss text that is five times as costly per character.
Questions to put to your team
First, do we normalise text before it reaches a model? Tendo's canonicalise() function in namespace-guard, rebuilt for version 0.23, rewrites only words that show signs of tampering, so genuine Russian, Turkish or Sámi words are left alone. On his flooded contract it brought token counts to within 3% of the clean version in under 2 ms.
Second, are our input limits measured in tokens rather than characters? Third, is there a spending cap for each customer, so that a single input stream cannot move the bill? Fourth, who sees the alert when cost per document departs from its baseline?
This is one data point, not a trend. But it shifts the question from whether the model can be deceived to whether the invoice can be. A defence built only around wrong answers would have scored perfectly here and still missed the cost.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: paultendo.github.io.





