- Anthropic's Claude Haiku 5.5 costs about 75% less to run on average than Haiku 4.5, and Anthropic says it resists prompt injection better than earlier Haiku models.
- Help Net Security reports it trails Sonnet 5.5 and Opus 5.5 on one Gray Swan benchmark, with most remaining weakness in graphical computer use.
- Before putting a small model in front of email, web pages or customers, limit what it can do if a hidden command gets through.
Cheaper models get put in front of more outside text
When an AI model gets much cheaper, companies often stop asking whether it is smart enough. They start asking where else they can use it. Each new place can be another source of text the company does not control.
That is the lesson in Anthropic's launch of Claude Haiku 5.5. Anthropic says the small model costs around 75% less to run than Haiku 4.5 on average. It pitches the model for summaries, classification, live customer support and browser use. Many of those jobs, such as customer support and browser use, mean reading content from outside the company.
Anthropic says the cut is 50% for longer requests. It adds that 90% of requests to Haiku 4.5 were in the shorter group. One customer quoted by Anthropic says a single feature makes about 8 million model calls a week. At that volume, price helps decide which workloads get automated.
How a hidden command works
Prompt injection hides instructions inside material an AI reads, such as an email or a webpage. The aim is to make the AI act against the wishes of the person using it.
The attacker does not need to break into anything. They only need to put text where an AI will read it.
Think of a mailroom clerk who reads every letter and sometimes obeys instructions written inside one. The more mail you hand the clerk, the more chances a forged note has to work.
What the tests show, and what they leave open
Help Net Security reports that Anthropic says no earlier Haiku model has held up as well against prompt injection. Against attackers who adapt their tactics, in coding and computer-use settings, Haiku 5.5 performed at about the level of Anthropic's top models.
A separate benchmark from Gray Swan tells a less comfortable story. On it, Haiku 5.5 fell behind Sonnet 5.5 and Opus 5.5. Help Net Security says most of its remaining weakness was in graphical computer use.
That matters because Anthropic also calls the model well suited to computer use and browser use. Graphical computer use and browser use are not the same task. But they are close neighbours, and the weak spot sits next to the pitch.
Refusing a malicious request from a user is a different test from resisting a command hidden in a webpage. Both results are progress. Neither covers the other.
One caveat applies to a different set of tests. In the evaluations of harmful requests and sensitive conversations, Anthropic left out extra live defences, such as real-time probes and monitoring. Those tests still found weaknesses in conversations about self-harm and eating disorders. Anthropic advises developers using its API to add their own safeguards.
Help Net Security does not say whether the prompt-injection tests included those live defences. That leaves one open question about how the injection results would look in production.
Offensive skill is up from Haiku 4.5, but well below the larger models
Help Net Security reports that Haiku 5.5 is better than Haiku 4.5 at finding vulnerabilities and writing exploits. Anthropic measured this with the model's cybersecurity safeguards switched off.
In one test on known flaws in Chrome's V8 engine, the model got a program to run code of its choosing in four of 410 runs. On the ExploitGym benchmark it beat Sonnet 5. It still stayed well below the other 5.5-family models.
Help Net Security cautions that these figures may not match what users can do under standard safeguards. Anthropic says those safeguards block penetration testing and other attacker-style techniques. They allow more defensive work than Sonnet 5.5's safeguards do. Qualifying security professionals can apply to its Cyber Verification Program for fewer restrictions.
What to ask your team
First, list every workflow where a small model reads text from outside the company. Customer messages, inbound email, web pages and uploaded documents all count.
Second, for each one, ask what the model can do if it obeys a hidden command. Can it send data, change records or click through a site? A model that only drafts a summary is a smaller risk than one that acts.
Third, ask your vendor whether its injection tests ran with the protections you will actually use. It is unclear whether Anthropic's prompt-injection tests included its production protections. Plan to add your own safeguards either way.
Fourth, keep human approval on actions that move money, data or access. A lower price is a reason to automate more work. It is not a reason to remove the check.
A model that resists hidden commands better is good news. The risk now depends less on the model and more on how much you let it touch.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Anthropic.





