- Cisco Talos reports malware authors are embedding plain-language instructions in code to steer AI-assisted analysis, and it found the trick works about 35 percent of the time.
- The risk sits in any pipeline where an AI reads text an attacker can write.
- Talos advises treating extracted text as evidence, not as a directive, and flagging instructions aimed at analysis systems.
When you cannot fix the part, control what it can touch
Pierre Cadieux of Cisco Talos opens this week's Threat Source newsletter with a story from an earlier job. He ran security at a financial institution. One department printed the checks the bank used to pay other institutions and its customers. Its systems were out of patch compliance, and he went to find out why.
The check-printing software would not run on current operating systems. The printer cards needed older hardware ports. The environment had zero tolerance for downtime. As one employee told him, they did not want to be the reason "someone's grandma doesn't get her check."
His answer was not a patch. He put the devices on their own network and blocked internet access to and from them. That also made them less likely to show up when an attacker mapped the internal network. The devices stayed vulnerable, but the odds of something going wrong fell.
The lesson here is that some components cannot be made trustworthy. For those, the discipline is to limit what they can reach and what can reach them. The same newsletter carries a research finding that applies that idea to AI.
Malware that talks to the analyst
Talos is disclosing new findings from its CAIRN research. Malware authors are embedding natural-language instructions in their code to evade AI-assisted analysis. Talos calls the trend "A3: AI-Analysis Evasion."
Over 18 months, Talos tracked techniques that range from simple comments telling an AI to ignore a file to "template spraying." Talos describes template spraying as designed to trick specific large language models. The newsletter summary does not explain how it works.
Talos says the methods are cheap and appear across all levels of malware sophistication. Some A3 families pair the trick with serious capability. MANTLEMAZE, for example, abuses vulnerable drivers to disable endpoint detection and response (EDR) software from kernel space. EDR is the monitoring software on laptops and servers. Kernel space is the most privileged part of an operating system.
How the trick works
An AI-assisted analysis tool reads text pulled from a suspicious file, such as comments and strings. A language model has no firm line between text to judge and text to obey. An attacker writes a sentence addressed to the AI, and the model may treat it as an order. This is prompt injection.
Thirty-five percent is not a majority. It is also far from zero, and attackers pay almost nothing to try. Talos's summary does not say which models or how many samples sit behind the figure. The full blog lists sample hashes.
The same logic as the check printers
Talos's advice mirrors the bank's. Anyone building or using AI-assisted pipelines should treat text extracted from a sample as evidence, never as a system directive. That is segmentation for instructions. The model may read the attacker's words, but they should have no authority over it.
Talos also sees an opening for defenders. The evasion instructions must be written in plaintext, which gives a stable detection surface. Security teams can flag imperative language aimed at analysis systems inside binaries as a suspicious signal. The attacker's trick becomes a tell.
Questions to put to your team
Where in our tooling does an AI read content an attacker can write? Is that content kept separate from the instructions the model receives? Can an AI verdict alone close an alert, or does a second check confirm it?
Do we scan for instructions addressed to analysis systems? Which systems sit out of patch for good business reasons, and are they isolated and documented?
The bank could not patch its printers, so it controlled their reach. Your AI analyst can be talked to, so control what its inputs are allowed to say.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Cisco Talos.





