- Cisco Talos found malware in four families that embeds plain-language instructions meant to steer AI analysis tools, across 84 samples from January 2025 to July 2026.
- The tricks are cheap and uneven in effect, but they show attackers expect AI in the defender's triage process.
- Leaders should ask whether their AI triage tools treat text inside a file as evidence, not as instruction.
Malware is starting to write to the machine reading it
For years, malware authors hid their code from human analysts. Now some are writing to the software that assists those analysts. Cisco Talos, in research published October 8, describes malware that carries plain sentences addressed to AI models. Talos calls this class "A3: AI-Analysis Evasion."
The lesson here is simple. When attackers write notes to your AI tool, they are telling you where they think your defenses now sit. That is a useful signal, and it is also a design test for any team that has put a language model into its triage process.
How the trick works
Many analysis pipelines pull readable text out of a suspicious file. They then hand that text to a language model and ask for a verdict or an explanation. The model receives two things at once: the analyst's question and the file's contents.
Talos explains the weakness. If the file contains a sentence that looks like an instruction, a model that does not keep the two apart can treat the file's words as orders. This echoes the old problem of SQL injection, where a system mistook data for commands.
This layer is separate from packers or anti-debug checks, which target the binary itself. A3 targets the text-reading step that sits above the binary.
From a copied comment to template spraying
The first, FRUITSHELL, was a simple PowerShell reverse shell. It opened with a two-line comment aimed at AI tools. The comment does nothing when the script runs. It exists only to be read by an automated pipeline.
Talos then found the same comment word for word in nine additional scripts from at least four distinct actors. Some belong to a more capable collection that Talos tracks as ROZESHELL. Talos concludes the comment is not a sign of low skill. It reads as a free add-on that capable authors adopt too.
PLOTSAFE generates the sentence from a template, with different filler words in each build. The Go compiler removes unused strings, so the author wrote a 29-byte dummy function to keep them in. Talos says this shows the generator was built on purpose.
HOLLOWCLAD goes further. It repeats one instruction across seven chat-template formats used by different language models. If one copy matches the template a scanner uses, the model may read it as a genuine system instruction. HOLLOWCLAD also adds about 80 fake protector section names and notes meant to intimidate the model. Talos notes the binary is not packed at all.
MANTLEMAZE sprays templates too, but with a different intimidation strategy. It invents government contracts, certifications and NDA numbers to push the model toward refusing. Talos says the aim appears to be touching a topic that model guardrails prohibit.
Does it work?
Talos tested this. It took anti-analysis strings from seven malware families and placed each into every sample body. Five local language models triaged each pair, three times over. That gave 135 matched pairs per string.
The cheapest approach, a direct instruction to ignore the sample, worked almost universally. The more complex options had little effect and often made the models more suspicious. Talos calls the overall impact "a mixed bag."
The percentages need careful reading. They are net shifts in verdict across matched pairs: pairs that moved toward benign, minus pairs that moved toward malicious, divided by all pairs. On that measure, the best techniques reached about 35%. It is not a plain success rate.
These results come from local models in Talos's own test, not from every product on the market. Your tools may behave differently, which is a reason to test them.
A weakness for the attacker, and a test for you
The instructions have to be plain text. Talos argues this gives defenders a stable detection surface. Legitimate software has no reason to tell an analyzer to refuse analysis or to claim government contracts. Talos also says conventional detection is unaffected.
The deeper risk is in how AI tools are built. Talos's core defense is to treat text inside a sample as evidence and not as instruction. A pipeline that blurs that line hands the attacker a voice in the verdict.
Questions to put to your security team
First, does any AI tool in our triage process read text from untrusted files? If so, who tested it against embedded instructions?
Second, does the prompt clearly separate the analyst's question from the file's contents? Can extracted text ever be read as a system directive?
Third, do we flag instructions addressed to an AI inside a file as a suspicious signal in their own right?
Talos calls this trend "interesting, but not alarming." Its larger point is that AI-assisted security is now an active adversarial environment. Attackers are writing for the machine in your workflow, so that workflow should be built to doubt what it reads.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Cisco Talos.





