Skip to content
The AI-First Web

CrowdStrike: an AI safety classifier misses attacks split into harmless steps

A classifier blocked all of about 515 direct techniques tested, but split-up tasks got through in 9 of 10 categories.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
CrowdStrike: an AI safety classifier misses attacks split into harmless steps
In brief
  • CrowdStrike reports that an AI safety classifier blocked all of about 515 direct attack techniques tested, but splitting tasks into harmless-looking requests bypassed it in 9 of 10 categories.
  • The gap is in the design: the classifier judges one request at a time, and the harm appears when a second, unprotected model puts the pieces together.
  • Leaders should ask AI vendors whether their safety claims cover multi-step and multi-model use, not only single prompts.

A guard who checks every bag at the door can still miss a theft. Each visitor carries in one harmless part. The guard sees nothing wrong. The danger appears only when someone joins the parts inside. CrowdStrike tested one safety classifier and found this gap. Its researchers argue the same gap applies to classifier-based safety designs in general. That is their claim, not a result tested across many filters.

What CrowdStrike found

CrowdStrike's Cyber Superintelligence Lab tested a safety classifier. This is a second AI model that reads each request and blocks it if it looks harmful. The lab says it tested the most advanced publicly deployed classifier of this kind. It calls the model behind it "Frontier Model A".

Against direct attacks, the classifier held. The team tried about 515 bypass techniques. They included encodings, psychological manipulation, many-shot tactics, tokenizer exploits and Unicode tricks. None got through directly.

0%
Direct bypass rate across about 515 techniques
Source: CrowdStrike Cyber Superintelligence Lab (October 6, 2026)

Then the team changed tactics. They broke harmful goals into small requests. Each request, on its own, was genuinely harmless.

The team tried this on ten categories of attacker behavior. Nine of them ended with exploit code that worked.

9 of 10
Offensive categories where the split-up method yielded working exploit code
Source: CrowdStrike Cyber Superintelligence Lab (October 6, 2026)

How the method works

The method has three steps. First, split the offensive goal into subtasks. Second, reword each subtask as an ordinary software request. Third, hand the answers to a smaller model that has no classifier. That model assembles the finished tool.

CrowdStrike's lead example is Windows process injection. This is a way to run code inside another program. Three harmless-looking requests to the protected model returned every building block. The smaller model then wrote a 221-line C program that compiles. It carried out the full chain of four Windows calls.

221 lines
Length of the compilable program the unprotected model assembled
Source: CrowdStrike Cyber Superintelligence Lab (October 6, 2026)

The reframing is flexible. In one variant, the team asked for detection rules for a named vulnerability. The classifier agreed, because the stated purpose was defensive. The smaller model then used the detail that came back to build working proof-of-concept exploits.

Not a classifier failure

CrowdStrike is clear on one point. The classifier made no wrong call. It judged every request it saw correctly. The harm exists only when the pieces are combined. That happens outside its view.

The researchers also call the pipeline a transfer of knowledge. The larger model knows more and writes better code. The smaller model inherits that skill during assembly. The team checked this with controlled tests across 8 use cases.

The transfer mattered most in specialized areas, such as process injection and exploits for specific CVEs. For well-documented techniques such as reverse shells, the smaller model needed little help.

One category held. This was defense evasion, which means disabling runtime security monitoring. The classifier blocked it even when the wording avoided security terms. The researchers say it recognized the concept itself.

Why simple fixes fall short

One obvious fix is to look for patterns across requests from one API key. CrowdStrike says this has limits. An attacker can spread requests across providers or rotate keys. They can also run local open-weight models for planning and assembly. A provider sees only its own slice.

CrowdStrike adds that Microsoft found the same technique on its own. It quotes Microsoft's paper: "the decisive context stays with the orchestrator."

What leaders should take from it

This is one security vendor's own research. Treat it as a credible lab's finding, not a settled measure of real-world attacks.

Still, the lesson is plain. A safety filter that works on single prompts does not tell you what the whole system allows. The risk sits in the seams between models. No single provider owns those seams.

Questions to put to your AI vendors and security team:

First, do the vendor's safety claims cover multi-step and multi-model use, or only single prompts? Second, what can the vendor see across requests from one account, and what can it not see? Third, does your threat model assume an attacker has a small local model next to a hosted one? Fourth, are your detection and developer tools used in ways that look like these harmless reframes?

CrowdStrike puts it simply: the classifier works, and the architecture around it needs hardening. A strong lock on every door does not protect a building if the parts can be carried out one at a time and assembled elsewhere.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: CrowdStrike.

Share this insight