- Cloudflare says its first single-agent security prototype made unsupported claims, so ordinary code now gathers evidence and sets scope before any model runs.
- The lesson is that an AI agent earns trust when its authority shrinks, not when its prompt grows.
- Buyers of AI security tools should ask who gathers evidence, how claims are checked, and how a failed lookup is recorded.
A prompt is not a fence
Security alerts seldom come one at a time. One event can trigger a spike, and an analyst must work out which alerts belong together. Cloudflare calls this the alert paradox. On October 7 it explained how AI agents now handle part of that work in its Managed Defense service.
The useful part of the post is the failed first attempt. Cloudflare gave one general-purpose agent the full investigation. The analysis helped, but it also contained claims the evidence did not back up.
Cloudflare says the cause was partly design. Telemetry, detector descriptions, policies and threat intelligence all sat in one prompt. Their separate roles blurred together.
The lesson here is that trust comes from narrowing what an agent may decide. A longer, more careful prompt does not achieve that. Cloudflare says a prompt cannot be relied on to keep an agent inside its limits.
Three ways the first version went wrong
Cloudflare names three recurring problems. The first it calls context becoming authority. An alert only suggests that something may be wrong. A broad agent could treat that suggestion as settled fact.
The second is scope drift. The agent could pull data from the wrong account, the wrong time period or the wrong source.
The third is that failure vanished from view. A lookup that timed out could look the same as a lookup that found nothing. For a security team, that gap decides whether an all-clear can be believed.
How the harness works
Cloudflare moved evidence collection and scope control into ordinary application code. This code runs before any model analysis. The first half of the system uses no AI agents at all.
Instead, fixed lookups with versioned API calls collect the basics. These are who the customer is, past detections, normal traffic, what the defenses did, and what the network saw. Every record carries a tag for its origin, version and time.
This has a practical payoff. If agents fetched their own data, two runs could disagree because the inputs changed. With a fixed snapshot, the same case can be replayed. Any difference then comes from interpretation, not retrieval.
Next comes triage. Clef, Cloudflare's open-source decision model running on Workers AI, compares each alert with its history. Alerts likely to be false positives skip the specialist agents. Known high-volume noise is marked passive and stays out of the active queue.
Alerts that need review go to a coordinator, which runs four specialist agents in parallel. They cover traffic analysis, customer context, global telemetry and threat intelligence. A synthesis agent then merges their findings into one advisory. It cannot fetch new evidence. It also cannot pick a classification outside an approved list.
Evidence must be admitted, as in a court
Before analysis, the system assembles a versioned evidence package. It records what the case is about, the period under review, the approved evidence, the rules in force, where the data came from and what could not be checked.
Specialists must point to items in that package. Ordinary software then tests each pointer. Is the item real? Does it belong to this case? Does it support the claim? Findings that fail are corrected or logged as limits.
A courtroom works the same way. A witness may rely only on evidence the court has admitted. The judge does not accept the witness's word that it exists.
Gaps are recorded, not hidden. Suppose Cloudflare's network-wide view is unavailable. The system can still say what looks odd for one customer. It cannot say whether others see the same thing. When evidence falls short, it offers no label and no recommended outcome.
Cloudflare adds that the model has no authority to cross customer boundaries or act in place of the analyst.
What the post does not show
This is Cloudflare describing its own system. The post gives no figures for time saved, accuracy or false-positive rates.
Access is limited. Cloudflare calls it an early beta, offered for qualifying application-security alerts and cases within Managed Defense.
Cloudflare also says the human analyst stays responsible for the decision and any fix. Each alert and case carries the evidence behind the recommendation. The analyst can accept the advice or change it.
Questions to put to any vendor selling security agents
Who collects the evidence: ordinary code or the model? Can an investigation be replayed from the same inputs? Does every claim cite a specific item, and does software check the citation? When a lookup fails, does the record say "not checked"? What can the agent not decide, and who signs off on actions?
An AI agent earns trust through what it is not allowed to decide. The strongest claim a vendor can make is a short list of those limits, plus the code that enforces them.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Cloudflare.





