Skip to content
Security & Trust

In tests, AI-run decoy servers kept attack bots busy far longer than scripts

HoneyVAL researchers measured how long fake web systems hold AI attackers, and what that costs the defender

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
In tests, AI-run decoy servers kept attack bots busy far longer than scripts
In brief
  • In the researchers' benchmark, HoneyVAL's naive AI-powered decoy web servers kept AI attackers engaged for 82.6 requests on average, against 30.6 for the researchers' own scripted baseline mocks.
  • Asking a decoy to convince the attacker it was secure lengthened the exchange but raised the rate at which attackers spotted it.
  • Leaders should ask what a decoy costs per interaction, how often it is detected, and who acts on what it learns.

Security effort has traditionally focused on keeping attackers out. A research project called HoneyVAL looks at the opposite move: let the attacker in, but into a fake system, and keep them there. The question it raises is simple. When the attacker is software, what does it cost to waste its time?

What the researchers tested

HoneyVAL is a test bench built by Vero and colleagues for "honeypots", which are decoy systems designed to look real. The bench covers HTTP honeypots, meaning decoys that answer web requests. It uses 16 simulated web application backends, so every test runs against the same fixed targets.

An AI hacking agent plays the attacker. The researchers track four things: how long the exchange lasts, how often the agent detects the decoy, what the decoy costs to run, and how fast it responds. Response speed matters because a slow or oddly timed reply can give a decoy away.

Two control tasks check the results. One tests the attacking agent against real versions of the applications. The other tests the decoy against a functional test suite for the application it imitates. The code is open source on GitHub.

82.6
Average requests per interaction, naive LLM-powered honeypots
Source: HoneyVAL, Vero and colleagues (honeyval.xyz, accessed October 10, 2026)
30.6
Average requests per interaction, custom rule-based baseline mocks
Source: HoneyVAL, Vero and colleagues (honeyval.xyz, accessed October 10, 2026)

What they found

The team tested what it calls naive decoys, driven by a language model. In the benchmark, these held attackers' attention well beyond the scripted mocks, which the researchers wrote as a baseline. The AI attackers were also less able to tell the decoys from real systems.

The project summary does not give detection rates or running costs in the text we reviewed. Both are tracked by the benchmark, and both decide whether the approach pays off.

Why a language model changes the decoy

A rule-based mock answers only what its authors scripted. An attacker who probes outside the script meets an error or silence. A language-model decoy writes each reply as the request arrives, so it can respond to unexpected probing. The summary does not say why the gap appeared, but that design difference is the obvious one to examine.

Longer engagement matters because the researchers treat interaction length as a stand-in for what the defender learns about the attacker.

The trade-off built into deception

Tuning a decoy has a price. When the researchers prompted a honeypot to "convince" the attacker that the simulated application was secure, interactions grew longer. Detection rates also rose.

This is one tuning experiment, not a rule. It shows that a longer exchange and a stealthier decoy can pull in different directions.

The same idea on the phone

Wired's Kernel Panic newsletter describes a related effort against phone scams, which is a different setting from web honeypots. The Australian company Apate runs AI bots that talk to scammers and keep the conversation going. Its founder, Dali Kaafar, says the aim is to build "the perfect victims for scammers".

Kaafar gave Wired some figures. He put the number of bots at about 350,000 and said calls often last beyond two hours. He also said the system has gathered over 250,000 items of intelligence on fraudsters, such as scam URLs and money mule accounts. These are the company's own figures. Wired's reporters tried a demo and did not get the AI to invest in their cryptocurrency pitch after six minutes.

350,000
Bots in Apate's platform, per its founder
Source: Dali Kaafar, Apate, as reported by Wired Kernel Panic (October 10, 2026)

The lesson: deception could become a contest of machines

The lesson here is that deception could become a software contest. Each fake reply costs the defender compute. Each wasted exchange costs the attacker time and, if the attacker is also software, compute too. That makes the economics the real question, and HoneyVAL tracks running cost for that reason.

Kaafar argues that the benefit goes to people who would otherwise be targeted. In his words, a minute a scammer spends on a bot is a minute not spent on a real person. That is his argument about phone scams. HoneyVAL is a simulated benchmark and reports no real-world victim benefit. Whether a web decoy that absorbs an automated attack helps anyone in the same way is untested there.

Questions to put to your security team

Do we run any decoys on internet-facing services? If so, who reads what they collect, and how quickly does it reach people who can block the attacker?

If we trial an AI decoy, what does each interaction cost, and how often do attackers spot it? Does anyone own the trade-off between keeping attackers engaged and being detected?

Who approves deliberate deception, and has legal and risk review looked at it? A decoy is only worth its running cost if someone acts on what it learns.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Vero and colleagues (HoneyVAL research, hosted at honeyval.xyz).

Share this insight