- Anthropic says an actor used Claude to rebuild flagged implants. ReversingLabs saw five stager builds with different hashes but one behavior profile.
- RL argues hash blocklists go stale faster when rebuilding is cheap. In this case, a hunt on behavior found files that hashes could not.
- Leaders should ask how far back their telemetry reaches and whether vendor reports are re-checked against it on the day they land.
Changing a file is cheap. Changing what the file does is expensive. That gap is the useful idea in a new analysis from security vendor ReversingLabs (RL), and it matters to anyone who pays for detection.
Anthropic's September 2026 Threat Intelligence Report describes an actor it calls GTG-20006. Anthropic calls the attribution "consistent with public reporting linking the actor to Midnight Blizzard." Anthropic says that when security products flagged the actor's implants, the actor "used Claude to systematically identify, modify and redeploy the detected artifacts."
RL then ran the two malware hashes from that report against its own telemetry. RL sells threat intelligence, so read its conclusions with that in mind. Its findings are specific and checkable.
What the malware actually does
RL describes two pieces of malware. The first is a PowerShell stager, a small script that prepares the ground for later tools. Per RL, it runs out of sight, unpacks a concealed payload and executes it. It also pulls browsing history using code lifted from the public PowerSploit and Empire projects. Its results go to a staging server.
The second is a backdoor written in Go. RL says it registers as a Windows service under a misleading name, so it survives reboots. It also slips into a different running process and harvests data from browsers and Teams.
None of this is new. MITRE ATT&CK, a public catalog of attacker techniques, has listed such behaviors for years, RL points out. Its conclusion is that the rebuilds brought no new tricks.
Why hashes go stale and behavior does not
A hash is a fingerprint of one exact file. Change one byte and the fingerprint changes. If a model can rebuild a flagged file each time a scanner catches it, every published hash ages faster. That loop is what Anthropic describes.
RL's data fits that picture, though RL says it cannot show whether AI produced it. Over a 25-day window, RL collected five separate versions of the stager. No two shared a fingerprint. Fuzzy hashing, which scores how alike two files are, grouped them into two code variants. A team that relies on hashes would have to start over with each version.
RL then wrote a hunt around the operator's beacon labels, the tags the stager sends home. The hunt returned 7 malicious files and no goodware.
Signals about how a file was built did not help. The Go backdoor's import hash, a fingerprint of the code libraries it uses, matches 12,000 goodware files and 6,900 malicious ones in RL telemetry over the past year. Go programs compiled the same way share traits whatever they do, so this fingerprint cannot separate good from bad. A similarity search on the backdoor returns about 19,500 files across dozens of families.
Reputation lagged the evidence
RL's second finding concerns slow-moving reputation scores. Anthropic names 14 servers. Thirteen looked low risk to outside reputation services. RL had tied only one to malware, and it carried a single outside flag before the report appeared.
The two lure shortcuts are files with no code at all. Antivirus engines do not flag them, and RL's own cloud verdict still labels them goodware. Only their destination makes them harmful.
Anthropic's appendix lists domains, and RL now rates 61 of them malicious. RL's telemetry had already flagged 48 before the report came out. Outside sources remain slow: a typical domain draws flags from just 2 of roughly 47.
Memory is the control
Think of a getaway car. A thief can repaint it and swap the plates in an afternoon. The driver's habits and the route to the safe house stay the same. Investigators who keep old records can still match the habits.
That is the lesson here. If rebuilding is cheap, the defender's edge shifts to the records behind the questions. RL frames four: has this been seen before, what does it do, what infrastructure does it touch, and what else links to it.
A caution on RL's headline claim. Its records show the Go backdoor first seen on July 6 and the stager on July 10. RL first saw the files about two months before the report and rated them malicious on August 1, weeks before Anthropic published. Seeing a file is not the same as understanding the campaign. The wider map came only after the report supplied the hashes. Pivots from those two hashes then turned up 7 further files, 8 URLs and 2 links between domains and servers that Anthropic's appendix lacks.
This is also one case. It does not show a general rule about AI and detection.
Questions for your security team
Do we re-check every vendor report against stored telemetry the day it lands? How far back does that data reach? Both seed files here sat in RL's corpus for two months, so a useful answer depended on history.
How long do hash-based blocks stay useful in our environment? Can we measure it? Do we write detections around capabilities, not only file fingerprints?
When a server looks low risk, do we check which files and domains are linked to it before we accept that? Here, reputation scores trailed the file evidence.
Hashes tell you which paint a file wears. Telemetry history tells you whose car it is.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: ReversingLabs.





