Skip to content
The AI-First Web

Attackers asked OpenAI's model to unlock its own protected reasoning

OpenAI says its encryption held, but a bug let data cross between conversations. It says other models share the flaw.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Attackers asked OpenAI's model to unlock its own protected reasoning

AI-generated image for WebPulse. About our images

Key finding

Prompts in the July 24-25 peak: 16,000 from 4,000 users (Source: OpenAI blog post, as reported by CyberScoop (September 30, 2026))

Encryption cannot help when the key works in the wrong place

Encryption keeps data from people who lack the key. It offers little if a system that holds the key will use it in the wrong context. That is the idea at the centre of a case OpenAI disclosed this week.

OpenAI reports that it ended an organised effort to copy how its models reason. The practice is called distillation. The company stressed that the operators "did not break our encryption." Instead, it said, they "manipulated model interactions" so that protected reasoning appeared in a form the requester could see.

What OpenAI says happened

OpenAI says it has traced the activity back to July 1, when it was low-level. It grew through July. On July 24 and 25, OpenAI counted 16,000 prompts from 4,000 users that matched what it called a "relevant extraction pattern." By July 28 the number of suspicious users had reached 15,000. That day, OpenAI said, it fully disrupted the operation.

16,000 from 4,000 users
Prompts in the July 24-25 peak
Source: OpenAI blog post, as reported by CyberScoop (September 30, 2026)
15,000
Suspicious users by July 28
Source: OpenAI blog post, as reported by CyberScoop (September 30, 2026)

How the method worked

OpenAI described the technique as "novel." The attackers took encrypted reasoning data from one conversation. They then opened a separate conversation and asked the model to decrypt that content and write it out as plain text.

OpenAI's fix points to the weak spot. It closed a bug that let encrypted data from one conversation be decrypted in a different one. OpenAI has not described the mechanism further.

The aim, per CyberScoop's account, is to copy a model's capabilities and training data. CyberScoop reports that security experts at Google and other firms say this often involves thousands of accounts bought on black or gray markets. Those accounts are then flooded with millions of prompts.

What is not established

OpenAI links a "core cluster" of the activity to individuals working for Moonshot AI, a China-based rival. CyberScoop notes the blog post gives no technical evidence for that attribution. OpenAI also said it is unclear whether all the activity is related.

OpenAI told CyberScoop it would share nothing more "for security reasons." CyberScoop has asked Moonshot AI for comment. The source does not report a reply.

Two further points matter. OpenAI's position is that other AI models share the same weakness. It has passed details to industry groups, including the Frontier Model Forum. Separately, outside researchers had flagged a comparable weakness to OpenAI in August.

The lesson: a decryption ability was not tied to the conversation that created the data

This is our reading of the case, not a finding in the source. The sourced fact is the bug fix: encrypted data from one conversation could be decrypted in another. That suggests a missing scoping check. The ability to decrypt was not bound to the conversation that produced the data.

There is another possible reading. The model's own judgment about what to reveal may have been the failure. The source does not establish that. It says only that the attackers manipulated model interactions.

A second observation is also our interpretation. The signal OpenAI describes is volume and pattern: many prompts matching one extraction pattern. It is not a break-in alarm. A security team watching only for intrusions would have had little to see.

Questions to put to your team

If your company builds on or exposes AI models, start with these.

First, which of our protections rely on the model choosing not to reveal something, and which are enforced by code outside the model? Second, can data from one user session be fed into another, and what happens when it is? Third, do we monitor for many accounts sending similar prompts, not just for unauthorised access? Fourth, what does our AI vendor tell us about incidents like this one, and how fast?

OpenAI says nobody broke its encryption. The harder lesson is that nobody had to.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: CyberScoop.

Share this insight