- MIT Technology Review reports that AI refusal is a trained, probabilistic habit layered over knowledge the model keeps. How it works inside is still a hypothesis.
- Safeguards cost money and can be bypassed. One Anthropic filter type raised computing costs by 24%, and Italian researchers got past two dozen models using poems.
- Treat a vendor's refusals as one control with an error rate. Add your own limits on data and actions, and ask who sets the line.
When a chatbot says no, it looks like judgment. It is not. MIT Technology Review describes refusal as a trained habit. The model still holds the knowledge underneath. For anyone buying or deploying AI, the lesson is plain. Refusal is a control with an error rate. Manage it like one.
How a model learns to say no
It starts with people. In 2022, OpenAI hired dozens of red-teamers. These are outside testers who try to break a model. Paul Röttger, then finishing a PhD on online extremism, was one of them.
He told the magazine that testers could ask anything they thought deserved a refusal. They logged the results in an Excel sheet. When he asked for an Al Qaeda recruitment post, the model wrote it. A few months later, he asked for an Al Qaeda pamphlet and the model refused.
Companies now add more steps. They reward a model for refusing harmful prompts. They punish it for refusing harmless ones. Often, other models run this training.
Then they wrap the model in smaller models called classifiers. Some read the user's prompt and block dangerous ones. Others read the answer and block harmful ones.
Nobody can fully point to where the no lives
Inside the model, a refusal appears as a pattern of internal signals, called activations. They light up when a prompt resembles ones the model was trained to decline. Researcher Andy Arditi has shown that if those signals are removed, the model stops refusing.
Jannes Elstner co-wrote a Google-funded study on this. He told the magazine that you may think you have found every part behind a refusal. Other hidden parts may still play a role.
We can see that a model said no. We know training caused it. How it decides is, in the article's words, "at best, a hypothesis." Elstner's answer to the unease: "We need refusal whether we understand it or not."
Layers of safeguards, each with holes
The industry's answer is the Swiss cheese model. Each layer has gaps. Stack enough layers and, the hope goes, the gaps stop lining up.
The layers cost money. Anthropic has put a number on it. One filter type raised the computing bill for its chatbots by 24%. Anthropic and others are now moving to cheaper probes that watch the model's internal activations.
Classifiers can be updated within weeks when there is something new to refuse. They are still statistical tools. Harvard researcher Ryan McBain asked major models the same risky question about suicide, again and again. They generally refused. But "every so often, they won't."
Why the knowledge stays, and why attackers go after the refusal
Steven Adler worked on safety at OpenAI until 2024. He says dangerous knowledge is tangled up with useful knowledge. Strip it from the training data, and the model can still rebuild it from related facts. Cutting the skill itself would leave the model much less smart.
So attackers target the refusal itself. Tricking a model into revealing what it knows is called jailbreaking. The article gives two examples.
Earlier this year, Italian researchers got past the guardrails of 24 popular models. They wrote their questions as poems. Last year, another team showed a "refuse, then comply" attack. The model says "Sorry, I can't do that." Then it gives the forbidden answer anyway.
Who draws the line is a business question
Where refusal should stop is a judgment call. Zico Kolter, an OpenAI board member and Gray Swan cofounder, said "where you draw the line is a huge question." Today AI companies draw it. The article says they do so with utmost secrecy.
Governments will soon draw their own lines. The Pentagon has pushed frontier model companies for fewer refusals. The author warns that governments could use refusal to block legitimate speech.
For a company, this cuts both ways. A tool may block something your staff need. The article notes that some users ask about system weaknesses so they can patch them. A tool may also let through something your policies forbid. Either way, the vendor sets the rule, not you.
What leaders should ask
This is WebPulse's view of what follows. Do not treat a vendor's refusal as your safety plan. Put these questions to your teams and suppliers.
First, what happens when the model fails to refuse? Limit what any AI system can reach or do. Then a bad answer has a bounded cost.
Second, has anyone tested our own use cases with our own prompts, repeatedly? McBain's finding shows that a single pass proves little.
Third, who sets the refusal line? How are changes announced? What is our route when a legitimate task is blocked?
Fourth, are outputs logged and reviewed? Then we find the failures before anyone else does.
The article calls refusal the load-bearing wall of AI safety. A wall can be inspected. This one is a statistical habit that its builders are still learning to read. Build the rest of the structure so it can stand if the wall gives way.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: MIT Technology Review.





