- Mistral AI says its new open model, Mistral Large 4, completes a security task that some closed models decline. The scores are mostly Mistral's own.
- A security team that rents AI also rents the vendor's rules, and a refusal during an incident acts like an outage.
- Test refusal rates on your own tasks, decide who would own usage policy if you self-host, and wait for independent results.
Rented AI comes with someone else's rules
A security team that rents its AI also rents someone else's rules. Those rules can change. They can also say no in the middle of an incident.
That is the argument behind Mistral AI's launch of Mistral Large 4, which it calls ML4. On October 6, the company opened a public preview. It says it will release the model's weights by the end of October. Weights are the trained numbers that let a company run the model on its own systems.
What Mistral announced
ML4 has 1 trillion parameters, but only 49 billion are active for any one request. This is a mixture-of-experts design. Each task is routed to a part of the model rather than the whole thing, which keeps running costs lower than the total size suggests.
The company says the training run used 3,800 Grace Blackwell chips from NVIDIA. They sit in datacenters Mistral owns in Europe. Customers who try the preview are served from those same sites.
The refusal problem
Mistral's central claim is about refusals. Defending software often starts by proving that a flaw is real. Mistral says the safety filters on closed models can block exactly that work. It argues that a tool pulled away during an active incident is itself a danger.
Mistral points to one part of the Artificial Analysis Cyber Index. A model is given a known flaw in open-source code. It must first show the flaw works, then write the fix.
On that part, Mistral reports 82% for ML4 and calls it the top result among all models. It also says Claude Opus 5.5 and GPT-6 Astra, among other closed models, land close to zero. Mistral's explanation is that they decline the job.
On Cybench, a 40-exercise set built from security competitions, Mistral says ML4 completes 93% of the challenges.
How to read these numbers
Treat them as a vendor's claims about its own product. The Cyber Index is an independent evaluation, but the comparison with closed models appears in Mistral's post. We have not seen a separate write-up from the index's authors.
The model is also a preview. Mistral says architecture details, more benchmarks and its post-training method will follow with the weights.
One distinction matters. A near-zero score from refusal measures a vendor's policy, not a model's skill. Mistral's own account is that the closed models decline the task. Whether that policy is a flaw or a safeguard depends on who is asking.
Control comes with responsibility
Before release, Mistral is testing ML4 in real settings with security leaders, selected partners and government authorities. Those groups use a version with fewer content limits and wider cyber abilities. So the public preview is not the only version in use.
Mistral also reports that ML4 refuses malicious cyber requests more often than any other open model it compared. It cites JailbreakBench, StrongREJECT and AgentHarm. It adds that the model resists 93.3% of attacks on Lakera's B3 benchmark. These too are Mistral's figures.
The post says threat actors increasingly jailbreak closed models for offensive work. It gives no data for that claim.
This shows the trade-off. Mistral pitches open weights as letting organisations run security work "under their own policies." Once the weights are out, the filters belong to whoever deploys the model. Capability and control are one purchase. So is the responsibility for how the tool is used.
Questions to put to your team
Ask who can switch off the AI tools your incident responders use, and what happens when they do. A refusal at 2 a.m. is an outage.
Ask vendors for refusal rates on your real tasks, not on a public benchmark. Test them yourself on non-sensitive examples.
If self-hosting is on the table, ask who will own the usage policy, the logging and the audit trail. Ask what hardware a model of this size needs. Mistral's post does not say.
Wait for the weights and for independent results before changing any contract.
Control over a security tool is not a line on a spec sheet. It is a duty you take on.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Mistral AI.





