- SemiAnalysis checked 857 releases from nine Chinese AI labs. Developers published a matching safety result for 31 (3.6%), and only 9 (1.1%) had one at launch.
- The analysts argue China's rules govern applications and outputs, not frontier capability. Following the rules therefore says little about how dangerous a model is.
- Ask any AI vendor for version-specific safety results dated at or before release, and treat a missing result as an open question.
A rulebook is not a test result. That is the lesson in a new SemiAnalysis study of Chinese AI labs, and it reaches well beyond China. A company can follow every rule written for it and still never show what its model can do. If you choose which AI to build on, that gap matters.
What the researchers counted
SemiAnalysis built its own list of 857 model releases from nine Chinese developers. The list runs from 2021 to 15 September 2026. Four are large firms: ByteDance, Alibaba, Tencent and Baidu. Five younger companies complete the set: DeepSeek, Moonshot, Zhipu (Z.ai), MiniMax and StepFun.
For each release, the team read the developer's own model cards, release notes and technical reports. It set a strict bar. A release scored only if the developer published real findings about that exact model. Topics included jailbreaks, toxicity, privacy, refusals and dangerous capability. Saying a model was "safety-trained" earned nothing.
The team also named the limits of its work. "Not found" means not found in the materials checked. It does not mean "not tested." Companies label variants at different levels of detail, so the per-company rates are only indicative.
What the results show
Another 16 releases were documented only after launch. The typical wait was 42 days. The slowest case took 349 days: DeepSeek-R1 only has a checked safety appendix from January 2026. A further 813 releases, 94.9% of the total, have no safety disclosure at all.
Output grew fast. Releases rose from 3 in Q1 2023 to 101 in Q3 2025. Disclosure did not follow. No quarter had more than 7 releases with any safety result. No quarter had more than 3 with a result at launch.
No single company explains the gap. Alibaba published the most releases, 238, yet only 7 have a result. Its count includes every Qwen size and snapshot. Tencent has 1 in 133 and Baidu 1 in 49. The younger firms did somewhat better. They account for 20 results across 317 releases (6.3%). The four large firms account for 11 across 540 (2%). Even so, SemiAnalysis says no company does this routinely.
Reasoning models are the fastest-advancing category. Of these, 93% have no published results. SemiAnalysis also found that no Chinese frontier text model has launched with tests for dangerous capabilities in the areas named in the IDAIS statements. Zhipu's GLM-5.3 cyber-capability note came closest.
Why compliance says little about frontier safety
SemiAnalysis lists mandates issued from September 2025 to October 2026. They cover content labeling, AI-companion services, agent rules, ethics review and inspection powers. The analysts say each one governs what AI says or what it does to people.
None requires a lab to test frontier capability. A lab could meet every rule on the list without running a single dangerous-capability evaluation. That is the report's link between the rules and the disclosure data. The report does not say the rules caused the low disclosure rate. It notes that no policy milestone leaves a mark on the disclosure lines.
Other regimes draw the line differently. The EU attaches systemic-risk duties at 10²⁵ FLOP, a measure of training compute. California's SB 53 applies once training reaches 10²⁶ operations. It asks developers for a frontier framework and for incident reports within 15 days. SemiAnalysis finds nothing comparable in China.
The week of 14 September shows the pattern. China's AI Safety Governance Framework 3.0 added agent and embodied-intelligence risks. Authorities then issued application guides for education, health care and broadcasting. A draft agent guide followed, covering identity, least privilege and human-approval steps. That made seven texts in five days. On the frontier itself, SemiAnalysis sees no change.
The analysts add that China's real approach is built around speed. They point to the Framework's first stated priority: promoting AI innovation and development. They also write that nobody in China's AI industry is slowing down, and nobody is asking them to.
What the leaders say
SemiAnalysis also logged what senior people at the nine labs have said about safety. It found 65 public items between January 2023 and 21 September 2026. Only 15 came from founders, CEOs or chief scientists who engaged with frontier safety. Nine of those voiced concern and proposed something. Six of the nine came from Zhipu alone. At five labs, the founder or CEO had said nothing.
What to ask your team
The report does not compare Chinese labs with Western ones. It does not show who is safer. It shows what evidence exists, and you can apply the same test to any vendor.
First, list every model your products and staff use, and who built it. Second, ask for a safety result tied to the exact version you run, dated at or before release. A result for a flagship does not cover a smaller size or a later snapshot.
Third, treat a missing result as an open question. Run your own tests before the model touches customer data. Fourth, if the model powers an agent, limit its permissions and require human approval for sensitive actions.
Compliance tells you a vendor followed the rules. Only a dated, version-specific test tells you what the model can do.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: SemiAnalysis.





