← Latest reporting

A model-safety claim needs a release-level evidence ledger

SemiAnalysis found public, model-specific safety results for 31 of 857 Chinese AI releases it reviewed. Procurement teams should ask for evidence tied to the exact version, not a lab-wide assurance.

Policy, Standards and GovernanceAI Capability Frontier
A flat blue field of release tiles contrasts with a small group linked to red evidence tabs.
Conceptual illustration generated with AI under editorial direction; it does not depict a real event.

What happened

SemiAnalysis counted 857 releases from nine Chinese developers through 15 September 2026 and found 31 with a public safety result matched to that release; nine had results at or before launch.

Why it matters

The study measures public disclosure, not whether private testing occurred. Its decision value is the release-level matching rule: a general safety statement cannot establish evidence for the model actually being deployed.

SemiAnalysis published an original release census on 8 October, covering 857 identifiable releases from nine Chinese developers between 2021 and 15 September 2026. It counted a disclosure only when a quantitative or substantive safety result could be matched to a named release. On that rule, 31 releases, or 3.6%, had a public result and nine, or 1.1%, had one at or before launch. Reuters independently reported the findings on 9 October.

Match evidence to the deployed version

The useful control is not a percentage target. It is a ledger joining the exact model identifier, release date, weights or API snapshot, evaluation method, result date and any deployment restriction. A statement that a model family was evaluated should not automatically cover a smaller variant, later snapshot or differently tuned service.

SemiAnalysis explicitly says “not found” means no public result in the materials checked, not that a developer ran no private test. Release naming also differs across companies, so per-company rates are not rankings. Procurement and assurance teams should retain both limitations instead of converting the census into a claim about comparative safety.

Ask a supplier to populate the ledger for the version in the contract. Test one rollback and one silent model update: can the organisation show which evidence still applies, who accepted the gap and when a fresh result is due? Record a failed match as an explicit assurance gap rather than treating the absence of public evidence as evidence of danger. If the answer lives in marketing copy or a generic model card, the evidence is not yet bound to the production decision.