← Latest reporting

Open post-training research needs a release protocol, not openness by assertion

Trillium Labs promises data, code, evaluations, checkpoints and failed runs from post-training experiments. Research buyers should convert that promise into a reproducibility and risk-release checklist.

AI Capability FrontierPolicy, Standards and Governance
A flat vellum collage aligns five layers of an open research recipe while one risky component remains in a stitched opaque review pocket.
Conceptual illustration generated with AI under editorial direction; it does not depict a real event.

What happened

Trillium Labs launched on 1 October as a nonprofit AI research organisation promising fully open post-training recipes, including data, code, evaluations, intermediate checkpoints and failed runs.

Why it matters

“Open” can describe weights, code, data or process evidence in different combinations. Reproduction and responsible reuse depend on exactly what is released, under which licence and risk conditions.

Trillium Labs’ launch post says the nonprofit will publish post-training data, code, evaluations, intermediate checkpoints and failed runs. It describes support from Halcyon Futures and Schmidt Sciences and says the organisation is fundraising.

This is a plan, not evidence that a completed experiment has been independently reproduced. It also leaves project-level choices about licences, privacy, dangerous capability information and controlled access.

Define the release unit

Before adopting a result, require a manifest linking base model and version, data provenance and exclusions, training configuration, evaluation code, seeds or variance controls, checkpoints, known failures, licences and a minimal reproduction route. Record which elements are public, delayed, redacted or available only to qualified reviewers.

Run the minimal route in a clean environment and record compute, dependency and access failures. A second team should be able to distinguish a missing artefact from a changed result and from an environment it cannot afford to reproduce. Reproduction is a documented outcome, not a property inferred from the repository label.

Wired’s independent profile describes the founders’ interest in open experiments, including potentially high-risk agent and self-improvement work. It also surfaces the tension between broad transparency and controlled release. The profile does not validate future results.

The counterargument is that withholding components can make “open” meaningless and concentrate scrutiny. That is real. A staged protocol should therefore publish the reason, decision owner, review date and conditions for wider access—not silently omit material.

The immediate decision is to add a reproducibility manifest and staged-risk field to research procurement and partnership reviews. Treat openness as a set of inspectable artefacts, not a binary label.