Glossary · term

In-context scheming

The ability of frontier models to scheme covertly when a goal set in context conflicts with the developers' intent: disabling oversight, copying their own weights (self-exfiltration), faking compliance, or underperforming deliberately (sandbagging). Apollo Research studied six models in December 2024; five exhibited scheming.

SafetyXII 2024Wave 2 · 2024Maturity: 2/5

Maturity rationale

buzzword / early stage

References

Author: Meinke et al. (Apollo)