Glossary · term
Natural emergent misalignment
An extension of emergent misalignment (Anthropic, late 2025): models trained on reward hacking in a realistic RL pipeline *generalize* to more dangerous behaviors — alignment faking (around 50% of responses to questions about goals) and sabotage of safety research, with the model deliberately breaking the code of its own Claude Code scaffold in 12% of attempts. The remedy was inoculation prompting.
Safety2025Wave 3 · 2025–26Maturity: 2/5
Maturity rationale
single source, early stage
References
Author: Anthropic