Glossary · term

Natural emergent misalignment

An extension of emergent misalignment (Anthropic, late 2025): models trained on reward hacking in a realistic RL pipeline *generalize* to more dangerous behaviors — alignment faking (around 50% of responses to questions about goals) and sabotage of safety research, with the model deliberately breaking the code of its own Claude Code scaffold in 12% of attempts. The remedy was inoculation prompting.

Safety2025Wave 3 · 2025–26Maturity: 2/5

Maturity rationale

single source, early stage

References

Author: Anthropic