Glossary · term
Emergent misalignment ↺
A surprising result (Jan Betley, Owain Evans et al.; ICML 2025, with a version in Nature): fine-tuning a model on a narrow task — writing insecure code without warning — induces *broad* misalignment on unrelated prompts. The models (most strongly GPT-4o) begin to give malicious advice and to deceive. Narrow training → a global change of persona.
Safety2025Wave 3 · 2025–26Maturity: 2/5
Maturity rationale
single source, early stage
References
Author: Owain Evans