LessWrong AI
· Communities
Can Recursive Self-Report Probing Detect Emergent Misalignment?
In this post, I summarize the findings from my work, which I did as part of the BlueDot AI Safety Course. The full code is available here. BackgroundBetley et al. (2025) showed that fine-tuning an LLM on insecure code not only learns to write just insecure code, but the model also starts expressing harmful values, dism