LessWrong AI
· Communities
I don't think Claude is misaligned in 'Agentic Misalignment Summer 2026 – Motivated Mislabeling'
Anthropic recently published Agentic Misalignment Summer 2026The "whistleblowing" scenario has already been examined and found problematic. I started taking a look at the transcripts for some others. As far as I can tell, the objective of each agentic misalignment evaluation was to simulate a corrupted principal (inclu