LessWrong AI
· Communities
Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model
(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread, slightly expanded for LessWrong.)In one experiment, Sonnet 5 descr