Skip to content
LessWrong AI · Communities

When Role-playing, Do Models Believe What They Say?

TL;DRWhen a model role-plays a persona, does it only change what it says, or also what it internally represents as true?To study this, we induce personas in five ways: prompting, in-context learning (ICL), supervised fine-tuning (SFT), Open Character Training (OCT), and Emergent Misalignment (EM). We measure internaliz