LessWrong AI
· Communities
The Geometry of Yes: Mapping Sycophancy Inside an LLM's Emotion Space
SummaryLLMs have internal emotion representations that causally shape their behaviour. This was recently observed in Claude: positive emotions like happy and loving are linked to sycophancy, and you can steer the model by directly manipulating these directions in activation space. But some questions were left open. Her