LessWrong AI
· Communities
Internal State Control is a General Property of LLMs
tl;dr:Lindsey 2025 found models can modulate their internal states: when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept is more present than when instructed to not think about it.Internal state controllability appears to be a general property of LLMs: the ef