Skip to content
LessWrong AI · Communities

NLAs read thoughts beyond the J-space

TLDR:On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked if it sees a hidden thought, the model says "No, let's move on"; the NLA reads "elephants", "secrecy", "love"!I reproduced Anthropic's