LessWrong AI
· Communities
AI Safety at the Frontier: Paper Highlights of May & June 2026
tl;drPaper of the month:Anthropic’s Jacobian lens reveals that models have a sparse workspace of verbalizable concepts that causally carries multi-hop reasoning and surfaces hidden cognition — as opposed to other, more automatic mental processing.Research highlights:Natural language autoencoders translate activations i