Skip to content
LessWrong AI · Communities

AI Safety at the Frontier: Paper Highlights of May & June 2026

tl;drPaper of the month:Anthropic’s Jacobian lens reveals that models have a sparse workspace of verbalizable concepts that causally carries multi-hop reasoning and surfaces hidden cognition — as opposed to other, more automatic mental processing.Research highlights:Natural language autoencoders translate activations i