Skip to content
LessWrong AI · Communities

Tracing causal structure in LLM-generated text: a different lens on the Dallas circuit

The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM.I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of hidden activations. The resulting attribution graph looks much more structured than I expected:Below is an