Skip to content
LessWrong AI · Communities

Inoculate or Reflect? Two training interventions under prompting, steering, and patching

Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models, contains a small experiment near the end that we found more interesting than the main findings.The technique is called Counterfactual Reflection Training (CRT). The model is fed a partial transcript in its context window,