Skip to content
LessWrong AI · Communities

Measuring Activation Control in LLMs

TL;DRInspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitor