Skip to content
LessWrong AI · Communities

Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes

Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data.TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can catch a sleeper agent about to defect despite knowing that directly asking the model fails completely. F