Skip to content
LessWrong AI · Communities

Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift

In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness of a model without deploying it, but we cannot safely deploy it until we know it is not scheming. We expect sufficiently