Skip to content
LessWrong AI · Communities

Self-monitoring doesn't scale (without these 3 countermeasures)

The untrusted monitoring protocol, as defined and evaluated in the AI control literature [1, 2, 3], looks like this:In comparison, the current monitoring setups at most frontier labs look like this:Before criticizing this, I first want to say that monitoring >> no monitoring. A year ago the 2nd box didn't meaningfully