Skip to content
LessWrong AI · Communities

We should consider how long monitoring is reliable for during RL

Epistemic status: I am new to AI Safety and am writing blogs to gain context. This blog post was formed from discussions with Aidan Ewart and Jonathan Bostock, but they do not necessarily endorse this post.TL;DRGiven recent examples of AI misbehaviour during training episodes, AI companies might want to start using mon