Skip to content
LessWrong AI · Communities

Necessity Protects Chain of Thought Monitoring by Prevention, Not Disclosure

Preface: The case for reading chain-of-thought is that it is cheap, scalable and simple, it's just sit and read what the model wrote and catch it before it does something terrible. The case against is that we don't have any guarantee the text is the reason. So I tried to look for a boundary by taking some class of task