Skip to content
LessWrong AI · Communities

Eliciting hidden knowledge from monitors with NLAs

Aleksandr Bowkis* and David Africa*TL;DRChain of thought (CoT) monitorability may be fragile, and natural language autoencoders (NLAs) may provide a helpful, decorrelated monitoring surface.We tried to read NLAs from the monitor itself, where the NLA readout surfaces what the monitor internally represents while judging