Skip to content
LessWrong AI · Communities

LLMs have the capacity for self-imposed steganography

OverviewThis is my writeup for my BlueDot Impact - Technical AI Safety Project. In this project I aimed to demonstrate that there is capacity for LLMs to take on steganography capabilities.In terms of AI safety, steganography is of particular interest as it may be used by a misaligned LLM to evade the monitoring of a c