Skip to content
LessWrong AI · Communities

Have models report provable security bugs in their environment

AIs are often deployed with limited permissions. They aren't allowed to reach the internet. Are given a limited set of files they can read or write. Aren't supposed to be able to read the held out evaluation test set. This could be during deployment or in training.Currently, when these guarantees fail, we find out only