LessWrong AI
· Communities
Have models report provable security bugs in their environment
AIs are often deployed with limited permissions. They aren't allowed to reach the internet. Are given a limited set of files they can read or write. Aren't supposed to be able to read the held out evaluation test set. This could be during deployment or in training.Currently, when these guarantees fail, we find out only