Skip to content
LessWrong AI · Communities

Foundation Models for Oversight

Cross-posted from the Transluce blog. To oversee an AI model, we'd ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldn't admit to if asked directly? Does the model treat a user differently once it infers something about their identity