LessWrong AI
· Communities
Item Response Theory for AI Safety
TLDR:Many important decisions for safety depend on or are influenced by benchmark scores. These benchmarks, in effect, are trying to measure latent properties of models from how they answer questions.Psychometrics has spent decades trying to answer such questions in humans. We can probably steal some tools from this li