Skip to content
LessWrong AI · Communities

A Score Is Not Understanding: toward a richer toolkit for model evaluations

We must take great care not to ignore the things that are not easily quantified - Brian Christian, The Alignment ProblemIntroductionModel evaluations have a problem. This isn't news [1]- the AI safety and research fields have known for years that current approaches to assessing the capability and safety of frontier mod