LessWrong AI
· Communities
Flipping the eval on its head
An eval is a product. Typically, its 1 x n or k x n where there are n samples and 1 or k different language models. This briefing will argue that we’d like to see k x n x m evals, or however many dimensions.This post is pitching an ambitious way to spend tokens on cyberhardening. If its not viable at current capabiliti