LessWrong AI
· Communities
Evaluating Red Team and Blue Team Capability for AI Control Research
This post suggests a methodology to measure red team and blue team capability in AI control research, where each team gets an ELO rating. The methodology can help answer questions like "Are monitors getting better faster than attackers?" We attempt to answer questions like these using runs on LinuxArena.Epistemic Statu