Skip to content
Alignment Forum · Communities

Should we benchmark conceptual capabilities using judgment prediction tasks?

A bunch of conceptual reasoning tasks involve very subjective judgments, which makes them poorly suited for benchmarking AI capabilities. For example, it seems unreasonable to benchmark how well AIs can predict the probability of misaligned AI takeover. Perhaps instead we should measure capabilities by explicitly instr