Skip to content
X · @emollick · X / Twitter

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans Va…

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humansValidated benchmarks need to have human (ideally multiple humans) baselines. It is increasingly hard & pricey to do, but important