LessWrong AI
· Communities
Three years of progress in 500 lines of code
TL;DRThere is some consensus that LLMs are bad at hard-to-verify tasks. The question is whether models are getting better at them over time. As a motivating example, I gave one research-reproduction task to 12 models spanning three years of progress, to illustrate how (1) what looked like an emergent capability was a p