Skip to content
arXiv cs.AI · Papers

Measuring AI Ability to Complete Long Software Tasks

arXiv:2503.14499v4 Announce Type: replace Abstract: Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typi