Skip to content
arXiv cs.CL · Papers

PACE: A Proxy for Agentic Capability Evaluation

arXiv:2607.02032v1 Announce Type: cross Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual c