Skip to content
arXiv cs.CL · Papers

Evo-Bench: Can Language Models Improve Agent Harness?

arXiv:2608.09096v2 Announce Type: replace Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However,