arXiv cs.CL
· Papers
Evo-Bench: Can Language Models Improve Agent Harness?
arXiv:2608.09096v2 Announce Type: replace Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However,