Skip to content
arXiv cs.AI · Papers

Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents

arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs determine how models invoke tools, maintain interaction history, and recover from failures. Consequently, effective match