Skip to content
arXiv cs.AI · Papers

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

arXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces. Even nominally open-ended tasks can often be solved by retrieving a well-known recipe and tuning a