arXiv cs.AI
· Papers
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
arXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces. Even nominally open-ended tasks can often be solved by retrieving a well-known recipe and tuning a