Skip to content
arXiv cs.CL · Papers

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

arXiv:2607.06008v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire executi