r/MachineLearning
· Communities
New LLM Coordination Benchmark – Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R]
Can LLM agents coordinate in long-horizon, open-ended worlds? We evaluate 13 modern LLMs in a new benchmark where agents must work together to explore, communicate, trade resources, craft tools, build structures, and fight mobs. TL;DR: Most agents struggle, averaging only ~6% normalised return. Yet on the hardest setti