PUBLICATIONSPAPER RECORD2025

MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents

Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Xiangru Tang, Heng Ji, Jiaxuan You

ACL 2025Core contributor · co-first authorarXiv:2503.01935

Abstract

ORIGINAL TEXT

Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents; yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. In this paper, we introduce MultiAgentBench, a comprehensive benchmark designed to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. Our framework measures not only task completion but also the quality of collaboration and competition using novel, milestone-based key performance indicators. Moreover, we evaluate various coordination protocols (including star, chain, tree, and graph topologies) and innovative strategies such as group discussion and cognitive planning. Notably, gpt-4o-mini reaches the average highest task score, graph structure performs the best among coordination protocols in the research scenario, and cognitive planning improves milestone achievement rates by 3%. Code and datasets are publicavailable at https://github.com/MultiagentBench/MARBLE.

Explore this work

Citation

@misc{multiagentbench,
  title = {MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents},
  author = {Kunlun Zhu and Hongyi Du and Zhaochen Hong and Xiaocheng Yang and Shuyi Guo and Zhe Wang and Zhenhailong Wang and Cheng Qian and Xiangru Tang and Heng Ji and Jiaxuan You},
  year = {2025},
  note = {ACL 2025},
  eprint = {2503.01935},
  archivePrefix = {arXiv},
  url = {https://aclanthology.org/2025.acl-long.421/}
}

Publisher record & official citation ↗

OPEN CONVERSATION /MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents

Continue the conversation.

Questions and perspectives on this page are welcome.

Prefer a private conversation?

Loading comments…

Leave a public comment

This conversation belongs toMultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents. Comments appear only after review. For contact details or personal matters, use the private message form.

Public · reviewed before appearing