Collaboration and competition
MultiAgentBench evaluates LLM teams through MARBLE’s interactive environments, including research, coding, database work, social interaction, Minecraft, and Werewolf. It measures task outcomes alongside milestone-based indicators of collaboration and competition, making intermediate coordination visible rather than reducing a run to its final answer.
Structure matters
The benchmark compares star, chain, tree, and graph communication structures, alongside group discussion and cognitive planning. This makes both the connections between agents and their coordination strategy explicit experimental choices.
Its findings are scenario-dependent. For example, the paper reports that graph coordination performs best in the research setting; that is not a claim that one topology is always best or that adding agents necessarily improves performance.
Paper and contribution
Hongyi Du is a core contributor and co-first author, including work on Werewolf Arena and theory-of-mind evaluation. This brief follows MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents, published at ACL 2025, and its contribution appendix. MARBLE contains the code and datasets.

Loading comments…