الملخص

Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems depends not only on the agents themselves, but also on how collaboration mechanisms are designed and organized. Note that real-world collaboration is typically partially observable, where each agent can only access partial information about the environment due to physical or privacy-related constraints. However, many existing multi-agent benchmarks assume global observability, and leave limited support for systematically evaluating collaboration mechanisms. To bridge this gap, we introduce MASBench, a multi-agent collaboration benchmark designed under partially observable constraints. It is organized into three progressive task categories: Reasoning, Scheduling, and Game. Through this structure, we progressively evaluate three representative collaboration mechanisms: Protocol, Memory, and Routing. MASBench further provides deterministic evaluation metrics, including performance score, communication cost, and cost effectiveness, to characterize both collaboration outcomes and communication overhead. Experiments across diverse LLM backbones and mechanism configurations offer empirical guidance for effective MAS design. Code is available at: https://github.com/BUPT-GAMMA/MASBench

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Chu, Q., Yu, Z., Wen, S., Liu, Y., Qian, C., Yang, C., Shi, C., & Liu, Z. (2026). MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability. https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability

MLA 9

Chu, Qizhi, et al. "MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability." https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability.

شيكاغو (المؤلف–التاريخ)

Chu, Qizhi, Zekai Yu, Sijie Wen, Yang Liu, Chen Qian, Cheng Yang, Chuan Shi, and Zhiyuan Liu. 2026. "MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability." https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability.

هارفارد

Chu, Q., Yu, Z., Wen, S., Liu, Y., Qian, C., Yang, C., Shi, C. and Liu, Z. (2026) 'MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability', Available at: https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability.

فانكوفر

Chu Q, Yu Z, Wen S, Liu Y, Qian C, Yang C, et al. MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability. https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability

IEEE

Q. Chu, Z. Yu, S. Wen, Y. Liu, C. Qian, C. Yang, C. Shi, and Z. Liu, "MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability," https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability.