الملخص
Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems depends not only on the agents themselves, but also on how collaboration mechanisms are designed and organized. Note that real-world collaboration is typically partially observable, where each agent can only access partial information about the environment due to physical or privacy-related constraints. However, many existing multi-agent benchmarks assume global observability, and leave limited support for systematically evaluating collaboration mechanisms. To bridge this gap, we introduce MASBench, a multi-agent collaboration benchmark designed under partially observable constraints. It is organized into three progressive task categories: Reasoning, Scheduling, and Game. Through this structure, we progressively evaluate three representative collaboration mechanisms: Protocol, Memory, and Routing. MASBench further provides deterministic evaluation metrics, including performance score, communication cost, and cost effectiveness, to characterize both collaboration outcomes and communication overhead. Experiments across diverse LLM backbones and mechanism configurations offer empirical guidance for effective MAS design. Code is available at: https://github.com/BUPT-GAMMA/MASBench
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Chu, Q., Yu, Z., Wen, S., Liu, Y., Qian, C., Yang, C., Shi, C., & Liu, Z. (2026). MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability. https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability
MLA 9
Chu, Qizhi, et al. "MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability." https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability.
شيكاغو (المؤلف–التاريخ)
Chu, Qizhi, Zekai Yu, Sijie Wen, Yang Liu, Chen Qian, Cheng Yang, Chuan Shi, and Zhiyuan Liu. 2026. "MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability." https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability.
هارفارد
Chu, Q., Yu, Z., Wen, S., Liu, Y., Qian, C., Yang, C., Shi, C. and Liu, Z. (2026) 'MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability', Available at: https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability.
فانكوفر
Chu Q, Yu Z, Wen S, Liu Y, Qian C, Yang C, et al. MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability. https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability
IEEE
Q. Chu, Z. Yu, S. Wen, Y. Liu, C. Qian, C. Yang, C. Shi, and Z. Liu, "MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability," https://omanscience.com/ar/articles/masbench-benchmarking-llm-based-multi-agent-collaboration-under-partial-observability.