Abstract

Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-aware metrics. We further present TRISS, a coordination-ready navigation system coupling an LLM-based subtask scheduler, a shared topological memory that turns each agent's exploration into team knowledge, and a conflict-aware execution mechanism that realizes simultaneous intentions as collision-free routes. Extensive experiments establish TRISS as a comprehensive baseline and reveal substantial room for improvement across scheduling, planning, and execution, highlighting the challenges of coordinating under MAVLN task constraints. Project page: https://xyz9911.github.io/mavln.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Xu, Y., & Liu, Z. (2026). Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method. https://omanscience.com/en/articles/systematic-multi-agent-vision-and-language-navigation-formulation-benchmark-and-method

MLA 9

Xu, Yunzhe, and Zhe Liu. "Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method." https://omanscience.com/en/articles/systematic-multi-agent-vision-and-language-navigation-formulation-benchmark-and-method.

Chicago (author–date)

Xu, Yunzhe, and Zhe Liu. 2026. "Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method." https://omanscience.com/en/articles/systematic-multi-agent-vision-and-language-navigation-formulation-benchmark-and-method.

Harvard

Xu, Y. and Liu, Z. (2026) 'Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method', Available at: https://omanscience.com/en/articles/systematic-multi-agent-vision-and-language-navigation-formulation-benchmark-and-method.

Vancouver

Xu Y, Liu Z. Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method. https://omanscience.com/en/articles/systematic-multi-agent-vision-and-language-navigation-formulation-benchmark-and-method

IEEE

Y. Xu, and Z. Liu, "Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method," https://omanscience.com/en/articles/systematic-multi-agent-vision-and-language-navigation-formulation-benchmark-and-method.