الملخص

Goal-oriented Vision-and-Language Navigation (VLN) requires agents to locate and reach targets described in natural language without prescribed routes. Air--ground collaboration is valuable for tasks requiring both wide-area search and fine-grained localization. However, systematic study of goal-oriented air--ground collaborative VLN remains limited by the lack of large-scale, diverse benchmarks and two core challenges: 1) substantial differences between aerial and ground views, together with useful observations becoming unavailable as navigation proceeds, make it difficult to maintain spatially consistent context across platforms and over time; and 2) asymmetric spatial observability makes ground perception locally detailed but spatially limited and aerial perception broad but locally coarse, limiting the reliability of single-platform planning. To address these limitations, we introduce AirGroundVLN, a benchmark containing 10,281 navigation episodes and 955 target instances across 19 Unreal Engine environments, with seen/unseen splits and an aerial-visibility protocol for systematic evaluation. Alongside the benchmark, we propose AG-CoNAV, a trainable reference framework comprising two key components: Spatiotemporally Anchored Collaborative Memory (SACM) and Aerial-Guided Regional-to-Local Planning (AGRLP). SACM maintains and retrieves spatially consistent historical context across aerial and ground observations. Meanwhile, AGRLP combines regional aerial guidance with fine-grained ground navigation. Extensive experiments demonstrate the effectiveness of AG-CoNAV and establish AirGroundVLN as a comprehensive benchmark for future exploration.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zeng, Z., Wu, Q., Suo, W., Wu, M., Zhang, B., Yu, H., & Wang, P. (2026). AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation. https://omanscience.com/ar/articles/airgroundvln-a-large-scale-benchmark-for-goal-oriented-air-ground-collaborative-vision-and-language-navigation

MLA 9

Zeng, Zhenxuan, et al. "AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation." https://omanscience.com/ar/articles/airgroundvln-a-large-scale-benchmark-for-goal-oriented-air-ground-collaborative-vision-and-language-navigation.

شيكاغو (المؤلف–التاريخ)

Zeng, Zhenxuan, Qingle Wu, Wei Suo, Maojia Wu, Bairong Zhang, Hangzheng Yu, and Peng Wang. 2026. "AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation." https://omanscience.com/ar/articles/airgroundvln-a-large-scale-benchmark-for-goal-oriented-air-ground-collaborative-vision-and-language-navigation.

هارفارد

Zeng, Z., Wu, Q., Suo, W., Wu, M., Zhang, B., Yu, H. and Wang, P. (2026) 'AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation', Available at: https://omanscience.com/ar/articles/airgroundvln-a-large-scale-benchmark-for-goal-oriented-air-ground-collaborative-vision-and-language-navigation.

فانكوفر

Zeng Z, Wu Q, Suo W, Wu M, Zhang B, Yu H, et al. AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation. https://omanscience.com/ar/articles/airgroundvln-a-large-scale-benchmark-for-goal-oriented-air-ground-collaborative-vision-and-language-navigation

IEEE

Z. Zeng, Q. Wu, W. Suo, M. Wu, B. Zhang, H. Yu, and P. Wang, "AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation," https://omanscience.com/ar/articles/airgroundvln-a-large-scale-benchmark-for-goal-oriented-air-ground-collaborative-vision-and-language-navigation.