الملخص
Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and executing continuous low-level actions. Existing methods often depend on LiDAR, panoramic cameras, or extra sensors; separate geometric-mapping and semantic-navigation visual representations can cause long-trajectory spatial-semantic inconsistencies. We propose LG-VLN, a monocular zero-shot framework with shared visual features and LangGraph-based state orchestration. An online feed-forward 3D reconstruction network predicts depth, camera poses, and dense point clouds for agent-pose estimation and global map fusion. Geometry and navigation share dense CleanDIFT features: semantic consistency rejects incorrect inter-frame correspondences, while target-instance constraints define visual references whose similarity combines with local BLIP-2 image-text relevance to form a semantic value map. LangGraph represents instruction parsing, geometric perception, semantic value updates, path planning, action execution, and failure recovery as a directed state graph with conditional transitions, persistent state, and modular recovery mechanisms. On a fixed 550-episode subset of the R2R-CE val-unseen split, LG-VLN achieves 21.3% success and 12.1% success weighted by path length. Ablations show shared semantic features improve navigation, further boosted by combining visual similarity and image-text relevance. Results establish shared visual representations and explicit state orchestration as effective for zero-shot VLN-CE using monocular RGB alone. Code will be publicly released for reproducibility.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zhao, J., Qiu, Y., Zhang, Z., Zhao, Z., & Xie, J. (2026). LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration. https://omanscience.com/ar/articles/lg-vln-a-zero-shot-vision-and-language-navigation-framework-with-langgraph-state-orchestration
MLA 9
Zhao, Jianhe, et al. "LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration." https://omanscience.com/ar/articles/lg-vln-a-zero-shot-vision-and-language-navigation-framework-with-langgraph-state-orchestration.
شيكاغو (المؤلف–التاريخ)
Zhao, Jianhe, Yanhua Qiu, Zhiyu Zhang, Zibo Zhao, and Jinhua Xie. 2026. "LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration." https://omanscience.com/ar/articles/lg-vln-a-zero-shot-vision-and-language-navigation-framework-with-langgraph-state-orchestration.
هارفارد
Zhao, J., Qiu, Y., Zhang, Z., Zhao, Z. and Xie, J. (2026) 'LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration', Available at: https://omanscience.com/ar/articles/lg-vln-a-zero-shot-vision-and-language-navigation-framework-with-langgraph-state-orchestration.
فانكوفر
Zhao J, Qiu Y, Zhang Z, Zhao Z, Xie J. LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration. https://omanscience.com/ar/articles/lg-vln-a-zero-shot-vision-and-language-navigation-framework-with-langgraph-state-orchestration
IEEE
J. Zhao, Y. Qiu, Z. Zhang, Z. Zhao, and J. Xie, "LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration," https://omanscience.com/ar/articles/lg-vln-a-zero-shot-vision-and-language-navigation-framework-with-langgraph-state-orchestration.