Abstract
Language-guided navigation for terrestrial-aerial bimodal robots requires selecting routes and locomotion modes that match scene context and task intent. Generated videos can represent such motion sequences, but recovering metrically consistent navigation references from them is challenging because of scale ambiguity and axis-dependent geometric distortions. We present TADreamer, a zero-shot framework that grounds video-imagined navigation in measured geometry without task-specific training or fine-tuning. A vision-language model translates onboard observations and instructions into navigation prompts, selects valid generated videos, and provides corrective feedback when regeneration is needed. The selected video is reconstructed into 3D waypoints annotated with terrestrial or aerial modes. A two-stage calibration procedure uses field-of-view constraints to initialize scale estimation, then refines axis-dependent scales, rotation, and translation by registering the reconstructed point cloud to measured geometry. The calibrated waypoints and mode labels guide a planner that incorporates measured geometry for robot execution. Real-world experiments demonstrate navigation across seven indoor and outdoor scenarios. With five candidates per round, usable videos are obtained within two rounds in all seven scenarios. On the calibration observations, our method reduces mean absolute depth error by 87.7% and mean absolute relative depth error by 86.3% compared with NavDreamer.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Li, X., Lai, T., Huang, X., Pang, R., Shen, S., Chen, J., Pan, Z., Xu, C., Gao, F., & Cao, Y. (2026). TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination. https://omanscience.com/en/articles/tadreamer-zero-shot-language-guided-3d-navigation-for-terrestrial-aerial-bimodal-robots-via-video-imagination
MLA 9
Li, Xiangyu, et al. "TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination." https://omanscience.com/en/articles/tadreamer-zero-shot-language-guided-3d-navigation-for-terrestrial-aerial-bimodal-robots-via-video-imagination.
Chicago (author–date)
Li, Xiangyu, Tiancheng Lai, Xijie Huang, Ruitian Pang, Siqi Shen, Juncheng Chen, Zaisheng Pan, Chao Xu, Fei Gao, and Yanjun Cao. 2026. "TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination." https://omanscience.com/en/articles/tadreamer-zero-shot-language-guided-3d-navigation-for-terrestrial-aerial-bimodal-robots-via-video-imagination.
Harvard
Li, X., Lai, T., Huang, X., Pang, R., Shen, S., Chen, J., Pan, Z., Xu, C., Gao, F. and Cao, Y. (2026) 'TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination', Available at: https://omanscience.com/en/articles/tadreamer-zero-shot-language-guided-3d-navigation-for-terrestrial-aerial-bimodal-robots-via-video-imagination.
Vancouver
Li X, Lai T, Huang X, Pang R, Shen S, Chen J, et al. TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination. https://omanscience.com/en/articles/tadreamer-zero-shot-language-guided-3d-navigation-for-terrestrial-aerial-bimodal-robots-via-video-imagination
IEEE
X. Li, T. Lai, X. Huang, R. Pang, S. Shen, J. Chen, Z. Pan, C. Xu, F. Gao, and Y. Cao, "TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination," https://omanscience.com/en/articles/tadreamer-zero-shot-language-guided-3d-navigation-for-terrestrial-aerial-bimodal-robots-via-video-imagination.