Abstract
Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning. Project page: https://jiaming.im/VideoSocNav
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wang, J., Nguyen, D. T., Chen, J., Shcherbyna, V., Liu, D., Shen, Z., & Soh, H. (2026). Learning Social Navigation from Internet Videos in the Policy State Space. https://omanscience.com/en/articles/learning-social-navigation-from-internet-videos-in-the-policy-state-space
MLA 9
Wang, Jiaming, et al. "Learning Social Navigation from Internet Videos in the Policy State Space." https://omanscience.com/en/articles/learning-social-navigation-from-internet-videos-in-the-policy-state-space.
Chicago (author–date)
Wang, Jiaming, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna, Diwen Liu, Zhengcheng Shen, and Harold Soh. 2026. "Learning Social Navigation from Internet Videos in the Policy State Space." https://omanscience.com/en/articles/learning-social-navigation-from-internet-videos-in-the-policy-state-space.
Harvard
Wang, J., Nguyen, D. T., Chen, J., Shcherbyna, V., Liu, D., Shen, Z. and Soh, H. (2026) 'Learning Social Navigation from Internet Videos in the Policy State Space', Available at: https://omanscience.com/en/articles/learning-social-navigation-from-internet-videos-in-the-policy-state-space.
Vancouver
Wang J, Nguyen DT, Chen J, Shcherbyna V, Liu D, Shen Z, et al. Learning Social Navigation from Internet Videos in the Policy State Space. https://omanscience.com/en/articles/learning-social-navigation-from-internet-videos-in-the-policy-state-space
IEEE
J. Wang, D. T. Nguyen, J. Chen, V. Shcherbyna, D. Liu, Z. Shen, and H. Soh, "Learning Social Navigation from Internet Videos in the Policy State Space," https://omanscience.com/en/articles/learning-social-navigation-from-internet-videos-in-the-policy-state-space.