Abstract
Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the goal. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons. The agent executes one action and replans with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Kong, D., Sun, G., Cheng, S., Xie, S., Pang, B., Xie, J., Geng, T., Ding, C., & Wu, Y. N. (2026). Learning to Plan from Random Exploration. https://omanscience.com/en/articles/learning-to-plan-from-random-exploration
MLA 9
Kong, Deqian, et al. "Learning to Plan from Random Exploration." https://omanscience.com/en/articles/learning-to-plan-from-random-exploration.
Chicago (author–date)
Kong, Deqian, Guangyan Sun, Sheng Cheng, Sirui Xie, Bo Pang, Jianwen Xie, Tony Geng, Caiwen Ding, and Ying Nian Wu. 2026. "Learning to Plan from Random Exploration." https://omanscience.com/en/articles/learning-to-plan-from-random-exploration.
Harvard
Kong, D., Sun, G., Cheng, S., Xie, S., Pang, B., Xie, J., Geng, T., Ding, C. and Wu, Y. N. (2026) 'Learning to Plan from Random Exploration', Available at: https://omanscience.com/en/articles/learning-to-plan-from-random-exploration.
Vancouver
Kong D, Sun G, Cheng S, Xie S, Pang B, Xie J, et al. Learning to Plan from Random Exploration. https://omanscience.com/en/articles/learning-to-plan-from-random-exploration
IEEE
D. Kong, G. Sun, S. Cheng, S. Xie, B. Pang, J. Xie, T. Geng, C. Ding, and Y. N. Wu, "Learning to Plan from Random Exploration," https://omanscience.com/en/articles/learning-to-plan-from-random-exploration.