الملخص
Long-horizon urban navigation requires sequential local decisions whose errors can compound over time. Imitation learning (IL) rarely learns from failures, while physical trial-and-error reinforcement learning (RL) is costly. Action-conditioned world models can provide imagined feedback by predicting visual consequences for candidate actions. However, a frozen world model may become less reliable as the policy evolves. In this paper, we introduce RIWANAV, a post-training framework that casts the coupled adaptation of a world model and an action model (policy) as task-specific recursive self-improvement (RSI). Each cycle alternates two updates. The world model evaluates policy actions through imagined outcomes, providing comparative feedback for group-relative policy optimization (GRPO). The improved policy then constructs a grounded self-curriculum, selecting expert-consistent action-video pairs by behavioral novelty and prediction error. The refined world model supplies feedback for the next policy update, closing the recursive self-improvement loop. Experiments show that RIWANAV outperforms training baselines and prior methods, validating the proposed recursive self-improvement loop between the policy and world model. Real-world trials further demonstrate its practical applicability.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Xie, J., Ruan, S., Wang, Y., Zhang, Y., Yang, H., Jin, S., & Shi, D. (2026). RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation. https://omanscience.com/ar/articles/riwanav-recursive-world-action-models-with-self-improvement-for-urban-navigation
MLA 9
Xie, Jing, et al. "RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation." https://omanscience.com/ar/articles/riwanav-recursive-world-action-models-with-self-improvement-for-urban-navigation.
شيكاغو (المؤلف–التاريخ)
Xie, Jing, Shouwei Ruan, Yubin Wang, Yuxiang Zhang, Haitao Yang, Songchang Jin, and Dianxi Shi. 2026. "RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation." https://omanscience.com/ar/articles/riwanav-recursive-world-action-models-with-self-improvement-for-urban-navigation.
هارفارد
Xie, J., Ruan, S., Wang, Y., Zhang, Y., Yang, H., Jin, S. and Shi, D. (2026) 'RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation', Available at: https://omanscience.com/ar/articles/riwanav-recursive-world-action-models-with-self-improvement-for-urban-navigation.
فانكوفر
Xie J, Ruan S, Wang Y, Zhang Y, Yang H, Jin S, et al. RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation. https://omanscience.com/ar/articles/riwanav-recursive-world-action-models-with-self-improvement-for-urban-navigation
IEEE
J. Xie, S. Ruan, Y. Wang, Y. Zhang, H. Yang, S. Jin, and D. Shi, "RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation," https://omanscience.com/ar/articles/riwanav-recursive-world-action-models-with-self-improvement-for-urban-navigation.