الملخص
Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Chen, Y., Zhang, Y., Witbrock, M., & Hu, S. (2026). Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching. https://omanscience.com/ar/articles/loop-free-inverse-reinforcement-learning-via-sequential-value-recovery-with-q-score-matching
MLA 9
Chen, Yang, et al. "Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching." https://omanscience.com/ar/articles/loop-free-inverse-reinforcement-learning-via-sequential-value-recovery-with-q-score-matching.
شيكاغو (المؤلف–التاريخ)
Chen, Yang, Yitan Zhang, Michael Witbrock, and Shuyue Hu. 2026. "Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching." https://omanscience.com/ar/articles/loop-free-inverse-reinforcement-learning-via-sequential-value-recovery-with-q-score-matching.
هارفارد
Chen, Y., Zhang, Y., Witbrock, M. and Hu, S. (2026) 'Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching', Available at: https://omanscience.com/ar/articles/loop-free-inverse-reinforcement-learning-via-sequential-value-recovery-with-q-score-matching.
فانكوفر
Chen Y, Zhang Y, Witbrock M, Hu S. Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching. https://omanscience.com/ar/articles/loop-free-inverse-reinforcement-learning-via-sequential-value-recovery-with-q-score-matching
IEEE
Y. Chen, Y. Zhang, M. Witbrock, and S. Hu, "Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching," https://omanscience.com/ar/articles/loop-free-inverse-reinforcement-learning-via-sequential-value-recovery-with-q-score-matching.