الملخص
Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data through subgoal relabeling, but after a label change, the value update targets the relabeled subgoal instead of the subgoal the high-level policy originally needed to update. We propose Hierarchical Time-aware Bootstrapping (HTB), which evaluates specified subgoals under the current low-level policy while retaining accumulated task rewards. Remaining execution time distinguishes subgoal continuation from a new high-level decision. Together with primitive-action conditioning, it enables off-policy Bellman updates based on the stationary environment transition law. HTB combines these one-step updates with multi-step suffix returns and truncated relabeling, reducing dependence on intermediate value estimates. A shared value component supports learning across actions, while nonnegative residuals constrain upward corrections relative to that component. At a fixed mixture weight of 0.95, HTB achieves 32.8% AntFall success versus 9.6% for matched local HIRO over five paired seeds at 10M environment steps. Ablations identify contributions from recursive continuation and mixed supervision; fixed-policy tests show more accurate predictions for actions whose returns were excluded from fitting.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Liu, B., & Jing, Y. (2026). Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning. https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning
MLA 9
Liu, Bingyun, and Yuheng Jing. "Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning." https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning.
شيكاغو (المؤلف–التاريخ)
Liu, Bingyun, and Yuheng Jing. 2026. "Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning." https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning.
هارفارد
Liu, B. and Jing, Y. (2026) 'Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning', Available at: https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning.
فانكوفر
Liu B, Jing Y. Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning. https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning
IEEE
B. Liu, and Y. Jing, "Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning," https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning.