الملخص

Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data through subgoal relabeling, but after a label change, the value update targets the relabeled subgoal instead of the subgoal the high-level policy originally needed to update. We propose Hierarchical Time-aware Bootstrapping (HTB), which evaluates specified subgoals under the current low-level policy while retaining accumulated task rewards. Remaining execution time distinguishes subgoal continuation from a new high-level decision. Together with primitive-action conditioning, it enables off-policy Bellman updates based on the stationary environment transition law. HTB combines these one-step updates with multi-step suffix returns and truncated relabeling, reducing dependence on intermediate value estimates. A shared value component supports learning across actions, while nonnegative residuals constrain upward corrections relative to that component. At a fixed mixture weight of 0.95, HTB achieves 32.8% AntFall success versus 9.6% for matched local HIRO over five paired seeds at 10M environment steps. Ablations identify contributions from recursive continuation and mixed supervision; fixed-policy tests show more accurate predictions for actions whose returns were excluded from fitting.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Liu, B., & Jing, Y. (2026). Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning. https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning

MLA 9

Liu, Bingyun, and Yuheng Jing. "Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning." https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning.

شيكاغو (المؤلف–التاريخ)

Liu, Bingyun, and Yuheng Jing. 2026. "Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning." https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning.

هارفارد

Liu, B. and Jing, Y. (2026) 'Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning', Available at: https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning.

فانكوفر

Liu B, Jing Y. Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning. https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning

IEEE

B. Liu, and Y. Jing, "Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning," https://omanscience.com/ar/articles/hierarchical-time-aware-bootstrapping-for-off-policy-subgoal-value-learning.