الملخص

Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We observe, however, that an explicit hierarchy does not by itself determine how stable the resulting temporal abstraction is: the learned boundary policy may replace the subgoal almost every turn, making it effectively transient, or retain a subgoal after it has stopped being appropriate. We call this temporal abstraction instability. We propose Stable Temporal Abstraction via Constrained Optimization (STAC), a constrained boundary-policy optimization method that represents premature replanning and stale persistence as constraint costs. STAC applies the resulting Lagrangian costs only to the sampled boundary decision, leaving the underlying algorithm's rewards, critic targets, subgoal advantages, and primitive-action advantages unchanged. Across two backbones and two benchmarks, STAC improves success over a strong hierarchical baseline by $8.1$ and $7.9$ points on ALFWorld and WebShop with Qwen3-0.6B, and by $23.5$ and $15.8$ points with Llama-3.2-1B-Instruct.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Hamidi, S. M., Cheng, Y., Xu, Y., Zhou, Z., & Geramifard, A. (2026). Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents. https://omanscience.com/ar/articles/hierarchical-reinforcement-learning-with-stable-temporal-abstraction-for-language-model-agents

MLA 9

Hamidi, Shayan Mohajer, et al. "Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents." https://omanscience.com/ar/articles/hierarchical-reinforcement-learning-with-stable-temporal-abstraction-for-language-model-agents.

شيكاغو (المؤلف–التاريخ)

Hamidi, Shayan Mohajer, Yize Cheng, Yuanda Xu, Zhengze Zhou, and Alborz Geramifard. 2026. "Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents." https://omanscience.com/ar/articles/hierarchical-reinforcement-learning-with-stable-temporal-abstraction-for-language-model-agents.

هارفارد

Hamidi, S. M., Cheng, Y., Xu, Y., Zhou, Z. and Geramifard, A. (2026) 'Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents', Available at: https://omanscience.com/ar/articles/hierarchical-reinforcement-learning-with-stable-temporal-abstraction-for-language-model-agents.

فانكوفر

Hamidi SM, Cheng Y, Xu Y, Zhou Z, Geramifard A. Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents. https://omanscience.com/ar/articles/hierarchical-reinforcement-learning-with-stable-temporal-abstraction-for-language-model-agents

IEEE

S. M. Hamidi, Y. Cheng, Y. Xu, Z. Zhou, and A. Geramifard, "Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents," https://omanscience.com/ar/articles/hierarchical-reinforcement-learning-with-stable-temporal-abstraction-for-language-model-agents.