الملخص
Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in $τ$-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Guo, H., Liu, F., Wang, Y., Qi, Y., Xiong, H., Sun, F., & Du, M. (2026). Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation. https://omanscience.com/ar/articles/environmental-feedback-modeling-matters-rethinking-feedback-treatment-in-agentic-hindsight-self-distillation
MLA 9
Guo, Hangxi, et al. "Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation." https://omanscience.com/ar/articles/environmental-feedback-modeling-matters-rethinking-feedback-treatment-in-agentic-hindsight-self-distillation.
شيكاغو (المؤلف–التاريخ)
Guo, Hangxi, Fengyuan Liu, Yue Wang, Yuhua Qi, Haoyi Xiong, Fei Sun, and Mengnan Du. 2026. "Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation." https://omanscience.com/ar/articles/environmental-feedback-modeling-matters-rethinking-feedback-treatment-in-agentic-hindsight-self-distillation.
هارفارد
Guo, H., Liu, F., Wang, Y., Qi, Y., Xiong, H., Sun, F. and Du, M. (2026) 'Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation', Available at: https://omanscience.com/ar/articles/environmental-feedback-modeling-matters-rethinking-feedback-treatment-in-agentic-hindsight-self-distillation.
فانكوفر
Guo H, Liu F, Wang Y, Qi Y, Xiong H, Sun F, et al. Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation. https://omanscience.com/ar/articles/environmental-feedback-modeling-matters-rethinking-feedback-treatment-in-agentic-hindsight-self-distillation
IEEE
H. Guo, F. Liu, Y. Wang, Y. Qi, H. Xiong, F. Sun, and M. Du, "Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation," https://omanscience.com/ar/articles/environmental-feedback-modeling-matters-rethinking-feedback-treatment-in-agentic-hindsight-self-distillation.