الملخص

Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Iavorskaia, M., Dietz, C., Albrecht, S., & Khadiv, M. (2026). Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation. https://omanscience.com/ar/articles/res-hil-human-guided-residual-reinforcement-learning-for-sample-efficient-dexterous-manipulation

MLA 9

Iavorskaia, Mariia, et al. "Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation." https://omanscience.com/ar/articles/res-hil-human-guided-residual-reinforcement-learning-for-sample-efficient-dexterous-manipulation.

شيكاغو (المؤلف–التاريخ)

Iavorskaia, Mariia, Christian Dietz, Sebastian Albrecht, and Majid Khadiv. 2026. "Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation." https://omanscience.com/ar/articles/res-hil-human-guided-residual-reinforcement-learning-for-sample-efficient-dexterous-manipulation.

هارفارد

Iavorskaia, M., Dietz, C., Albrecht, S. and Khadiv, M. (2026) 'Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation', Available at: https://omanscience.com/ar/articles/res-hil-human-guided-residual-reinforcement-learning-for-sample-efficient-dexterous-manipulation.

فانكوفر

Iavorskaia M, Dietz C, Albrecht S, Khadiv M. Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation. https://omanscience.com/ar/articles/res-hil-human-guided-residual-reinforcement-learning-for-sample-efficient-dexterous-manipulation

IEEE

M. Iavorskaia, C. Dietz, S. Albrecht, and M. Khadiv, "Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation," https://omanscience.com/ar/articles/res-hil-human-guided-residual-reinforcement-learning-for-sample-efficient-dexterous-manipulation.