Abstract
Recent advances in robot learning have enabled generalist control policies capable of completing a wide range of tasks. However, their performance degrades when deployed in unseen environments, making it critical to detect failures and teach recovery behaviors. Existing runtime monitoring methods often require task- and policy-specific training or hyperparameter tuning, limiting cross-task deployment and introducing additional overhead during iterative policy updates. We present Reward-DAgger, a robot-gated interactive imitation learning framework that uses dense progress signals from a general-purpose reward model to determine when human intervention is needed. Our approach is agnostic to the underlying policy architecture, requires no access to policy internals, and can be applied across tasks without retuning the gating mechanism. Our results show that Reward-DAgger achieves a better failure-detection accuracy-latency tradeoff than existing runtime monitoring baselines. Across eight simulated and real-world tasks, Reward-DAgger consistently improves the downstream policy's autonomous success rate throughout interactive learning and achieves strong return on human effort, outperforming the baselines in most settings. Importantly, the same gating configuration is used across tasks without task-specific hyperparameter tuning, demonstrating transfer across tasks, environments, and policy architectures. Code and videos are available at https://liralab.usc.edu/reward-dagger.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Li, R., Korkmaz, Y., & Bıyık, E. (2026). Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models. https://omanscience.com/en/articles/reward-dagger-robot-gated-interactive-imitation-learning-with-general-purpose-progress-based-reward-models
MLA 9
Li, Ryan, et al. "Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models." https://omanscience.com/en/articles/reward-dagger-robot-gated-interactive-imitation-learning-with-general-purpose-progress-based-reward-models.
Chicago (author–date)
Li, Ryan, Yigit Korkmaz, and Erdem Bıyık. 2026. "Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models." https://omanscience.com/en/articles/reward-dagger-robot-gated-interactive-imitation-learning-with-general-purpose-progress-based-reward-models.
Harvard
Li, R., Korkmaz, Y. and Bıyık, E. (2026) 'Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models', Available at: https://omanscience.com/en/articles/reward-dagger-robot-gated-interactive-imitation-learning-with-general-purpose-progress-based-reward-models.
Vancouver
Li R, Korkmaz Y, Bıyık E. Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models. https://omanscience.com/en/articles/reward-dagger-robot-gated-interactive-imitation-learning-with-general-purpose-progress-based-reward-models
IEEE
R. Li, Y. Korkmaz, and E. Bıyık, "Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models," https://omanscience.com/en/articles/reward-dagger-robot-gated-interactive-imitation-learning-with-general-purpose-progress-based-reward-models.