الملخص
Learning from human demonstrations is a reliable way to teach robots new tasks, but the gains from each additional demonstration shrink as the policy improves. Continued improvement can instead come from supervised deployment, where an operator places the objects and intervenes when the policy fails. We ask how to maximize improvement from a fixed budget of supervised episodes on high-precision manipulation tasks with wide ranges of object placements. We observe that failures can concentrate in a small subset of initial states, so uniform collection spends much of the operator's time on states the policy already handles. Mulligan makes the initial-state distribution a decision, starting each round's episodes at observed failures and untried states. To further improve data efficiency, we augment interactive imitation learning with a value function trained on all data, including failures that imitation discards. Across three real-world tasks evaluated on 2,550 held-out, blinded episodes and two simulated tasks, Mulligan outperforms uniform initial-state sampling at matched collection budgets, and combined with value-based action selection, HiL-IDQL+Mulligan, improves final real-task success by 10-34 percentage points. With operator interventions, the human-robot team completes 98% of collection episodes, remaining productive while the policy learns. Videos, code, and data are available at https://mulligan.page/.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Ankile, L., Dong, P., Bhowmik, R., Muppidi, A., Yuan, D. D., Song, S., & Finn, C. (2026). Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning. https://omanscience.com/ar/articles/mulligan-performance-guided-data-collection-for-efficient-on-robot-learning
MLA 9
Ankile, Lars, et al. "Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning." https://omanscience.com/ar/articles/mulligan-performance-guided-data-collection-for-efficient-on-robot-learning.
شيكاغو (المؤلف–التاريخ)
Ankile, Lars, Perry Dong, Rohan Bhowmik, Aneesh Muppidi, David D. Yuan, Shuran Song, and Chelsea Finn. 2026. "Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning." https://omanscience.com/ar/articles/mulligan-performance-guided-data-collection-for-efficient-on-robot-learning.
هارفارد
Ankile, L., Dong, P., Bhowmik, R., Muppidi, A., Yuan, D. D., Song, S. and Finn, C. (2026) 'Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning', Available at: https://omanscience.com/ar/articles/mulligan-performance-guided-data-collection-for-efficient-on-robot-learning.
فانكوفر
Ankile L, Dong P, Bhowmik R, Muppidi A, Yuan DD, Song S, et al. Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning. https://omanscience.com/ar/articles/mulligan-performance-guided-data-collection-for-efficient-on-robot-learning
IEEE
L. Ankile, P. Dong, R. Bhowmik, A. Muppidi, D. D. Yuan, S. Song, and C. Finn, "Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning," https://omanscience.com/ar/articles/mulligan-performance-guided-data-collection-for-efficient-on-robot-learning.