الملخص

Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which first trains VLA policies to remain functional under missing visual inputs. At inference time, MINT selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable. Experiments on $π_{0.5}$ and GR00T N1.5 show that MINT significantly improves task success under camera loss over the original models. Experiments on AgiBot G2 further demonstrate the real-robot deployment under camera loss. The benchmark is available at https://minglejiang.github.io/Mail-Bench/

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Jiang, M., Xu, R., Wang, Y., & Xu, C. (2026). Learning to Act under Visual Interruptions with Vision-Language-Action Models. https://omanscience.com/ar/articles/learning-to-act-under-visual-interruptions-with-vision-language-action-models

MLA 9

Jiang, Mingle, et al. "Learning to Act under Visual Interruptions with Vision-Language-Action Models." https://omanscience.com/ar/articles/learning-to-act-under-visual-interruptions-with-vision-language-action-models.

شيكاغو (المؤلف–التاريخ)

Jiang, Mingle, Rui Xu, Yunke Wang, and Chang Xu. 2026. "Learning to Act under Visual Interruptions with Vision-Language-Action Models." https://omanscience.com/ar/articles/learning-to-act-under-visual-interruptions-with-vision-language-action-models.

هارفارد

Jiang, M., Xu, R., Wang, Y. and Xu, C. (2026) 'Learning to Act under Visual Interruptions with Vision-Language-Action Models', Available at: https://omanscience.com/ar/articles/learning-to-act-under-visual-interruptions-with-vision-language-action-models.

فانكوفر

Jiang M, Xu R, Wang Y, Xu C. Learning to Act under Visual Interruptions with Vision-Language-Action Models. https://omanscience.com/ar/articles/learning-to-act-under-visual-interruptions-with-vision-language-action-models

IEEE

M. Jiang, R. Xu, Y. Wang, and C. Xu, "Learning to Act under Visual Interruptions with Vision-Language-Action Models," https://omanscience.com/ar/articles/learning-to-act-under-visual-interruptions-with-vision-language-action-models.