Abstract

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by combining semantic knowledge from pretrained vision-language models with expressive action-generation policies. Diffusion-based action generators are particularly effective for modeling temporally coherent action chunks, but these chunks are typically executed open-loop after inference. This limits responsiveness when objects move, contacts change, or the scene evolves during execution. We propose VLA-Feedback, a two-timescale architecture that combines low-frequency diffusion planning with high-frequency visual feedback. Rather than fully denoising an action chunk before execution, VLA-Feedback retains its final denoising step as a lightweight feedback interface, allowing each action to be corrected using the latest observation before it is executed. This design preserves the expressiveness of the diffusion planner while enabling real-time action correction without rerunning the full vision-language diffusion model. VLA-Feedback matched GR00T on static LIBERO tasks while improving average success on dynamic simulation tasks from 27.5% to 85.0%. On real-robot tasks, it improved average success from 51% to 73%. Additional materials can be found on our project page: https://vla-feedback.github.io.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ji, Y., Zhou, X., Sentis, L., & Seo, M. (2026). Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs. https://omanscience.com/en/articles/catch-me-if-you-can-real-time-feedback-denoising-for-responsive-vlas

MLA 9

Ji, Yiheng, et al. "Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs." https://omanscience.com/en/articles/catch-me-if-you-can-real-time-feedback-denoising-for-responsive-vlas.

Chicago (author–date)

Ji, Yiheng, Xingru Zhou, Luis Sentis, and Mingyo Seo. 2026. "Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs." https://omanscience.com/en/articles/catch-me-if-you-can-real-time-feedback-denoising-for-responsive-vlas.

Harvard

Ji, Y., Zhou, X., Sentis, L. and Seo, M. (2026) 'Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs', Available at: https://omanscience.com/en/articles/catch-me-if-you-can-real-time-feedback-denoising-for-responsive-vlas.

Vancouver

Ji Y, Zhou X, Sentis L, Seo M. Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs. https://omanscience.com/en/articles/catch-me-if-you-can-real-time-feedback-denoising-for-responsive-vlas

IEEE

Y. Ji, X. Zhou, L. Sentis, and M. Seo, "Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs," https://omanscience.com/en/articles/catch-me-if-you-can-real-time-feedback-denoising-for-responsive-vlas.