Abstract
Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design. DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone. DIVA preserves the full visual token sequence and requires no external grounding supervision. On LIBERO, DIVA improves OpenVLA-OFT from 96.6% to 98.0% average success and raises its zero-shot LIBERO-Plus score from 69.6 to 72.6. Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Feng, K., Sun, G., Wang, Z., He, Y., Shen, Z., & Li, A. (2026). DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies. https://omanscience.com/en/articles/diva-dual-space-intent-aware-visual-attenuation-for-vision-language-action-policies
MLA 9
Feng, Kaixi, et al. "DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies." https://omanscience.com/en/articles/diva-dual-space-intent-aware-visual-attenuation-for-vision-language-action-policies.
Chicago (author–date)
Feng, Kaixi, Guoheng Sun, Ziyao Wang, Yexiao He, Zheyu Shen, and Ang Li. 2026. "DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies." https://omanscience.com/en/articles/diva-dual-space-intent-aware-visual-attenuation-for-vision-language-action-policies.
Harvard
Feng, K., Sun, G., Wang, Z., He, Y., Shen, Z. and Li, A. (2026) 'DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies', Available at: https://omanscience.com/en/articles/diva-dual-space-intent-aware-visual-attenuation-for-vision-language-action-policies.
Vancouver
Feng K, Sun G, Wang Z, He Y, Shen Z, Li A. DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies. https://omanscience.com/en/articles/diva-dual-space-intent-aware-visual-attenuation-for-vision-language-action-policies
IEEE
K. Feng, G. Sun, Z. Wang, Y. He, Z. Shen, and A. Li, "DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies," https://omanscience.com/en/articles/diva-dual-space-intent-aware-visual-attenuation-for-vision-language-action-policies.