Abstract
Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth. Our analysis reveals that RAW-to-RGB processing materially shapes both action prediction and manipulation success, with different ISP dimensions exerting substantially different effects. Guided by these findings, we introduce RawVLA, a streaming neural ISP that adaptively renders RAW observations for frozen VLA policies while concentrating its capacity on the imaging factors relevant to embodied behavior. We further present RawVLA-Bench, a RAW-domain manipulation benchmark to expose image processing as an explicit evaluation variable across clean and adverse acquisition conditions. Experiments on RawVLA-Bench show that RawVLA preserves performance under standard conditions while substantially improving robustness under degraded imaging, establishing adaptive RAW processing as an effective interface between physical cameras and embodied policies.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Liu, S., Zhou, H., Qian, L., Fang, Y., Hou, X., Zhou, Q., Gu, L., Sui, W., Yang, J., & Cui, Z. (2026). RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation. https://omanscience.com/en/articles/rawvla-embodied-neural-image-signal-processor-for-robotic-manipulation
MLA 9
Liu, Shuhong, et al. "RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation." https://omanscience.com/en/articles/rawvla-embodied-neural-image-signal-processor-for-robotic-manipulation.
Chicago (author–date)
Liu, Shuhong, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, and Ziteng Cui. 2026. "RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation." https://omanscience.com/en/articles/rawvla-embodied-neural-image-signal-processor-for-robotic-manipulation.
Harvard
Liu, S., Zhou, H., Qian, L., Fang, Y., Hou, X., Zhou, Q., Gu, L., Sui, W., Yang, J. and Cui, Z. (2026) 'RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation', Available at: https://omanscience.com/en/articles/rawvla-embodied-neural-image-signal-processor-for-robotic-manipulation.
Vancouver
Liu S, Zhou H, Qian L, Fang Y, Hou X, Zhou Q, et al. RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation. https://omanscience.com/en/articles/rawvla-embodied-neural-image-signal-processor-for-robotic-manipulation
IEEE
S. Liu, H. Zhou, L. Qian, Y. Fang, X. Hou, Q. Zhou, L. Gu, W. Sui, J. Yang, and Z. Cui, "RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation," https://omanscience.com/en/articles/rawvla-embodied-neural-image-signal-processor-for-robotic-manipulation.