Abstract

Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, W., Shi, Q., & Zhou, Y. (2026). StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies. https://omanscience.com/en/articles/statemem-single-state-residual-memory-with-adaptive-inference-for-vision-language-action-policies

MLA 9

Li, Wenzhuo, et al. "StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies." https://omanscience.com/en/articles/statemem-single-state-residual-memory-with-adaptive-inference-for-vision-language-action-policies.

Chicago (author–date)

Li, Wenzhuo, Qiongfeng Shi, and Yi Zhou. 2026. "StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies." https://omanscience.com/en/articles/statemem-single-state-residual-memory-with-adaptive-inference-for-vision-language-action-policies.

Harvard

Li, W., Shi, Q. and Zhou, Y. (2026) 'StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies', Available at: https://omanscience.com/en/articles/statemem-single-state-residual-memory-with-adaptive-inference-for-vision-language-action-policies.

Vancouver

Li W, Shi Q, Zhou Y. StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies. https://omanscience.com/en/articles/statemem-single-state-residual-memory-with-adaptive-inference-for-vision-language-action-policies

IEEE

W. Li, Q. Shi, and Y. Zhou, "StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies," https://omanscience.com/en/articles/statemem-single-state-residual-memory-with-adaptive-inference-for-vision-language-action-policies.