Abstract

Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address this, we propose \textbf{Optimus-R}, a memory-centric VLA framework that formulates robotic adaptation as explicit query-skill memory tuning. Optimus-R introduces: (i) An \textbf{Inline Memory Interface for skill extraction}. It inserts learnable memory tokens into the VLA prefix stream, allowing the backbone to derive control-aware query and skill representations within the native action-conditioning pathway. (ii) A \textbf{Query-Skill Memory Bank for skill learning}. It externalizes skills into query prototypes for deciding \emph{what} to retrieve and skill values for specifying \emph{how} to act, supporting skill reuse and expansion with limited parameter updates. (iii) A lightweight \textbf{Bridge-and-Adapt mechanism for skill updating}. It aligns target-domain queries and skills with the existing memory space through a lightweight adapter and residual memory updates. Experiments on in-domain adaptation, cross-domain adaptation, and lifelong learning show that Optimus-R enables data-efficient skill learning while mitigating catastrophic forgetting.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, Z., Shao, R., Hu, B., Zhang, H., Jiang, D., & Nie, L. (2026). Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model. https://omanscience.com/en/articles/inline-memory-meets-reusable-skills-memory-centric-framework-for-vision-language-action-model

MLA 9

Li, Zaijing, et al. "Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model." https://omanscience.com/en/articles/inline-memory-meets-reusable-skills-memory-centric-framework-for-vision-language-action-model.

Chicago (author–date)

Li, Zaijing, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, and Liqiang Nie. 2026. "Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model." https://omanscience.com/en/articles/inline-memory-meets-reusable-skills-memory-centric-framework-for-vision-language-action-model.

Harvard

Li, Z., Shao, R., Hu, B., Zhang, H., Jiang, D. and Nie, L. (2026) 'Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model', Available at: https://omanscience.com/en/articles/inline-memory-meets-reusable-skills-memory-centric-framework-for-vision-language-action-model.

Vancouver

Li Z, Shao R, Hu B, Zhang H, Jiang D, Nie L. Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model. https://omanscience.com/en/articles/inline-memory-meets-reusable-skills-memory-centric-framework-for-vision-language-action-model

IEEE

Z. Li, R. Shao, B. Hu, H. Zhang, D. Jiang, and L. Nie, "Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model," https://omanscience.com/en/articles/inline-memory-meets-reusable-skills-memory-centric-framework-for-vision-language-action-model.