Preprint Open access
EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation
Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the polic …