Abstract

Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounter change. Yet current agents struggle to maintain such memory reliably. Our analysis traces this limitation to four key deficiencies: weak fine-grained visual memory, unreliable dynamic world-state tracking, failing to record world state revealed by interaction outcomes, and limited generalization from prior experience. However, existing benchmarks do not directly assess these memory capabilities during long-horizon embodied interaction. To address this gap, we introduce EmbodiedMemory-Bench (EMem-Bench), comprising 2,554 interactive episodes across four task families. EMem-Bench requires agents to build and update memory from interaction history, then use it to complete a later task by acting in the environment. We further present Embodied-Memorizer (EMem), an external memory system that organizes embodied experience into spatial, event, and scene memories. We also train EMem-8B, an 8B policy that manages and uses these memories. We evaluate a diverse range of open-source and proprietary MLLMs and representative multimodal memory systems. Results show that current models remain weak and uneven across the four challenges. Under matched backbones, EMem achieves the best overall performance among the evaluated memory systems and improves both open-source and proprietary models, while EMem-8B further improves over its backbone. Project page: https://zju-omniai.github.io/EmbodiedMemoryBench/

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Liang, L., Zhong, X., Pan, M., Zhou, X., Liu, X., Li, Q., Li, P., Chen, J., Zhang, X., & Zhang, W. (2026). EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks. https://omanscience.com/en/articles/embodiedmemory-bench-benchmarking-embodied-memory-for-long-horizon-embodied-tasks

MLA 9

Liang, Lizhou, et al. "EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks." https://omanscience.com/en/articles/embodiedmemory-bench-benchmarking-embodied-memory-for-long-horizon-embodied-tasks.

Chicago (author–date)

Liang, Lizhou, Xinyu Zhong, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Qinfeng Li, Peng Li, Jintao Chen, Xuhong Zhang, and Wenqi Zhang. 2026. "EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks." https://omanscience.com/en/articles/embodiedmemory-bench-benchmarking-embodied-memory-for-long-horizon-embodied-tasks.

Harvard

Liang, L., Zhong, X., Pan, M., Zhou, X., Liu, X., Li, Q., Li, P., Chen, J., Zhang, X. and Zhang, W. (2026) 'EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks', Available at: https://omanscience.com/en/articles/embodiedmemory-bench-benchmarking-embodied-memory-for-long-horizon-embodied-tasks.

Vancouver

Liang L, Zhong X, Pan M, Zhou X, Liu X, Li Q, et al. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks. https://omanscience.com/en/articles/embodiedmemory-bench-benchmarking-embodied-memory-for-long-horizon-embodied-tasks

IEEE

L. Liang, X. Zhong, M. Pan, X. Zhou, X. Liu, Q. Li, P. Li, J. Chen, X. Zhang, and W. Zhang, "EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks," https://omanscience.com/en/articles/embodiedmemory-bench-benchmarking-embodied-memory-for-long-horizon-embodied-tasks.