Abstract

Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.

Keywords

Subject

Publication details

DOI
10.1145/3767308.3835688
Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, Y., Fu, Y., Dai, Y., Gong, J., Qian, T., & Wang, X. (2026). MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding. https://doi.org/10.1145/3767308.3835688

MLA 9

Li, Yinying, et al. "MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding." https://doi.org/10.1145/3767308.3835688.

Chicago (author–date)

Li, Yinying, Yuqian Fu, Yulin Dai, Jingyu Gong, Tianwen Qian, and Xiaoling Wang. 2026. "MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding." https://doi.org/10.1145/3767308.3835688.

Harvard

Li, Y., Fu, Y., Dai, Y., Gong, J., Qian, T. and Wang, X. (2026) 'MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding', doi:10.1145/3767308.3835688.

Vancouver

Li Y, Fu Y, Dai Y, Gong J, Qian T, Wang X. MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding. doi:10.1145/3767308.3835688

IEEE

Y. Li, Y. Fu, Y. Dai, J. Gong, T. Qian, and X. Wang, "MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding," doi: 10.1145/3767308.3835688.