Abstract

Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between real-time perception and long-term memory. Retrieving historical information provides a natural remedy, yet historical recall is not uniformly beneficial: unnecessary history may introduce irrelevant context into current reasoning and interfere with native real-time perception. Effective streaming memory should therefore address not only what to remember, but also when and how to access it. To this end, we introduce FlashBack, a training-free framework for selective, multi-level memory in streaming vision-language models. Before retrieving history, FlashBack draws on the semantic understanding of the frozen streaming VLM to infer whether a query calls for historical evidence. This assessment determines whether inference remains on the Native trajectory or invokes an isolated Recall trajectory. The Recall trajectory combines recent context with retrieved long-term memory through a query-local Side-KV pathway, preserving local temporal continuity without modifying the persistent Native state. We instantiate FlashBack on StreamingVLM and Mage-VL-4B and evaluate it on OVO-Bench and StreamingBench. The results show improvements on several long-horizon and memory-dependent tasks while largely preserving real-time perception, with performance competitive with strong training-based streaming methods despite requiring no additional training. Our code will be announced later.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chen, Y., Yu, M., Wang, R. Q., Wang, B., Cao, X., Tang, C., Chen, J., & Gu, J. (2026). FlashBack: Knowing When to Remember in Streaming Vision-Language Models. https://omanscience.com/en/articles/flashback-knowing-when-to-remember-in-streaming-vision-language-models

MLA 9

Chen, Yi, et al. "FlashBack: Knowing When to Remember in Streaming Vision-Language Models." https://omanscience.com/en/articles/flashback-knowing-when-to-remember-in-streaming-vision-language-models.

Chicago (author–date)

Chen, Yi, MingMing Yu, Rui-Qi Wang, Boran Wang, Xiaohang Cao, Chu Tang, Jingmin Chen, and Jie Gu. 2026. "FlashBack: Knowing When to Remember in Streaming Vision-Language Models." https://omanscience.com/en/articles/flashback-knowing-when-to-remember-in-streaming-vision-language-models.

Harvard

Chen, Y., Yu, M., Wang, R. Q., Wang, B., Cao, X., Tang, C., Chen, J. and Gu, J. (2026) 'FlashBack: Knowing When to Remember in Streaming Vision-Language Models', Available at: https://omanscience.com/en/articles/flashback-knowing-when-to-remember-in-streaming-vision-language-models.

Vancouver

Chen Y, Yu M, Wang RQ, Wang B, Cao X, Tang C, et al. FlashBack: Knowing When to Remember in Streaming Vision-Language Models. https://omanscience.com/en/articles/flashback-knowing-when-to-remember-in-streaming-vision-language-models

IEEE

Y. Chen, M. Yu, R. Q. Wang, B. Wang, X. Cao, C. Tang, J. Chen, and J. Gu, "FlashBack: Knowing When to Remember in Streaming Vision-Language Models," https://omanscience.com/en/articles/flashback-knowing-when-to-remember-in-streaming-vision-language-models.