الملخص

Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Hu, X., Yu, J., Zhang, L., Zhuge, Y., & Lu, H. (2026). When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs. https://omanscience.com/ar/articles/when-to-retrieve-when-to-stay-uncertainty-aware-temporal-evidence-allocation-for-streaming-video-llms

MLA 9

Hu, Xiang, et al. "When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs." https://omanscience.com/ar/articles/when-to-retrieve-when-to-stay-uncertainty-aware-temporal-evidence-allocation-for-streaming-video-llms.

شيكاغو (المؤلف–التاريخ)

Hu, Xiang, Jiazuo Yu, Lu Zhang, Yunzhi Zhuge, and Huchuan Lu. 2026. "When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs." https://omanscience.com/ar/articles/when-to-retrieve-when-to-stay-uncertainty-aware-temporal-evidence-allocation-for-streaming-video-llms.

هارفارد

Hu, X., Yu, J., Zhang, L., Zhuge, Y. and Lu, H. (2026) 'When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs', Available at: https://omanscience.com/ar/articles/when-to-retrieve-when-to-stay-uncertainty-aware-temporal-evidence-allocation-for-streaming-video-llms.

فانكوفر

Hu X, Yu J, Zhang L, Zhuge Y, Lu H. When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs. https://omanscience.com/ar/articles/when-to-retrieve-when-to-stay-uncertainty-aware-temporal-evidence-allocation-for-streaming-video-llms

IEEE

X. Hu, J. Yu, L. Zhang, Y. Zhuge, and H. Lu, "When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs," https://omanscience.com/ar/articles/when-to-retrieve-when-to-stay-uncertainty-aware-temporal-evidence-allocation-for-streaming-video-llms.