Abstract

Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Huang, Y., Zhang, Y., Wang, Y., Lu, J., Dong, Q., Wang, H., Zeng, H., Zhang, M., & Fu, Y. (2026). RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding. https://omanscience.com/en/articles/racer-reflective-agent-coupling-query-interpretation-and-tool-based-retrieval-for-frame-selection-in-long-video-understanding

MLA 9

Huang, Yiyang, et al. "RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding." https://omanscience.com/en/articles/racer-reflective-agent-coupling-query-interpretation-and-tool-based-retrieval-for-frame-selection-in-long-video-understanding.

Chicago (author–date)

Huang, Yiyang, Yitian Zhang, Yizhou Wang, Jianglin Lu, Qihua Dong, Hailing Wang, Huimin Zeng, Mingyuan Zhang, and Yun Fu. 2026. "RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding." https://omanscience.com/en/articles/racer-reflective-agent-coupling-query-interpretation-and-tool-based-retrieval-for-frame-selection-in-long-video-understanding.

Harvard

Huang, Y., Zhang, Y., Wang, Y., Lu, J., Dong, Q., Wang, H., Zeng, H., Zhang, M. and Fu, Y. (2026) 'RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding', Available at: https://omanscience.com/en/articles/racer-reflective-agent-coupling-query-interpretation-and-tool-based-retrieval-for-frame-selection-in-long-video-understanding.

Vancouver

Huang Y, Zhang Y, Wang Y, Lu J, Dong Q, Wang H, et al. RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding. https://omanscience.com/en/articles/racer-reflective-agent-coupling-query-interpretation-and-tool-based-retrieval-for-frame-selection-in-long-video-understanding

IEEE

Y. Huang, Y. Zhang, Y. Wang, J. Lu, Q. Dong, H. Wang, H. Zeng, M. Zhang, and Y. Fu, "RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding," https://omanscience.com/en/articles/racer-reflective-agent-coupling-query-interpretation-and-tool-based-retrieval-for-frame-selection-in-long-video-understanding.