Abstract

Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned sequential stopping. At each causal checkpoint, CASE combines an auxiliary multiple-choice assessment of accumulated evidence with the host agent's execution state. From complete native trajectories, we construct a cost-aware target that compares answering now with stopping later along the same search path, accounting jointly for answer correctness and the full cost of continued reasoning. A lightweight Ridge regressor learns this decision gap and produces STOP/CONTINUE decisions. We evaluate three vision-language models with VideoSeek and AVP. On Video-MME, end-to-end accuracy changes by +0.67 percentage points on average while CASE reduces model-token use by 53.63%. The same frozen policies then transfer zero-shot to LongVideoBench and MLVU, with end-to-end accuracy changes of +3.38 and +4.58 points while saving 58.78% and 51.28% of model tokens, respectively. Across all agent-model-benchmark combinations, CASE attains the highest accuracy-efficiency Pareto-frontier coverage among the compared stopping methods (83.3%) at the selected operating points. Online execution preserves this favorable accuracy-efficiency trade-off and additionally reduces measured runtime by 54.1% on average. CASE provides a plug-in termination framework for long-video reasoning agents, enabling them to decide when further evidence acquisition is no longer worthwhile.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Du, Y., Liu, C., Liu, Z., Zheng, F., Wang, Z., Nie, J., & Huang, S. (2026). CASE: Cost-Aware Stopping for Efficient Long-Video Agents. https://omanscience.com/en/articles/case-cost-aware-stopping-for-efficient-long-video-agents

MLA 9

Du, Yiming, et al. "CASE: Cost-Aware Stopping for Efficient Long-Video Agents." https://omanscience.com/en/articles/case-cost-aware-stopping-for-efficient-long-video-agents.

Chicago (author–date)

Du, Yiming, Chenghao Liu, Zhiyuan Liu, Fangxing Zheng, Zhao Wang, Junnan Nie, and Songfang Huang. 2026. "CASE: Cost-Aware Stopping for Efficient Long-Video Agents." https://omanscience.com/en/articles/case-cost-aware-stopping-for-efficient-long-video-agents.

Harvard

Du, Y., Liu, C., Liu, Z., Zheng, F., Wang, Z., Nie, J. and Huang, S. (2026) 'CASE: Cost-Aware Stopping for Efficient Long-Video Agents', Available at: https://omanscience.com/en/articles/case-cost-aware-stopping-for-efficient-long-video-agents.

Vancouver

Du Y, Liu C, Liu Z, Zheng F, Wang Z, Nie J, et al. CASE: Cost-Aware Stopping for Efficient Long-Video Agents. https://omanscience.com/en/articles/case-cost-aware-stopping-for-efficient-long-video-agents

IEEE

Y. Du, C. Liu, Z. Liu, F. Zheng, Z. Wang, J. Nie, and S. Huang, "CASE: Cost-Aware Stopping for Efficient Long-Video Agents," https://omanscience.com/en/articles/case-cost-aware-stopping-for-efficient-long-video-agents.