الملخص

Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned sequential stopping. At each causal checkpoint, CASE combines an auxiliary multiple-choice assessment of accumulated evidence with the host agent's execution state. From complete native trajectories, we construct a cost-aware target that compares answering now with stopping later along the same search path, accounting jointly for answer correctness and the full cost of continued reasoning. A lightweight Ridge regressor learns this decision gap and produces STOP/CONTINUE decisions. We evaluate three vision-language models with VideoSeek and AVP. On Video-MME, end-to-end accuracy changes by +0.67 percentage points on average while CASE reduces model-token use by 53.63%. The same frozen policies then transfer zero-shot to LongVideoBench and MLVU, with end-to-end accuracy changes of +3.38 and +4.58 points while saving 58.78% and 51.28% of model tokens, respectively. Across all agent-model-benchmark combinations, CASE attains the highest accuracy-efficiency Pareto-frontier coverage among the compared stopping methods (83.3%) at the selected operating points. Online execution preserves this favorable accuracy-efficiency trade-off and additionally reduces measured runtime by 54.1% on average. CASE provides a plug-in termination framework for long-video reasoning agents, enabling them to decide when further evidence acquisition is no longer worthwhile.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Du, Y., Liu, C., Liu, Z., Zheng, F., Wang, Z., Nie, J., & Huang, S. (2026). CASE: Cost-Aware Stopping for Efficient Long-Video Agents. https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents

MLA 9

Du, Yiming, et al. "CASE: Cost-Aware Stopping for Efficient Long-Video Agents." https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents.

شيكاغو (المؤلف–التاريخ)

Du, Yiming, Chenghao Liu, Zhiyuan Liu, Fangxing Zheng, Zhao Wang, Junnan Nie, and Songfang Huang. 2026. "CASE: Cost-Aware Stopping for Efficient Long-Video Agents." https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents.

هارفارد

Du, Y., Liu, C., Liu, Z., Zheng, F., Wang, Z., Nie, J. and Huang, S. (2026) 'CASE: Cost-Aware Stopping for Efficient Long-Video Agents', Available at: https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents.

فانكوفر

Du Y, Liu C, Liu Z, Zheng F, Wang Z, Nie J, et al. CASE: Cost-Aware Stopping for Efficient Long-Video Agents. https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents

IEEE

Y. Du, C. Liu, Z. Liu, F. Zheng, Z. Wang, J. Nie, and S. Huang, "CASE: Cost-Aware Stopping for Efficient Long-Video Agents," https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents.