الملخص
Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned sequential stopping. At each causal checkpoint, CASE combines an auxiliary multiple-choice assessment of accumulated evidence with the host agent's execution state. From complete native trajectories, we construct a cost-aware target that compares answering now with stopping later along the same search path, accounting jointly for answer correctness and the full cost of continued reasoning. A lightweight Ridge regressor learns this decision gap and produces STOP/CONTINUE decisions. We evaluate three vision-language models with VideoSeek and AVP. On Video-MME, end-to-end accuracy changes by +0.67 percentage points on average while CASE reduces model-token use by 53.63%. The same frozen policies then transfer zero-shot to LongVideoBench and MLVU, with end-to-end accuracy changes of +3.38 and +4.58 points while saving 58.78% and 51.28% of model tokens, respectively. Across all agent-model-benchmark combinations, CASE attains the highest accuracy-efficiency Pareto-frontier coverage among the compared stopping methods (83.3%) at the selected operating points. Online execution preserves this favorable accuracy-efficiency trade-off and additionally reduces measured runtime by 54.1% on average. CASE provides a plug-in termination framework for long-video reasoning agents, enabling them to decide when further evidence acquisition is no longer worthwhile.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Du, Y., Liu, C., Liu, Z., Zheng, F., Wang, Z., Nie, J., & Huang, S. (2026). CASE: Cost-Aware Stopping for Efficient Long-Video Agents. https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents
MLA 9
Du, Yiming, et al. "CASE: Cost-Aware Stopping for Efficient Long-Video Agents." https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents.
شيكاغو (المؤلف–التاريخ)
Du, Yiming, Chenghao Liu, Zhiyuan Liu, Fangxing Zheng, Zhao Wang, Junnan Nie, and Songfang Huang. 2026. "CASE: Cost-Aware Stopping for Efficient Long-Video Agents." https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents.
هارفارد
Du, Y., Liu, C., Liu, Z., Zheng, F., Wang, Z., Nie, J. and Huang, S. (2026) 'CASE: Cost-Aware Stopping for Efficient Long-Video Agents', Available at: https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents.
فانكوفر
Du Y, Liu C, Liu Z, Zheng F, Wang Z, Nie J, et al. CASE: Cost-Aware Stopping for Efficient Long-Video Agents. https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents
IEEE
Y. Du, C. Liu, Z. Liu, F. Zheng, Z. Wang, J. Nie, and S. Huang, "CASE: Cost-Aware Stopping for Efficient Long-Video Agents," https://omanscience.com/ar/articles/case-cost-aware-stopping-for-efficient-long-video-agents.