الملخص

Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro Recall@0.5 from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Chen, J., Cui, B., Jia, R., Wang, Z., Wo, T., Sun, P., Huang, L., Xue, H., Yang, Y., & Hong, H. (2026). Grounding with Confidence: Controllable Generative Video Temporal Grounding. https://omanscience.com/ar/articles/grounding-with-confidence-controllable-generative-video-temporal-grounding

MLA 9

Chen, Jinhao, et al. "Grounding with Confidence: Controllable Generative Video Temporal Grounding." https://omanscience.com/ar/articles/grounding-with-confidence-controllable-generative-video-temporal-grounding.

شيكاغو (المؤلف–التاريخ)

Chen, Jinhao, Benlei Cui, Ruijian Jia, Ziheng Wang, Tianyu Wo, Pengfei Sun, Longtao Huang, Hui Xue, Yitong Yang, and Haiwen Hong. 2026. "Grounding with Confidence: Controllable Generative Video Temporal Grounding." https://omanscience.com/ar/articles/grounding-with-confidence-controllable-generative-video-temporal-grounding.

هارفارد

Chen, J., Cui, B., Jia, R., Wang, Z., Wo, T., Sun, P., Huang, L., Xue, H., Yang, Y. and Hong, H. (2026) 'Grounding with Confidence: Controllable Generative Video Temporal Grounding', Available at: https://omanscience.com/ar/articles/grounding-with-confidence-controllable-generative-video-temporal-grounding.

فانكوفر

Chen J, Cui B, Jia R, Wang Z, Wo T, Sun P, et al. Grounding with Confidence: Controllable Generative Video Temporal Grounding. https://omanscience.com/ar/articles/grounding-with-confidence-controllable-generative-video-temporal-grounding

IEEE

J. Chen, B. Cui, R. Jia, Z. Wang, T. Wo, P. Sun, L. Huang, H. Xue, Y. Yang, and H. Hong, "Grounding with Confidence: Controllable Generative Video Temporal Grounding," https://omanscience.com/ar/articles/grounding-with-confidence-controllable-generative-video-temporal-grounding.