Abstract
Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Ben-Ami, D., Cohen, K., & Baskin, C. (2026). Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness. https://omanscience.com/en/articles/have-i-seen-enough-frozen-video-language-models-encode-evidence-readiness
MLA 9
Ben-Ami, Dan, et al. "Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness." https://omanscience.com/en/articles/have-i-seen-enough-frozen-video-language-models-encode-evidence-readiness.
Chicago (author–date)
Ben-Ami, Dan, Kobi Cohen, and Chaim Baskin. 2026. "Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness." https://omanscience.com/en/articles/have-i-seen-enough-frozen-video-language-models-encode-evidence-readiness.
Harvard
Ben-Ami, D., Cohen, K. and Baskin, C. (2026) 'Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness', Available at: https://omanscience.com/en/articles/have-i-seen-enough-frozen-video-language-models-encode-evidence-readiness.
Vancouver
Ben-Ami D, Cohen K, Baskin C. Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness. https://omanscience.com/en/articles/have-i-seen-enough-frozen-video-language-models-encode-evidence-readiness
IEEE
D. Ben-Ami, K. Cohen, and C. Baskin, "Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness," https://omanscience.com/en/articles/have-i-seen-enough-frozen-video-language-models-encode-evidence-readiness.