Abstract

Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lin, J., Lao, Z., Cui, J., Zhao, L., Zhao, P., Zhang, D., Akbari, A., Qi, Y., Jiang, X., Wang, Y., Yu, H., & Peng, L. (2026). LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception. https://omanscience.com/en/articles/leap-learned-block-wise-evidence-retrieval-for-long-audio-video-perception

MLA 9

Lin, Juyi, et al. "LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception." https://omanscience.com/en/articles/leap-learned-block-wise-evidence-retrieval-for-long-audio-video-perception.

Chicago (author–date)

Lin, Juyi, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang, Heather Yu, and Liang Peng. 2026. "LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception." https://omanscience.com/en/articles/leap-learned-block-wise-evidence-retrieval-for-long-audio-video-perception.

Harvard

Lin, J., Lao, Z., Cui, J., Zhao, L., Zhao, P., Zhang, D., Akbari, A., Qi, Y., Jiang, X., Wang, Y., Yu, H. and Peng, L. (2026) 'LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception', Available at: https://omanscience.com/en/articles/leap-learned-block-wise-evidence-retrieval-for-long-audio-video-perception.

Vancouver

Lin J, Lao Z, Cui J, Zhao L, Zhao P, Zhang D, et al. LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception. https://omanscience.com/en/articles/leap-learned-block-wise-evidence-retrieval-for-long-audio-video-perception

IEEE

J. Lin, Z. Lao, J. Cui, L. Zhao, P. Zhao, D. Zhang, A. Akbari, Y. Qi, X. Jiang, Y. Wang, H. Yu, and L. Peng, "LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception," https://omanscience.com/en/articles/leap-learned-block-wise-evidence-retrieval-for-long-audio-video-perception.