الملخص

Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational overhead. Many existing methods prune visual tokens during or after vision transformer (ViT) encoding, leaving much of the encoding cost unaddressed. Some approaches prune patches before encoding but rely on learned auxiliary networks for patch selection, incurring additional training and inference overhead. To address these limitations, we propose FlashGaze, a training-free method that reduces spatiotemporal redundancy before ViT encoding without introducing auxiliary networks. FlashGaze uses pixel-space differences as a proxy for information loss and employs Quadtree Dynamic Programming to jointly optimize patch dropping, merging, and keeping under a fixed budget. Experiments on two MLLM backbones across multiple benchmarks demonstrate substantial efficiency gains while largely preserving accuracy. On Qwen3-VL-8B, FlashGaze retains 98% of the full-input baseline accuracy on LongVideoBench while achieving up to 5.4x and 17x speedups in ViT encoding and MLLM prefill, respectively, and reducing peak GPU memory usage by a factor of 1.8. These efficiency gains enable the model to process videos with more frames and higher resolutions on the same GPU hardware, unlocking video understanding at scales previously out of reach.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zhu, Z., Zhou, Y., Tan, L., Kang, J., Li, S., & Yang, X. (2026). FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding. https://omanscience.com/ar/articles/flashgaze-training-free-multi-scale-patch-pruning-for-efficient-video-understanding

MLA 9

Zhu, Ziye, et al. "FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding." https://omanscience.com/ar/articles/flashgaze-training-free-multi-scale-patch-pruning-for-efficient-video-understanding.

شيكاغو (المؤلف–التاريخ)

Zhu, Ziye, Yanghao Zhou, Lixing Tan, Jialiang Kang, Shuxuan Li, and Xiao Yang. 2026. "FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding." https://omanscience.com/ar/articles/flashgaze-training-free-multi-scale-patch-pruning-for-efficient-video-understanding.

هارفارد

Zhu, Z., Zhou, Y., Tan, L., Kang, J., Li, S. and Yang, X. (2026) 'FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding', Available at: https://omanscience.com/ar/articles/flashgaze-training-free-multi-scale-patch-pruning-for-efficient-video-understanding.

فانكوفر

Zhu Z, Zhou Y, Tan L, Kang J, Li S, Yang X. FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding. https://omanscience.com/ar/articles/flashgaze-training-free-multi-scale-patch-pruning-for-efficient-video-understanding

IEEE

Z. Zhu, Y. Zhou, L. Tan, J. Kang, S. Li, and X. Yang, "FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding," https://omanscience.com/ar/articles/flashgaze-training-free-multi-scale-patch-pruning-for-efficient-video-understanding.