نسخة أولية وصول مفتوح
FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding
Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational ove …