Authors

Mingjun Xiao

Publications 1

Preprint Open access

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

Haoyu Guo, Yuan Feng, Junlin Lv et al. · 2026

Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight enco …

Co-authors