الباحثون

Jufeng Yang

المنشورات 1

نسخة أولية وصول مفتوح

Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning

Yicheng Xue, Han Wu, Jufeng Yang وآخرون · 2026

Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as …

المؤلفون المشاركون