نسخة أولية وصول مفتوح
Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning
Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as …