الباحثون

Jianyuan Guo

المنشورات 2

نسخة أولية وصول مفتوح

Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning

Yicheng Xue, Han Wu, Jufeng Yang وآخرون · 2026

Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as …

نسخة أولية وصول مفتوح

Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking

Haoyang Luo, Linwei Tao, Jie Gui وآخرون · 2026

Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly at …

المؤلفون المشاركون