الباحثون

Dongdong Ren

المنشورات 2

نسخة أولية وصول مفتوح

TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning

Jing Wang, Zhiping Wu, Dongdong Ren وآخرون · 2026

Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then disc …

نسخة أولية وصول مفتوح

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

Zhiping Wu, Dongdong Ren, Yangchengyu Zhou وآخرون · 2026

Large vision-language models (VLMs) can be efficiently deployed under stringent memory and latency constraints through post training quantization (PTQ). However, most PTQ methods are designed for unimodal large language models (LLMs). These methods treat quantization errors as isotropic perturbations under the Euclidea …

المؤلفون المشاركون