الباحثون

Lin Gao

المنشورات 2

نسخة أولية وصول مفتوح

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

Yushuai Sun, Zikun Zhou, Lin Gao وآخرون · 2026

Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. …

نسخة أولية وصول مفتوح

Representation Dynamics Reveal Semantic Saliency and Similarity for Visual Token Pruning in MLLMs

Weixuan Li, Zikun Zhou, Xinyi Zhuang وآخرون · 2026

Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflec …

المؤلفون المشاركون