الباحثون

Minsoo Rhu

المنشورات 1

نسخة أولية وصول مفتوح

SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving

Gunho Park, Kyoungho Jeun, Juntaek Oh وآخرون · 2026

Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughp …

المؤلفون المشاركون