Abstract
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Park, G., Jeun, K., Oh, J., Shin, B., Park, B., & Rhu, M. (2026). SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving. https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving
MLA 9
Park, Gunho, et al. "SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving." https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving.
Chicago (author–date)
Park, Gunho, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin, Baeseong Park, and Minsoo Rhu. 2026. "SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving." https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving.
Harvard
Park, G., Jeun, K., Oh, J., Shin, B., Park, B. and Rhu, M. (2026) 'SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving', Available at: https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving.
Vancouver
Park G, Jeun K, Oh J, Shin B, Park B, Rhu M. SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving. https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving
IEEE
G. Park, K. Jeun, J. Oh, B. Shin, B. Park, and M. Rhu, "SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving," https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving.