Abstract

Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Park, G., Jeun, K., Oh, J., Shin, B., Park, B., & Rhu, M. (2026). SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving. https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving

MLA 9

Park, Gunho, et al. "SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving." https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving.

Chicago (author–date)

Park, Gunho, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin, Baeseong Park, and Minsoo Rhu. 2026. "SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving." https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving.

Harvard

Park, G., Jeun, K., Oh, J., Shin, B., Park, B. and Rhu, M. (2026) 'SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving', Available at: https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving.

Vancouver

Park G, Jeun K, Oh J, Shin B, Park B, Rhu M. SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving. https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving

IEEE

G. Park, K. Jeun, J. Oh, B. Shin, B. Park, and M. Rhu, "SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving," https://omanscience.com/en/articles/slimwise-decoupling-expert-pruning-across-prefill-and-decode-for-efficient-moe-serving.