Abstract

Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identify a structural pathology of fine-grained upcycling: when fine-grained experts are derived from a single source model, naive routing collapses and downstream accuracy drops to near-random (e.g., on Qwen3-1.7B, Drop-Upcycling-fine-grained reaches only 23.2% average accuracy across 15 benchmarks, essentially matching from-scratch training at 22.2%, while the same method's coarse-grained variant reaches 50.2%). We propose DivMoE, the first framework achieving fine-grained MoE Upcycling with structurally-balanced routing. DivMoE introduces domain-specialized fine-grained expert initialization, deriving experts from dense models that have undergone domain-adaptive continual pre-training, and diversity-constrained routing, a hard structural constraint guaranteeing that each token activates experts from distinct domain groups. Across two base models and 15 benchmarks, DivMoE consistently outperforms six upcycling baselines (55.6% vs. 51.6% for the strongest baseline on Qwen3-1.7B) and strictly improves over the dense base model on every benchmark after Stage 2 continual pre-training -- closing the regression gap that has plagued prior fine-grained upcycling. After supervised fine-tuning on a public reasoning mixture, our 12B-parameter DivMoE model matches Moonlight-MoE (16B) at 64.5% average accuracy while outperforming a controlled NVIDIA-Upcycling baseline by 6.1 percentage points.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lou, Y., Yang, K., Zhang, G., Liu, Y., & You, Y. (2026). DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition. https://omanscience.com/en/articles/divmoe-fine-grained-moe-upcycling-via-cross-domain-expert-composition

MLA 9

Lou, Yuxuan, et al. "DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition." https://omanscience.com/en/articles/divmoe-fine-grained-moe-upcycling-via-cross-domain-expert-composition.

Chicago (author–date)

Lou, Yuxuan, Kai Yang, Geng Zhang, Yong Liu, and Yang You. 2026. "DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition." https://omanscience.com/en/articles/divmoe-fine-grained-moe-upcycling-via-cross-domain-expert-composition.

Harvard

Lou, Y., Yang, K., Zhang, G., Liu, Y. and You, Y. (2026) 'DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition', Available at: https://omanscience.com/en/articles/divmoe-fine-grained-moe-upcycling-via-cross-domain-expert-composition.

Vancouver

Lou Y, Yang K, Zhang G, Liu Y, You Y. DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition. https://omanscience.com/en/articles/divmoe-fine-grained-moe-upcycling-via-cross-domain-expert-composition

IEEE

Y. Lou, K. Yang, G. Zhang, Y. Liu, and Y. You, "DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition," https://omanscience.com/en/articles/divmoe-fine-grained-moe-upcycling-via-cross-domain-expert-composition.