الملخص
Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine. The router and mask scores are optimized jointly, while the underlying FFN weight values remain unchanged. This formulation supports neuron-structured, semi-structured, and unstructured experts within the same routing architecture. Our main configuration uses four 2:4 experts with top-2 routing, where two half-dense expert passes have the nominal FFN arithmetic of one dense pass, without requiring independent expert weight matrices. On five vision-language benchmarks with Qwen and Gemma backbones, this configuration achieves the highest performance among the compared baselines. Comparisons across mask granularities, routing interventions, and compute-matched controls distinguish the effects of learned connectivity from expert activation count. These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zhang, M., Bai, Y., Wang, Z., Huang, Y., Huang, Y., Wang, H., Zeng, H., & Fu, Y. (2026). MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling. https://omanscience.com/ar/articles/maskerade-token-routed-mask-experts-for-dense-to-moe-upcycling
MLA 9
Zhang, Mingyuan, et al. "MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling." https://omanscience.com/ar/articles/maskerade-token-routed-mask-experts-for-dense-to-moe-upcycling.
شيكاغو (المؤلف–التاريخ)
Zhang, Mingyuan, Yue Bai, Zhongruo Wang, Yupin Huang, Yiyang Huang, Hailing Wang, Huimin Zeng, and Yun Fu. 2026. "MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling." https://omanscience.com/ar/articles/maskerade-token-routed-mask-experts-for-dense-to-moe-upcycling.
هارفارد
Zhang, M., Bai, Y., Wang, Z., Huang, Y., Huang, Y., Wang, H., Zeng, H. and Fu, Y. (2026) 'MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling', Available at: https://omanscience.com/ar/articles/maskerade-token-routed-mask-experts-for-dense-to-moe-upcycling.
فانكوفر
Zhang M, Bai Y, Wang Z, Huang Y, Huang Y, Wang H, et al. MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling. https://omanscience.com/ar/articles/maskerade-token-routed-mask-experts-for-dense-to-moe-upcycling
IEEE
M. Zhang, Y. Bai, Z. Wang, Y. Huang, Y. Huang, H. Wang, H. Zeng, and Y. Fu, "MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling," https://omanscience.com/ar/articles/maskerade-token-routed-mask-experts-for-dense-to-moe-upcycling.