الملخص
Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $Θ(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of $Θ(h/r)$ more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\&A benchmarks after pretraining, while matching reasoning ability. Code available at https://github.com/Matheart/MoRE_code.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wong, H., Goel, S., & Boix-Adserà, E. (2026). MoRE: Scaling mixture of experts with hardware-aware low-rank routing. https://omanscience.com/ar/articles/more-scaling-mixture-of-experts-with-hardware-aware-low-rank-routing
MLA 9
Wong, Honam, et al. "MoRE: Scaling mixture of experts with hardware-aware low-rank routing." https://omanscience.com/ar/articles/more-scaling-mixture-of-experts-with-hardware-aware-low-rank-routing.
شيكاغو (المؤلف–التاريخ)
Wong, Honam, Surbhi Goel, and Enric Boix-Adserà. 2026. "MoRE: Scaling mixture of experts with hardware-aware low-rank routing." https://omanscience.com/ar/articles/more-scaling-mixture-of-experts-with-hardware-aware-low-rank-routing.
هارفارد
Wong, H., Goel, S. and Boix-Adserà, E. (2026) 'MoRE: Scaling mixture of experts with hardware-aware low-rank routing', Available at: https://omanscience.com/ar/articles/more-scaling-mixture-of-experts-with-hardware-aware-low-rank-routing.
فانكوفر
Wong H, Goel S, Boix-Adserà E. MoRE: Scaling mixture of experts with hardware-aware low-rank routing. https://omanscience.com/ar/articles/more-scaling-mixture-of-experts-with-hardware-aware-low-rank-routing
IEEE
H. Wong, S. Goel, and E. Boix-Adserà, "MoRE: Scaling mixture of experts with hardware-aware low-rank routing," https://omanscience.com/ar/articles/more-scaling-mixture-of-experts-with-hardware-aware-low-rank-routing.