Abstract
Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Cao, T., Shao, J., Qiu, Y., Atarashi, K., Kashima, H., & Zhao, Q. (2026). Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression. https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression
MLA 9
Cao, Tianxiao, et al. "Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression." https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression.
Chicago (author–date)
Cao, Tianxiao, Jiahe Shao, Yuning Qiu, Kyohei Atarashi, Hisashi Kashima, and Qibin Zhao. 2026. "Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression." https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression.
Harvard
Cao, T., Shao, J., Qiu, Y., Atarashi, K., Kashima, H. and Zhao, Q. (2026) 'Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression', Available at: https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression.
Vancouver
Cao T, Shao J, Qiu Y, Atarashi K, Kashima H, Zhao Q. Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression. https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression
IEEE
T. Cao, J. Shao, Y. Qiu, K. Atarashi, H. Kashima, and Q. Zhao, "Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression," https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression.