Abstract

Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Cao, T., Shao, J., Qiu, Y., Atarashi, K., Kashima, H., & Zhao, Q. (2026). Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression. https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression

MLA 9

Cao, Tianxiao, et al. "Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression." https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression.

Chicago (author–date)

Cao, Tianxiao, Jiahe Shao, Yuning Qiu, Kyohei Atarashi, Hisashi Kashima, and Qibin Zhao. 2026. "Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression." https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression.

Harvard

Cao, T., Shao, J., Qiu, Y., Atarashi, K., Kashima, H. and Zhao, Q. (2026) 'Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression', Available at: https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression.

Vancouver

Cao T, Shao J, Qiu Y, Atarashi K, Kashima H, Zhao Q. Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression. https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression

IEEE

T. Cao, J. Shao, Y. Qiu, K. Atarashi, H. Kashima, and Q. Zhao, "Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression," https://omanscience.com/en/articles/shared-low-rank-basis-factorization-for-data-free-mixture-of-experts-compression.