Abstract

Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss. Our method introduces a cluster confusion matrix for domain-aware dataset curation and applies dynamic batching for efficient training. Experiments show that L0-MoE achieves up to 2.5x speedup over dense models while maintaining competitive performance, outperforming existing LLM acceleration baselines.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, Z., Yang, J., Tao, Z., & Chen, M. (2026). Accelerating Dense LLMs via L0-regularized Mixture-of-Experts. https://omanscience.com/en/articles/accelerating-dense-llms-via-l0-regularized-mixture-of-experts

MLA 9

Zhang, Zhenyu, et al. "Accelerating Dense LLMs via L0-regularized Mixture-of-Experts." https://omanscience.com/en/articles/accelerating-dense-llms-via-l0-regularized-mixture-of-experts.

Chicago (author–date)

Zhang, Zhenyu, Jiudong Yang, Zhaowen Tao, and Meng Chen. 2026. "Accelerating Dense LLMs via L0-regularized Mixture-of-Experts." https://omanscience.com/en/articles/accelerating-dense-llms-via-l0-regularized-mixture-of-experts.

Harvard

Zhang, Z., Yang, J., Tao, Z. and Chen, M. (2026) 'Accelerating Dense LLMs via L0-regularized Mixture-of-Experts', Available at: https://omanscience.com/en/articles/accelerating-dense-llms-via-l0-regularized-mixture-of-experts.

Vancouver

Zhang Z, Yang J, Tao Z, Chen M. Accelerating Dense LLMs via L0-regularized Mixture-of-Experts. https://omanscience.com/en/articles/accelerating-dense-llms-via-l0-regularized-mixture-of-experts

IEEE

Z. Zhang, J. Yang, Z. Tao, and M. Chen, "Accelerating Dense LLMs via L0-regularized Mixture-of-Experts," https://omanscience.com/en/articles/accelerating-dense-llms-via-l0-regularized-mixture-of-experts.