Abstract
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific data scaling (TSDS) technique to enhance UniAE-MoE's understanding capabilities. On the XARES-LLM benchmark, UniAE-MoE attains a score of 0.802, achieving state-of-the-art performance. It also delivers top-tier performance in the official Interspeech 2026 Audio Encoder Capability Challenge, further demonstrating robust generalization across diverse audio tasks. Together, these results validate the effectiveness of UniAE-MoE for unified audio understanding across speech, music, and general audio domains.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Cai, S., Zhang, Z., Nie, Z., Peng, J., Xie, J., & Wu, Z. (2026). UniAE-MoE: A Unified Audio Encoder via Mixture of Experts. https://omanscience.com/en/articles/uniae-moe-a-unified-audio-encoder-via-mixture-of-experts
MLA 9
Cai, Shengbo, et al. "UniAE-MoE: A Unified Audio Encoder via Mixture of Experts." https://omanscience.com/en/articles/uniae-moe-a-unified-audio-encoder-via-mixture-of-experts.
Chicago (author–date)
Cai, Shengbo, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, and Zhiyong Wu. 2026. "UniAE-MoE: A Unified Audio Encoder via Mixture of Experts." https://omanscience.com/en/articles/uniae-moe-a-unified-audio-encoder-via-mixture-of-experts.
Harvard
Cai, S., Zhang, Z., Nie, Z., Peng, J., Xie, J. and Wu, Z. (2026) 'UniAE-MoE: A Unified Audio Encoder via Mixture of Experts', Available at: https://omanscience.com/en/articles/uniae-moe-a-unified-audio-encoder-via-mixture-of-experts.
Vancouver
Cai S, Zhang Z, Nie Z, Peng J, Xie J, Wu Z. UniAE-MoE: A Unified Audio Encoder via Mixture of Experts. https://omanscience.com/en/articles/uniae-moe-a-unified-audio-encoder-via-mixture-of-experts
IEEE
S. Cai, Z. Zhang, Z. Nie, J. Peng, J. Xie, and Z. Wu, "UniAE-MoE: A Unified Audio Encoder via Mixture of Experts," https://omanscience.com/en/articles/uniae-moe-a-unified-audio-encoder-via-mixture-of-experts.