الملخص

Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs with fewer than 200M parameters. We introduce a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM. Its architecture connects a compact CED-Small audio encoder to SmolLM2-135M through a frequency-merging mapper. With supervision drawn from ReasonAQA, AudioMCQ, and AVQA, the model undergoes three training stages: audio-language alignment (Stage 1), audio-dependent fine-tuning (Stage 2), and post-training (Stage 3) aimed at strengthening weak skills while retaining learned capabilities. Across five random seeds, Mizar achieves mean accuracies of 52.92% on MMAU, 42.42% on MMAR, and 36.02% on ADQA-clean, surpassing the previous best-performing ALM below 200M parameters on all three benchmarks. It also supports local inference on a single CPU: on questions from the MMAU benchmark, the mean latency from opening the audio file to generating a complete answer is 1.09 seconds. Code and checkpoints are available at https://github.com/KaiyangLi1992/Mizar_159M.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Li, K., Han, S., Tian, Y., & Ji, S. (2026). Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding. https://omanscience.com/ar/articles/mizar-a-159m-parameter-audio-language-model-for-audio-understanding

MLA 9

Li, Kaiyang, et al. "Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding." https://omanscience.com/ar/articles/mizar-a-159m-parameter-audio-language-model-for-audio-understanding.

شيكاغو (المؤلف–التاريخ)

Li, Kaiyang, Shaobo Han, Yue Tian, and Shihao Ji. 2026. "Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding." https://omanscience.com/ar/articles/mizar-a-159m-parameter-audio-language-model-for-audio-understanding.

هارفارد

Li, K., Han, S., Tian, Y. and Ji, S. (2026) 'Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding', Available at: https://omanscience.com/ar/articles/mizar-a-159m-parameter-audio-language-model-for-audio-understanding.

فانكوفر

Li K, Han S, Tian Y, Ji S. Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding. https://omanscience.com/ar/articles/mizar-a-159m-parameter-audio-language-model-for-audio-understanding

IEEE

K. Li, S. Han, Y. Tian, and S. Ji, "Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding," https://omanscience.com/ar/articles/mizar-a-159m-parameter-audio-language-model-for-audio-understanding.