الملخص
Audio language models must handle dozens of distinct skills, from pitch comparison and speaker counting to musical tempo estimation and emotion recognition. Joint training on all skills at once causes interference: gains on one skill often come at the cost of another. We propose \textbf{SkillFormer}, which decomposes audio understanding into skill-specific low-rank adapters and composes them at inference time through a learned router. The router examines the question to decide which adapters to activate and how much weight each should carry, so that a pitch query engages different parameters than a genre classification query. An alternating training schedule updates each adapter on its own skill cluster before jointly calibrating the router, preventing the gradient conflicts that arise in standard multi-task optimization. SkillFormer adds fewer than 4\% of the base model's parameters and requires no changes to the audio encoder or language backbone. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, it raises the average accuracy by 2.5 to 4.1 points, with balanced gains across perception, reasoning, and semantic subcategories.
الكلمات المفتاحية
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Seung-woo, L., Qi, B., Min-jun, K., & Won-young, J. (2026). SkillFormer: Skill-Decomposed Adaptation for Audio Language Models. https://omanscience.com/ar/articles/skillformer-skill-decomposed-adaptation-for-audio-language-models
MLA 9
Seung-woo, Lee, et al. "SkillFormer: Skill-Decomposed Adaptation for Audio Language Models." https://omanscience.com/ar/articles/skillformer-skill-decomposed-adaptation-for-audio-language-models.
شيكاغو (المؤلف–التاريخ)
Seung-woo, Lee, Bowen Qi, Kim Min-jun, and Jang Won-young. 2026. "SkillFormer: Skill-Decomposed Adaptation for Audio Language Models." https://omanscience.com/ar/articles/skillformer-skill-decomposed-adaptation-for-audio-language-models.
هارفارد
Seung-woo, L., Qi, B., Min-jun, K. and Won-young, J. (2026) 'SkillFormer: Skill-Decomposed Adaptation for Audio Language Models', Available at: https://omanscience.com/ar/articles/skillformer-skill-decomposed-adaptation-for-audio-language-models.
فانكوفر
Seung-woo L, Qi B, Min-jun K, Won-young J. SkillFormer: Skill-Decomposed Adaptation for Audio Language Models. https://omanscience.com/ar/articles/skillformer-skill-decomposed-adaptation-for-audio-language-models
IEEE
L. Seung-woo, B. Qi, K. Min-jun, and J. Won-young, "SkillFormer: Skill-Decomposed Adaptation for Audio Language Models," https://omanscience.com/ar/articles/skillformer-skill-decomposed-adaptation-for-audio-language-models.