الباحثون

Bowen Qi

المنشورات 2

نسخة أولية وصول مفتوح

SkillFormer: Skill-Decomposed Adaptation for Audio Language Models

Lee Seung-woo, Bowen Qi, Kim Min-jun وآخرون · 2026

Audio language models must handle dozens of distinct skills, from pitch comparison and speaker counting to musical tempo estimation and emotion recognition. Joint training on all skills at once causes interference: gains on one skill often come at the cost of another. We propose \textbf{SkillFormer}, which decomposes a …

نسخة أولية وصول مفتوح

AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models

Large audio language models answer questions about speech, sound, and music, yet their accuracy drops sharply on tasks that need fine-grained acoustic analysis. Judging which of two speakers has the higher pitch demands iterative signal-level reasoning that a content question does not. Current models spend the same com …

المؤلفون المشاركون