الباحثون

Hung-yi Lee

المنشورات 6

نسخة أولية وصول مفتوح

A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic

Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identif …

نسخة أولية وصول مفتوح

Can LLM Agents Automate Reinforcement Learning for Text-to-Speech?

Xuanjun Chen, Zixiong Su, Hao Shi وآخرون · 2026

Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate this question with AgenticTTS-Forge, a c …

نسخة أولية وصول مفتوح

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Kuan-Po Huang, Haohe Liu, Puyuan Peng وآخرون · 2026

Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal represe …

نسخة أولية وصول مفتوح

MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses

Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu وآخرون · 2026

Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution o …

نسخة أولية وصول مفتوح

Correlation-Guided Encoder Selection for Multi-Encoder Large Audio-Language Models

Pei-Jun Liao, Hung-Shin Lee, Wenze Ren وآخرون · 2026

Multi-encoder fusion extends Large Audio-Language Models (LALMs) beyond speech-centric recognition, but selecting encoders via intuition or exhaustive search often introduces redundant representations and inflates an already constrained compute budget. We propose CUES (Correlation-gUided Encoder Selection), a lightweig …

المؤلفون المشاركون