الباحثون

Hao Shi

المنشورات 3

نسخة أولية وصول مفتوح

Can LLM Agents Automate Reinforcement Learning for Text-to-Speech?

Xuanjun Chen, Zixiong Su, Hao Shi وآخرون · 2026

Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate this question with AgenticTTS-Forge, a c …

نسخة أولية وصول مفتوح

What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

Renping Zhou, Zanlin Ni, Zihao Fan وآخرون · 2026

World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirel …

نسخة أولية وصول مفتوح

Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition

Hao Shi, Yun Liu, Xuehao Yang وآخرون · 2026

Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an identical target, so their reading distinction is absent from the supervision interface and can …

المؤلفون المشاركون