الباحثون

Xilin Jiang

المنشورات 2

نسخة أولية وصول مفتوح

Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners

Xilin Jiang, Shun Zhang, Tejas Jayashankar وآخرون · 2026

We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical a …

نسخة أولية وصول مفتوح

Video Generation Models: A Survey of Post-Training and Alignment

Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni وآخرون · 2026

Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherenc …

المؤلفون المشاركون