الباحثون

Shinji Watanabe

المنشورات 4

نسخة أولية وصول مفتوح

Code-Switching Spoken Language Identification as Multi-Label Set Prediction

Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utteran …

نسخة أولية وصول مفتوح

YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio. We fi …

نسخة أولية وصول مفتوح

HaikuS2S: A Cascaded System For Responding In Verse

Devangi Sharma, Sophia Judicke, Glenda Tan وآخرون · 2026

Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation. However, these models do not model haiku's …

المؤلفون المشاركون