الباحثون

Ke Hu

المنشورات 3

نسخة أولية وصول مفتوح

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing appr …

نسخة أولية وصول مفتوح

Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

Full-duplex speech-to-speech (S2S) models enable natural conversational AI by allowing simultaneous listening and speaking. However, these models typically lack inherent user speech transcription, which is essential for applications such as conversation logging, accessibility features, and quality monitoring. In this w …

المؤلفون المشاركون