الملخص
Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder hidden states through a small connector into a frozen LLM and reads each question as a single-token distribution over its options. Nothing is decoded, and an 8-GPU node answers 80 decisions about eight utterances in about 0.1 s. With a last-layer connector, spoken QA stays close to reading the transcript (90% vs. 91%). DuplexJev also hears the speaker: gender and emotion accuracy both reach 90% (from 55% and 28%) with a cross-attention connector, whose spoken QA drops by only 1 point (83% to 82%). We train decisions with cross-entropy on the read-out answer token, instead of the usual transcript distillation, whose teacher never hears the voice, and keep distillation for content. Encoders and LLMs are interchangeable; we release weights, training recipe, a batched-inference pipeline for full-duplex serving and a bilingual spoken-QA set.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Jin, J., Ma, Z., Yin, M., Chen, J., Song, H., Pang, Z., & Zhang, X. (2026). Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript. https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript
MLA 9
Jin, Jie, et al. "Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript." https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript.
شيكاغو (المؤلف–التاريخ)
Jin, Jie, Ziyin Ma, Min Yin, Jinyu Chen, Haigang Song, Zhikun Pang, and Xiaowen Zhang. 2026. "Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript." https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript.
هارفارد
Jin, J., Ma, Z., Yin, M., Chen, J., Song, H., Pang, Z. and Zhang, X. (2026) 'Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript', Available at: https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript.
فانكوفر
Jin J, Ma Z, Yin M, Chen J, Song H, Pang Z, et al. Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript. https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript
IEEE
J. Jin, Z. Ma, M. Yin, J. Chen, H. Song, Z. Pang, and X. Zhang, "Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript," https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript.