الملخص

Full-duplex voice agents make many small, closed decisions, which current systems answer by slow autoregressive decoding. We propose DuplexJev, which feeds ASR-encoder hidden states through a small connector into a frozen LLM and reads each question as a single-token distribution over its options. Nothing is decoded, and an 8-GPU node answers 80 decisions about eight utterances in about 0.1 s. With a last-layer connector, spoken QA stays close to reading the transcript (90% vs. 91%). DuplexJev also hears the speaker: gender and emotion accuracy both reach 90% (from 55% and 28%) with a cross-attention connector, whose spoken QA drops by only 1 point (83% to 82%). We train decisions with cross-entropy on the read-out answer token, instead of the usual transcript distillation, whose teacher never hears the voice, and keep distillation for content. Encoders and LLMs are interchangeable; we release weights, training recipe, a batched-inference pipeline for full-duplex serving and a bilingual spoken-QA set.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Jin, J., Ma, Z., Yin, M., Chen, J., Song, H., Pang, Z., & Zhang, X. (2026). Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript. https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript

MLA 9

Jin, Jie, et al. "Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript." https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript.

شيكاغو (المؤلف–التاريخ)

Jin, Jie, Ziyin Ma, Min Yin, Jinyu Chen, Haigang Song, Zhikun Pang, and Xiaowen Zhang. 2026. "Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript." https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript.

هارفارد

Jin, J., Ma, Z., Yin, M., Chen, J., Song, H., Pang, Z. and Zhang, X. (2026) 'Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript', Available at: https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript.

فانكوفر

Jin J, Ma Z, Yin M, Chen J, Song H, Pang Z, et al. Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript. https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript

IEEE

J. Jin, Z. Ma, M. Yin, J. Chen, H. Song, Z. Pang, and X. Zhang, "Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript," https://omanscience.com/ar/articles/batched-speech-decisions-without-decoding-single-token-supervision-lets-a-frozen-llm-hear-beyond-the-transcript.