Abstract
Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent tasks. We propose a frontend-backend architecture where a duplex speech-to-text frontend learns to emit a delegation token and forwards streaming ASR transcripts to a text-based backend LLM for tool calls. Tool-call results from the backend are injected back into the frontend through a lightweight prefill-and-repeat mechanism and then synthesized using streaming TTS to the user. Our approach largely preserves regular duplex turn-taking, interruption handling, and low-latency interaction as it requires minimal modifications to the frontend model. In a single-turn tool-call evaluation, our system achieves 92-97% tool-call recall, competitive tool-call prediction performance, and 81.2% accuracy in rejecting irrelevant calls. When equipped with a larger backend (e.g., Qwen3-235B-A22B), our system achieves competitive results on Full-Duplex-Bench-V3 compared to open and closed source models, and significantly outperforms GPT-realtime-mini and Qwen3-Omni-30B-A3B-Instruct on EVA-Bench. These results demonstrate that backend delegation is an effective and modular approach for combining natural duplex speech interaction with strong agentic tool-call capabilities.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Hu, K., Deng, S., Chen, C., Rastorgueva, E., Casanova, E., Kumar, P., Choudhary, D., Srihari, N., Mahabaleshwarkar, A. S., Trinh, V. A., Essid, S., Olabiyi, O., & Chen, Z. (2026). A frontend-backend architecture for tool calls in full-duplex speech models. https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models
MLA 9
Hu, Ke, et al. "A frontend-backend architecture for tool calls in full-duplex speech models." https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models.
Chicago (author–date)
Hu, Ke, Slyne Deng, Chen Chen, Elena Rastorgueva, Edresson Casanova, Punit Kumar, Dharmendra Choudhary, Nikhil Srihari, Ameya Sunil Mahabaleshwarkar, Viet Anh Trinh, Slim Essid, Oluwatobi Olabiyi, and Zhehuai Chen. 2026. "A frontend-backend architecture for tool calls in full-duplex speech models." https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models.
Harvard
Hu, K., Deng, S., Chen, C., Rastorgueva, E., Casanova, E., Kumar, P., Choudhary, D., Srihari, N., Mahabaleshwarkar, A. S., Trinh, V. A., Essid, S., Olabiyi, O. and Chen, Z. (2026) 'A frontend-backend architecture for tool calls in full-duplex speech models', Available at: https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models.
Vancouver
Hu K, Deng S, Chen C, Rastorgueva E, Casanova E, Kumar P, et al. A frontend-backend architecture for tool calls in full-duplex speech models. https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models
IEEE
K. Hu, S. Deng, C. Chen, E. Rastorgueva, E. Casanova, P. Kumar, D. Choudhary, N. Srihari, A. S. Mahabaleshwarkar, V. A. Trinh, S. Essid, O. Olabiyi, and Z. Chen, "A frontend-backend architecture for tool calls in full-duplex speech models," https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models.