Abstract

Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent tasks. We propose a frontend-backend architecture where a duplex speech-to-text frontend learns to emit a delegation token and forwards streaming ASR transcripts to a text-based backend LLM for tool calls. Tool-call results from the backend are injected back into the frontend through a lightweight prefill-and-repeat mechanism and then synthesized using streaming TTS to the user. Our approach largely preserves regular duplex turn-taking, interruption handling, and low-latency interaction as it requires minimal modifications to the frontend model. In a single-turn tool-call evaluation, our system achieves 92-97% tool-call recall, competitive tool-call prediction performance, and 81.2% accuracy in rejecting irrelevant calls. When equipped with a larger backend (e.g., Qwen3-235B-A22B), our system achieves competitive results on Full-Duplex-Bench-V3 compared to open and closed source models, and significantly outperforms GPT-realtime-mini and Qwen3-Omni-30B-A3B-Instruct on EVA-Bench. These results demonstrate that backend delegation is an effective and modular approach for combining natural duplex speech interaction with strong agentic tool-call capabilities.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hu, K., Deng, S., Chen, C., Rastorgueva, E., Casanova, E., Kumar, P., Choudhary, D., Srihari, N., Mahabaleshwarkar, A. S., Trinh, V. A., Essid, S., Olabiyi, O., & Chen, Z. (2026). A frontend-backend architecture for tool calls in full-duplex speech models. https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models

MLA 9

Hu, Ke, et al. "A frontend-backend architecture for tool calls in full-duplex speech models." https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models.

Chicago (author–date)

Hu, Ke, Slyne Deng, Chen Chen, Elena Rastorgueva, Edresson Casanova, Punit Kumar, Dharmendra Choudhary, Nikhil Srihari, Ameya Sunil Mahabaleshwarkar, Viet Anh Trinh, Slim Essid, Oluwatobi Olabiyi, and Zhehuai Chen. 2026. "A frontend-backend architecture for tool calls in full-duplex speech models." https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models.

Harvard

Hu, K., Deng, S., Chen, C., Rastorgueva, E., Casanova, E., Kumar, P., Choudhary, D., Srihari, N., Mahabaleshwarkar, A. S., Trinh, V. A., Essid, S., Olabiyi, O. and Chen, Z. (2026) 'A frontend-backend architecture for tool calls in full-duplex speech models', Available at: https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models.

Vancouver

Hu K, Deng S, Chen C, Rastorgueva E, Casanova E, Kumar P, et al. A frontend-backend architecture for tool calls in full-duplex speech models. https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models

IEEE

K. Hu, S. Deng, C. Chen, E. Rastorgueva, E. Casanova, P. Kumar, D. Choudhary, N. Srihari, A. S. Mahabaleshwarkar, V. A. Trinh, S. Essid, O. Olabiyi, and Z. Chen, "A frontend-backend architecture for tool calls in full-duplex speech models," https://omanscience.com/en/articles/a-frontend-backend-architecture-for-tool-calls-in-full-duplex-speech-models.