Abstract

LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding generates these calls token by token, incurring substantial latency for requests involving multiple calls or many argument fields. The explicit argument structure offers opportunities for parallel generation, but later argument values may depend on preceding fields and calls, so independently generated values can differ from the target model's output. We present SchemaFill, a framework for efficient LLM tool calling through slot-parallel speculative decoding. SchemaFill generates future slot values concurrently as candidates, without requiring advance knowledge of the actual call sequence or argument values. Candidates spanning multiple fields and calls are concatenated for verification by the target model under the actual output prefix. Only verified tokens are committed, and the target supplies corrections when candidates disagree. This applies target verification while exploiting parallelism across slots and calls. On Glaive and BFCL, SchemaFill achieves up to a 4.05$\times$ improvement in end-to-end throughput over autoregressive decoding. Code is available at https://github.com/Czzzk/SchemaFill.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chen, Z. K., Li, S. Y., Zhan, D. C., & Ye, H. J. (2026). SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding. https://omanscience.com/en/articles/schemafill-efficient-llm-tool-calling-via-slot-parallel-speculative-decoding

MLA 9

Chen, Zhi-Kai, et al. "SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding." https://omanscience.com/en/articles/schemafill-efficient-llm-tool-calling-via-slot-parallel-speculative-decoding.

Chicago (author–date)

Chen, Zhi-Kai, Song-Yan Li, De-Chuan Zhan, and Han-Jia Ye. 2026. "SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding." https://omanscience.com/en/articles/schemafill-efficient-llm-tool-calling-via-slot-parallel-speculative-decoding.

Harvard

Chen, Z. K., Li, S. Y., Zhan, D. C. and Ye, H. J. (2026) 'SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding', Available at: https://omanscience.com/en/articles/schemafill-efficient-llm-tool-calling-via-slot-parallel-speculative-decoding.

Vancouver

Chen ZK, Li SY, Zhan DC, Ye HJ. SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding. https://omanscience.com/en/articles/schemafill-efficient-llm-tool-calling-via-slot-parallel-speculative-decoding

IEEE

Z. K. Chen, S. Y. Li, D. C. Zhan, and H. J. Ye, "SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding," https://omanscience.com/en/articles/schemafill-efficient-llm-tool-calling-via-slot-parallel-speculative-decoding.