نسخة أولية وصول مفتوح
SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding
LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding gen …