الباحثون

Han-Jia Ye

المنشورات 8

نسخة أولية وصول مفتوح

SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding

Zhi-Kai Chen, Song-Yan Li, De-Chuan Zhan وآخرون · 2026

LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding gen …

نسخة أولية وصول مفتوح

RACE: Residual-Aware Test-Time Adaptation for Neighbor-Rich Time-Series Foundation Model Forecasting

Hao-Nan Shi, Tong Wu, Chen-Cong Sun وآخرون · 2026

Time-series foundation models (TSFMs) perform strongly across forecasting tasks, but their per-series inference is ill-suited to neighbor-rich forecasting, where each query has access to related but nonidentical historical series. Continuous glucose monitoring (CGM) and Web/cloud workloads exemplify this setting: CGM t …

نسخة أولية وصول مفتوح

Consistent Plan-Act for Long-Horizon Agentic Tasks

Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang وآخرون · 2026

Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks, we prompt both agen …

نسخة أولية وصول مفتوح

Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate poo …

نسخة أولية وصول مفتوح

Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Guannan Lai, Gelin Bian, Hao-Xuan Ma وآخرون · 2026

Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision …

المؤلفون المشاركون