Preprint Open access
Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing
Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate poo …