Abstract
In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Xiao, X., Zhang, Y., Liu, C., Zhao, L., Chen, J., Zhao, T., Xu, X., Kim, Y., Wang, T., & Xu, M. (2026). FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents. https://omanscience.com/en/articles/forge-form-optimal-routing-of-grounded-evidence-for-frozen-llm-agents
MLA 9
Xiao, Xi, et al. "FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents." https://omanscience.com/en/articles/forge-form-optimal-routing-of-grounded-evidence-for-frozen-llm-agents.
Chicago (author–date)
Xiao, Xi, Yunbei Zhang, Chen Liu, Lin Zhao, Jialin Chen, Tianchen Zhao, Xiang Xu, Youngeun Kim, Tianyang Wang, and Min Xu. 2026. "FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents." https://omanscience.com/en/articles/forge-form-optimal-routing-of-grounded-evidence-for-frozen-llm-agents.
Harvard
Xiao, X., Zhang, Y., Liu, C., Zhao, L., Chen, J., Zhao, T., Xu, X., Kim, Y., Wang, T. and Xu, M. (2026) 'FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents', Available at: https://omanscience.com/en/articles/forge-form-optimal-routing-of-grounded-evidence-for-frozen-llm-agents.
Vancouver
Xiao X, Zhang Y, Liu C, Zhao L, Chen J, Zhao T, et al. FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents. https://omanscience.com/en/articles/forge-form-optimal-routing-of-grounded-evidence-for-frozen-llm-agents
IEEE
X. Xiao, Y. Zhang, C. Liu, L. Zhao, J. Chen, T. Zhao, X. Xu, Y. Kim, T. Wang, and M. Xu, "FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents," https://omanscience.com/en/articles/forge-form-optimal-routing-of-grounded-evidence-for-frozen-llm-agents.