Abstract

Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoning. Existing approaches largely prescribe these decisions through their harness; we instead shift them to the model. We introduce 1) Hermes, a family of simple, configurable harnesses that progressively varies model control over context allocation and reuse, and 2) Hermes-Learn, a two-stage framework for learning these capabilities. We find that capable models can exploit this flexibility to scale with additional inference-time compute, while smaller open-source models initially struggle to do so. Training with Hermes-Learn closes this gap, inducing adaptive contextual reasoning strategies that vary with both the problem and the progress of reasoning. These gains generalize across benchmarks and models, extrapolate beyond the inference-time compute seen during training, and transfer to complementary test-time scaling methods beyond Hermes.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, X., Goswami, M., Liu, H., Kanakaris, N., Huang, L., Jana, P., Blöbaum, P., & Jain, P. (2026). Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling. https://omanscience.com/en/articles/hermes-learning-contextual-reasoning-unlocks-test-time-scaling

MLA 9

Li, Xinyu, et al. "Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling." https://omanscience.com/en/articles/hermes-learning-contextual-reasoning-unlocks-test-time-scaling.

Chicago (author–date)

Li, Xinyu, Mononito Goswami, Hao Liu, Nikos Kanakaris, Langlin Huang, Prithwish Jana, Patrick Blöbaum, and Purak Jain. 2026. "Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling." https://omanscience.com/en/articles/hermes-learning-contextual-reasoning-unlocks-test-time-scaling.

Harvard

Li, X., Goswami, M., Liu, H., Kanakaris, N., Huang, L., Jana, P., Blöbaum, P. and Jain, P. (2026) 'Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling', Available at: https://omanscience.com/en/articles/hermes-learning-contextual-reasoning-unlocks-test-time-scaling.

Vancouver

Li X, Goswami M, Liu H, Kanakaris N, Huang L, Jana P, et al. Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling. https://omanscience.com/en/articles/hermes-learning-contextual-reasoning-unlocks-test-time-scaling

IEEE

X. Li, M. Goswami, H. Liu, N. Kanakaris, L. Huang, P. Jana, P. Blöbaum, and P. Jain, "Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling," https://omanscience.com/en/articles/hermes-learning-contextual-reasoning-unlocks-test-time-scaling.