Abstract
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but incorrect reasoning trajectories, whereas weak sharpening leaves the answer distribution diffuse. To resolve this trade-off, we introduce \textbf{Parallel Power Tempering (PPT)}, instantiating power-sharpened LLM sampling via parallel tempering. Running multiple \emph{interacting} replicas in parallel at different sharpening levels allows lower-power replicas to explore diverse reasoning trajectories and higher-power chains to further exploit higher-likelihood responses favored by the sharpened target. Specifically, we tailor \method{} to inference-time sampling by mitigating a truncation bias, identified in prior power samplers, and investigate effective swap strategies under finite memory and compute budgets. Extensive experimentation shows that \method{} substantially improves single-chain power-sharpened sampling and outperforms RL-post-trained models, producing higher-quality reasoning traces and even achieving performance comparable to frontier models.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Theodoropoulos, P., Jiang, N., Duan, X., Hasan, A., Nevmyvaka, Y., Theodorou, E. A., & Deng, W. (2026). Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling. https://omanscience.com/en/articles/explore-broadly-reason-sharply-push-small-models-toward-the-frontier-via-sampling
MLA 9
Theodoropoulos, Panagiotis, et al. "Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling." https://omanscience.com/en/articles/explore-broadly-reason-sharply-push-small-models-toward-the-frontier-via-sampling.
Chicago (author–date)
Theodoropoulos, Panagiotis, Nan Jiang, Xintong Duan, Ali Hasan, Yuriy Nevmyvaka, Evangelos A. Theodorou, and Wei Deng. 2026. "Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling." https://omanscience.com/en/articles/explore-broadly-reason-sharply-push-small-models-toward-the-frontier-via-sampling.
Harvard
Theodoropoulos, P., Jiang, N., Duan, X., Hasan, A., Nevmyvaka, Y., Theodorou, E. A. and Deng, W. (2026) 'Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling', Available at: https://omanscience.com/en/articles/explore-broadly-reason-sharply-push-small-models-toward-the-frontier-via-sampling.
Vancouver
Theodoropoulos P, Jiang N, Duan X, Hasan A, Nevmyvaka Y, Theodorou EA, et al. Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling. https://omanscience.com/en/articles/explore-broadly-reason-sharply-push-small-models-toward-the-frontier-via-sampling
IEEE
P. Theodoropoulos, N. Jiang, X. Duan, A. Hasan, Y. Nevmyvaka, E. A. Theodorou, and W. Deng, "Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling," https://omanscience.com/en/articles/explore-broadly-reason-sharply-push-small-models-toward-the-frontier-via-sampling.