Abstract
Reasoning models often continue generating after their answers have settled. Settle learns when to stop from answer stability in completed traces. It trains the existing end-of-reasoning token while keeping other predictions close to the base model, and requires only ordinary decoding at inference. On MATH-500 with Qwen3-4B, Settle reduces token count by 40% with a 0.5-percentage-point decrease in accuracy. It gains 6.16 percentage points over supervised fine-tuning on the same traces shortened at their first stable answer, at nearly identical token counts. Its stopping score predicts whether a correct answer will remain correct. Settle extends the accuracy-token-count Pareto frontier of the evaluated stopping methods.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Brown, R., Fu, Z., & Russell, C. (2026). Settle: Learning When to Stop Reasoning. https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning
MLA 9
Brown, Ryan, et al. "Settle: Learning When to Stop Reasoning." https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning.
Chicago (author–date)
Brown, Ryan, Zihao Fu, and Chris Russell. 2026. "Settle: Learning When to Stop Reasoning." https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning.
Harvard
Brown, R., Fu, Z. and Russell, C. (2026) 'Settle: Learning When to Stop Reasoning', Available at: https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning.
Vancouver
Brown R, Fu Z, Russell C. Settle: Learning When to Stop Reasoning. https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning
IEEE
R. Brown, Z. Fu, and C. Russell, "Settle: Learning When to Stop Reasoning," https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning.