Abstract

Reasoning models often continue generating after their answers have settled. Settle learns when to stop from answer stability in completed traces. It trains the existing end-of-reasoning token while keeping other predictions close to the base model, and requires only ordinary decoding at inference. On MATH-500 with Qwen3-4B, Settle reduces token count by 40% with a 0.5-percentage-point decrease in accuracy. It gains 6.16 percentage points over supervised fine-tuning on the same traces shortened at their first stable answer, at nearly identical token counts. Its stopping score predicts whether a correct answer will remain correct. Settle extends the accuracy-token-count Pareto frontier of the evaluated stopping methods.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Brown, R., Fu, Z., & Russell, C. (2026). Settle: Learning When to Stop Reasoning. https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning

MLA 9

Brown, Ryan, et al. "Settle: Learning When to Stop Reasoning." https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning.

Chicago (author–date)

Brown, Ryan, Zihao Fu, and Chris Russell. 2026. "Settle: Learning When to Stop Reasoning." https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning.

Harvard

Brown, R., Fu, Z. and Russell, C. (2026) 'Settle: Learning When to Stop Reasoning', Available at: https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning.

Vancouver

Brown R, Fu Z, Russell C. Settle: Learning When to Stop Reasoning. https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning

IEEE

R. Brown, Z. Fu, and C. Russell, "Settle: Learning When to Stop Reasoning," https://omanscience.com/en/articles/settle-learning-when-to-stop-reasoning.