الملخص

Verifier-guided decoding can prevent harmful reasoning steps from contaminating subsequent generation, but typically relies on an external learned verifier. We ask whether a language model can instead reject its own bad reasoning steps. We define a prefix's recoverability as the probability that the frozen generator can complete it correctly. Diagnostics show that adjacent recoverability changes are often difficult to resolve with practical Monte Carlo budgets, while same-prefix candidates exhibit a sparse low-recoverability tail. We introduce Self-Step Rejection (SSR), which trains a lightweight LoRA acceptance gate on the generator backbone while keeping the base model frozen. SSR uses confidence-qualified first-passage supervision: steps before the first resolved crossing of a root-relative recoverability barrier are accepted, the crossing step is rejected, and unresolved steps and suffixes are excluded. Training combines pointwise classification, same-prefix pairwise learning, and group-relative policy refinement using final-answer correctness. At inference, SSR accepts candidates or resamples from the unchanged prefix under rejection budgets, without an external learned verifier. Across three reasoning models and five mathematical reasoning benchmarks, SSR improves macro-average accuracy over single-pass decoding by 5.4--10.1 points using 1.21--1.40x as many generated tokens, and achieves the highest macro-average accuracy among evaluated step-level methods. Full-solution scaling methods require 4.47--8.27x the single-pass token cost for comparable performance.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Xiong, S., Liu, X., Jin, Y., Wang, X., & Gao, J. (2026). Can Language Models Learn to Reject Their Own Bad Reasoning Steps? https://omanscience.com/ar/articles/can-language-models-learn-to-reject-their-own-bad-reasoning-steps

MLA 9

Xiong, Siheng, et al. "Can Language Models Learn to Reject Their Own Bad Reasoning Steps?" https://omanscience.com/ar/articles/can-language-models-learn-to-reject-their-own-bad-reasoning-steps.

شيكاغو (المؤلف–التاريخ)

Xiong, Siheng, Xiaoze Liu, Yiqiao Jin, Xiaoqian Wang, and Jing Gao. 2026. "Can Language Models Learn to Reject Their Own Bad Reasoning Steps?" https://omanscience.com/ar/articles/can-language-models-learn-to-reject-their-own-bad-reasoning-steps.

هارفارد

Xiong, S., Liu, X., Jin, Y., Wang, X. and Gao, J. (2026) 'Can Language Models Learn to Reject Their Own Bad Reasoning Steps?', Available at: https://omanscience.com/ar/articles/can-language-models-learn-to-reject-their-own-bad-reasoning-steps.

فانكوفر

Xiong S, Liu X, Jin Y, Wang X, Gao J. Can Language Models Learn to Reject Their Own Bad Reasoning Steps? https://omanscience.com/ar/articles/can-language-models-learn-to-reject-their-own-bad-reasoning-steps

IEEE

S. Xiong, X. Liu, Y. Jin, X. Wang, and J. Gao, "Can Language Models Learn to Reject Their Own Bad Reasoning Steps?," https://omanscience.com/ar/articles/can-language-models-learn-to-reject-their-own-bad-reasoning-steps.