Abstract

Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING: (Solver-guided Process Rewards for Novel LogIcal ReasoNing Step Generation). SPRING uses SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. It introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, it designs process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation across three logical reasoning benchmarks, ZebraLogic, AR-LSAT, and Knights and Knaves, and four LLMs shows that SPRING consistently outperforms base LLMs, outcome-only reward baselines, and Logic-LM. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and strongest outcome-only baseline, respectively. On AR-LSAT, it improves overall accuracy by up to 64.93 and 12.14 points, respectively. On Knights and Knaves, SPRING achieves up to 93.14 puzzle accuracy and 96.05 person accuracy.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ali, M. A., Wang, W., Wang, H., & Raza, M. (2026). Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning. https://omanscience.com/en/articles/rewarding-novel-deductions-solver-guided-process-supervision-for-logical-reasoning

MLA 9

Ali, Muhammad Asif, et al. "Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning." https://omanscience.com/en/articles/rewarding-novel-deductions-solver-guided-process-supervision-for-logical-reasoning.

Chicago (author–date)

Ali, Muhammad Asif, Wenqing Wang, Huan Wang, and Mohammad Raza. 2026. "Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning." https://omanscience.com/en/articles/rewarding-novel-deductions-solver-guided-process-supervision-for-logical-reasoning.

Harvard

Ali, M. A., Wang, W., Wang, H. and Raza, M. (2026) 'Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning', Available at: https://omanscience.com/en/articles/rewarding-novel-deductions-solver-guided-process-supervision-for-logical-reasoning.

Vancouver

Ali MA, Wang W, Wang H, Raza M. Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning. https://omanscience.com/en/articles/rewarding-novel-deductions-solver-guided-process-supervision-for-logical-reasoning

IEEE

M. A. Ali, W. Wang, H. Wang, and M. Raza, "Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning," https://omanscience.com/en/articles/rewarding-novel-deductions-solver-guided-process-supervision-for-logical-reasoning.