Abstract

We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules and their applicability, assessing conditions from evidence, combining judgments under rules and exceptions, and providing checkable justifications. These demands motivate a benchmark assessing both decisions and their stated grounds. We introduce RGDT-Bench, providing 202.1K condition-level supervision slots across four task tracks and eight supported task-probe combinations that vary access to supporting information. Label-blind extraction and deterministic checks produce labels for warrant completeness: source-referenced coverage and consistency of stated decision grounds. The benchmark attributes failures to four process layers: rule use, condition, evidence, and aggregation, and checks the final outcome. Among evaluable correct responses, warrant incompleteness averages 40.2% across six evaluated LLMs and supported task-probe combinations. Such warrant incompleteness poses potential safety risks and remains difficult to detect: the best of seventeen existing evaluators reaches only 57.69% (random: 50%) task-averaged area under the receiver operating characteristic curve (AUROC). To address this difficulty, we train a simple reward model with warrant supervision. It achieves 69.24% task-averaged AUROC among correct answers, exceeding the matched outcome-supervised baseline by 10.37 pp (percentage points) and the best existing evaluator by 11.55 pp. Beyond completeness assessment, the model outperforms both outcome-supervised baselines across nearly all response-selection comparisons, supporting RGDT-Bench's warrant supervision for RGDT reasoning.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhao, J., Xu, H., Zhang, H., Qian, S., Tang, Y., Wang, X., Sun, K., Wu, P., Zhong, S., & Wang, P. (2026). RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications. https://omanscience.com/en/articles/rgdt-bench-benchmarking-llm-reasoning-for-rule-governed-decisions-and-their-justifications

MLA 9

Zhao, Jianpeng, et al. "RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications." https://omanscience.com/en/articles/rgdt-bench-benchmarking-llm-reasoning-for-rule-governed-decisions-and-their-justifications.

Chicago (author–date)

Zhao, Jianpeng, Haihua Xu, Haoyang Zhang, Shuang Qian, Yixiang Tang, Xintao Wang, Kun Sun, Pei Wu, Shuhan Zhong, and Pengyang Wang. 2026. "RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications." https://omanscience.com/en/articles/rgdt-bench-benchmarking-llm-reasoning-for-rule-governed-decisions-and-their-justifications.

Harvard

Zhao, J., Xu, H., Zhang, H., Qian, S., Tang, Y., Wang, X., Sun, K., Wu, P., Zhong, S. and Wang, P. (2026) 'RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications', Available at: https://omanscience.com/en/articles/rgdt-bench-benchmarking-llm-reasoning-for-rule-governed-decisions-and-their-justifications.

Vancouver

Zhao J, Xu H, Zhang H, Qian S, Tang Y, Wang X, et al. RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications. https://omanscience.com/en/articles/rgdt-bench-benchmarking-llm-reasoning-for-rule-governed-decisions-and-their-justifications

IEEE

J. Zhao, H. Xu, H. Zhang, S. Qian, Y. Tang, X. Wang, K. Sun, P. Wu, S. Zhong, and P. Wang, "RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications," https://omanscience.com/en/articles/rgdt-bench-benchmarking-llm-reasoning-for-rule-governed-decisions-and-their-justifications.