الملخص

Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical correctness. This work proposes Function-Structured Graph Reinforcement Learning (FSG-RL), connecting subproblem graphs and Python implementations with multi-verifier feedback. The policy first learns to generate code from function graphs through supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) then optimizes the policy using answer-gated rewards and span-level credit assignment. The framework also supports teacher supervision and structured memory. A benchmark curated from Grade School Math 8K (GSM8K), MathQA, MATH, and Omni-MATH pairs public function graphs with private verification specifications. Under a unified evaluation protocol, GRPO improves final-answer accuracy from 43.25% to 67.50% and full solution success from 32.25% to 52.25% over SFT. Continued reinforcement learning (RL) with teacher supervision yields additional gains. The gains extend beyond producing correctly formatted code, supporting verifier-guided reinforcement learning for mathematical reasoning. Code is available at https://github.com/ZihanLiummyycc/FSG-RL.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Liu, Z., & Xie, X. (2026). Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning. https://omanscience.com/ar/articles/function-structured-reinforcement-learning-with-executable-verifiers-for-mathematical-reasoning

MLA 9

Liu, Zihan, and Xurong Xie. "Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning." https://omanscience.com/ar/articles/function-structured-reinforcement-learning-with-executable-verifiers-for-mathematical-reasoning.

شيكاغو (المؤلف–التاريخ)

Liu, Zihan, and Xurong Xie. 2026. "Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning." https://omanscience.com/ar/articles/function-structured-reinforcement-learning-with-executable-verifiers-for-mathematical-reasoning.

هارفارد

Liu, Z. and Xie, X. (2026) 'Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning', Available at: https://omanscience.com/ar/articles/function-structured-reinforcement-learning-with-executable-verifiers-for-mathematical-reasoning.

فانكوفر

Liu Z, Xie X. Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning. https://omanscience.com/ar/articles/function-structured-reinforcement-learning-with-executable-verifiers-for-mathematical-reasoning

IEEE

Z. Liu, and X. Xie, "Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning," https://omanscience.com/ar/articles/function-structured-reinforcement-learning-with-executable-verifiers-for-mathematical-reasoning.