نسخة أولية وصول مفتوح
Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding
On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signa …