الملخص

Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its dominant component and incorporate the donor's complementary component. With a fixed target for how much of the anchor update to retain, SURGE determines the block size from the weights without testing candidate policies. We evaluate two 1.5B mathematical-reasoning histories, DeepSeek and Nemotron, and one 7B coding history, OLMo. SURGE improves benchmark-average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor. It reaches 54.17% on DeepSeek AIME24 against a measured native maximum of 50.83%, and 83.7% on OLMo HumanEval+ against 82.8%. These gains exceed the observed training curves. Geometric controls support the importance of RL-update structure beyond weight displacement or token reduction alone. Each constructed model runs as a single policy. Our findings identify stored RL history as a reusable scaling resource: the capability available from a training run need not end at its best checkpoint.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Yang, B., Fan, J., Ma, H., Guo, R., & Liu, G. (2026). Does Scaling Reinforcement Learning Really Require More Training? https://omanscience.com/ar/articles/does-scaling-reinforcement-learning-really-require-more-training

MLA 9

Yang, Bangji, et al. "Does Scaling Reinforcement Learning Really Require More Training?" https://omanscience.com/ar/articles/does-scaling-reinforcement-learning-really-require-more-training.

شيكاغو (المؤلف–التاريخ)

Yang, Bangji, Jiajun Fan, Hongbo Ma, Ruihan Guo, and Ge Liu. 2026. "Does Scaling Reinforcement Learning Really Require More Training?" https://omanscience.com/ar/articles/does-scaling-reinforcement-learning-really-require-more-training.

هارفارد

Yang, B., Fan, J., Ma, H., Guo, R. and Liu, G. (2026) 'Does Scaling Reinforcement Learning Really Require More Training?', Available at: https://omanscience.com/ar/articles/does-scaling-reinforcement-learning-really-require-more-training.

فانكوفر

Yang B, Fan J, Ma H, Guo R, Liu G. Does Scaling Reinforcement Learning Really Require More Training? https://omanscience.com/ar/articles/does-scaling-reinforcement-learning-really-require-more-training

IEEE

B. Yang, J. Fan, H. Ma, R. Guo, and G. Liu, "Does Scaling Reinforcement Learning Really Require More Training?," https://omanscience.com/ar/articles/does-scaling-reinforcement-learning-really-require-more-training.