الباحثون

Ruihan Guo

المنشورات 2

نسخة أولية وصول مفتوح

How Much Can Language Models Gain from Test-Time Computation?

Bangji Yang, Jingyuan Li, Jiajun Fan وآخرون · 2026

How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that …

نسخة أولية وصول مفتوح

Does Scaling Reinforcement Learning Really Require More Training?

Bangji Yang, Jiajun Fan, Hongbo Ma وآخرون · 2026

Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL histor …

المؤلفون المشاركون