الباحثون

Hongbo Ma

المنشورات 4

نسخة أولية وصول مفتوح

How Much Can Language Models Gain from Test-Time Computation?

Bangji Yang, Jingyuan Li, Jiajun Fan وآخرون · 2026

How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that …

نسخة أولية وصول مفتوح

Does Scaling Reinforcement Learning Really Require More Training?

Bangji Yang, Jiajun Fan, Hongbo Ma وآخرون · 2026

Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL histor …

نسخة أولية وصول مفتوح

$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient

Hongbo Ma, Sansheng Cao, Jiajun Fan وآخرون · 2026

LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and …

المؤلفون المشاركون