الباحثون

Zhenyu Tian

المنشورات 1

نسخة أولية وصول مفتوح

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian وآخرون · 2026

Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure …

المؤلفون المشاركون