الباحثون

Xikun Zhang

المنشورات 1

نسخة أولية وصول مفتوح

Bellman Policy Optimization

Zhuoqing Song, Haotian Xu, Xikun Zhang وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to …

المؤلفون المشاركون