الباحثون

Weijie Liu

المنشورات 2

نسخة أولية وصول مفتوح

Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR

Kun Liang, Chenming Tang, Clive Bai وآخرون · 2026

Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges …

نسخة أولية وصول مفتوح

Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors

Yuchen Cai, Ding Cao, Qixiang Yin وآخرون · 2026

Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional e …

المؤلفون المشاركون