الباحثون

Zeyuan Liu

المنشورات 3

نسخة أولية وصول مفتوح

Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR

Kun Liang, Chenming Tang, Clive Bai وآخرون · 2026

Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges …

نسخة أولية وصول مفتوح

DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents

Hanyang Wang, Zeyuan Liu, Zhengyu Chen وآخرون · 2026

On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, …

المؤلفون المشاركون