الباحثون

Alvin Cheung

المنشورات 1

نسخة أولية وصول مفتوح

EasyPPO: Stabilizing the Critic Is Key

Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao وآخرون · 2026

A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for la …

المؤلفون المشاركون