الباحثون

Mayank Mishra

المنشورات 3

نسخة أولية وصول مفتوح

Feeling Through the Load: Compliant Quadruped Locomotion under Payload Interactions

Quadruped robots are increasingly expected to carry objects while moving through human environments. But what happens when a person interacts directly with the payload rather than with the robot? If the payload is unrestrained, the robot must distinguish intentional external interactions from ordinary payload motion, w …

نسخة أولية وصول مفتوح

EasyPPO: Stabilizing the Critic Is Key

Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao وآخرون · 2026

A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for la …

المؤلفون المشاركون