Preprint Open access
Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introdu …