نسخة أولية وصول مفتوح
PHIRL: Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning
Human demonstrations provide dense policy-level information but sometimes lack local precision. Human feedback presents accurate local critiques, but offers sparse evaluations rather than direct policy guidance. We propose Progress-Heuristicized Inverse Reinforcement Learning (PHIRL), a data-efficient framework that le …