Preprint Open access
PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learnin …