Abstract

Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Mousa, A., Rachman, R., Karavis, N., Caprio, M., & Allmendinger, R. (2026). PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots. https://omanscience.com/en/articles/promo-preference-conditioned-multi-objective-reinforcement-learning-for-quadrupedal-robots

MLA 9

Mousa, Amr, et al. "PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots." https://omanscience.com/en/articles/promo-preference-conditioned-multi-objective-reinforcement-learning-for-quadrupedal-robots.

Chicago (author–date)

Mousa, Amr, Rifny Rachman, Neil Karavis, Michele Caprio, and Richard Allmendinger. 2026. "PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots." https://omanscience.com/en/articles/promo-preference-conditioned-multi-objective-reinforcement-learning-for-quadrupedal-robots.

Harvard

Mousa, A., Rachman, R., Karavis, N., Caprio, M. and Allmendinger, R. (2026) 'PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots', Available at: https://omanscience.com/en/articles/promo-preference-conditioned-multi-objective-reinforcement-learning-for-quadrupedal-robots.

Vancouver

Mousa A, Rachman R, Karavis N, Caprio M, Allmendinger R. PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots. https://omanscience.com/en/articles/promo-preference-conditioned-multi-objective-reinforcement-learning-for-quadrupedal-robots

IEEE

A. Mousa, R. Rachman, N. Karavis, M. Caprio, and R. Allmendinger, "PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots," https://omanscience.com/en/articles/promo-preference-conditioned-multi-objective-reinforcement-learning-for-quadrupedal-robots.