Preprint Open access
Which Preferences to Train On? End-to-End Multi-Objective Alignment with an Adversarial Preference Distribution
Aligning large language models (LLMs) with human values is important for safe, efficient, and beneficial AI deployment. However, human values are multifaceted: helpfulness, harmlessness and humor trade off against one another, and different users want different trade-offs. Multi-objective alignment (MOA) addresses this …