Preprint Open access
Collaborative Personalized Preference Alignment for LLMs under Data Deficiency
Real-world users often exhibit highly heterogeneous preferences over multiple objectives for LLM responses. A lightweight aligner can tailor these responses to individual preferences, but scarce user-specific feedback makes personalized training difficult. Learning shared initializations across users can support few-sh …