Preprint Open access
Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show t …