الملخص

Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show that alignment with human preferences can make model behavior less human-like even when both preferences and responses come entirely from humans. We call this the Turing-test gap. We show that preference alignment preserves the human response distribution only under a restrictive condition, and find no consistent evidence that real human preferences satisfy it. Empirically, the loss of human-response likelihood increases with the strength of preference weighting, regardless of its direction, and the gap also appears under standard DPO. These results establish human-likeness as an explicit dimension of alignment rather than something assumed to follow from preference alignment.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Yuan, S., Lin, R., Li, M., Hong, G., Gu, J., Feng, L., Russell, C., & Liu, T. (2026). Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment. https://omanscience.com/ar/articles/are-human-aligned-models-models-of-humans-a-turing-test-gap-in-preference-alignment

MLA 9

Yuan, Suqin, et al. "Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment." https://omanscience.com/ar/articles/are-human-aligned-models-models-of-humans-a-turing-test-gap-in-preference-alignment.

شيكاغو (المؤلف–التاريخ)

Yuan, Suqin, Runqi Lin, Muyang Li, Guanzhe Hong, Jindong Gu, Lei Feng, Chris Russell, and Tongliang Liu. 2026. "Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment." https://omanscience.com/ar/articles/are-human-aligned-models-models-of-humans-a-turing-test-gap-in-preference-alignment.

هارفارد

Yuan, S., Lin, R., Li, M., Hong, G., Gu, J., Feng, L., Russell, C. and Liu, T. (2026) 'Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment', Available at: https://omanscience.com/ar/articles/are-human-aligned-models-models-of-humans-a-turing-test-gap-in-preference-alignment.

فانكوفر

Yuan S, Lin R, Li M, Hong G, Gu J, Feng L, et al. Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment. https://omanscience.com/ar/articles/are-human-aligned-models-models-of-humans-a-turing-test-gap-in-preference-alignment

IEEE

S. Yuan, R. Lin, M. Li, G. Hong, J. Gu, L. Feng, C. Russell, and T. Liu, "Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment," https://omanscience.com/ar/articles/are-human-aligned-models-models-of-humans-a-turing-test-gap-in-preference-alignment.