الملخص
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Shao, T. (2026). Language-model groups overstate consensus when replaying human deliberation on a reasoning task. https://omanscience.com/ar/articles/language-model-groups-overstate-consensus-when-replaying-human-deliberation-on-a-reasoning-task
MLA 9
Shao, Tengfei. "Language-model groups overstate consensus when replaying human deliberation on a reasoning task." https://omanscience.com/ar/articles/language-model-groups-overstate-consensus-when-replaying-human-deliberation-on-a-reasoning-task.
شيكاغو (المؤلف–التاريخ)
Shao, Tengfei. 2026. "Language-model groups overstate consensus when replaying human deliberation on a reasoning task." https://omanscience.com/ar/articles/language-model-groups-overstate-consensus-when-replaying-human-deliberation-on-a-reasoning-task.
هارفارد
Shao, T. (2026) 'Language-model groups overstate consensus when replaying human deliberation on a reasoning task', Available at: https://omanscience.com/ar/articles/language-model-groups-overstate-consensus-when-replaying-human-deliberation-on-a-reasoning-task.
فانكوفر
Shao T. Language-model groups overstate consensus when replaying human deliberation on a reasoning task. https://omanscience.com/ar/articles/language-model-groups-overstate-consensus-when-replaying-human-deliberation-on-a-reasoning-task
IEEE
T. Shao, "Language-model groups overstate consensus when replaying human deliberation on a reasoning task," https://omanscience.com/ar/articles/language-model-groups-overstate-consensus-when-replaying-human-deliberation-on-a-reasoning-task.