نسخة أولية وصول مفتوح
Measuring Collapse and Correction in Homogeneous-Panel LLM Debate
Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We intro …