الباحثون

Eunsol Choi

المنشورات 1

نسخة أولية وصول مفتوح

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

Yihuai Hong, Shauli Ravfogel, Chen Zhao وآخرون · 2026

Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignm …

المؤلفون المشاركون