الباحثون

Pepa Atanasova

المنشورات 1

نسخة أولية وصول مفتوح

From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness

Qianli Wang, Yilong Wang, Dennis Wei وآخرون · 2026

Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large lan …

المؤلفون المشاركون