Abstract

Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large language model's (LLM) CoT reasoning engage the same internal concepts that support the LLM's direct prediction, and do the shared concepts causally drive its answer? Encoding a prediction pass and a CoT pass with a single shared sparse autoencoder (SAE), a reliable approximator of the latent concepts LLMs use, makes their internal concepts directly comparable. We introduce three correlational metrics of concept-level alignment and a causal metric, $Δp$, which ablates the shared concepts and measures the drop in answer probability. Across five LLMs and four datasets, concept alignment is generally high, as indicated by the correlational metrics; yet these only identify which concepts are shared, not how much they causally contribute. $Δp$ fills this gap: causal faithfulness varies substantially with model depth, peaking at mid-to-late layers rather than the final ones, and model scale reshapes the layer-wise profile. Moreover, causally important shared concepts are not always verbalized in the CoT. These dissociations suggest that faithfulness cannot be reliably assessed from surface-level or representational correspondence alone; assessing it requires causal tests of whether the internal concepts underlying a CoT actually drive the model's prediction.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, Q., Wang, Y., Wei, D., Sun, J., Ostermann, S., Augenstein, I., Atanasova, P., & Feldhus, N. (2026). From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness. https://omanscience.com/en/articles/from-concept-alignment-to-causal-grounding-an-intervention-test-of-chain-of-thought-faithfulness

MLA 9

Wang, Qianli, et al. "From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness." https://omanscience.com/en/articles/from-concept-alignment-to-causal-grounding-an-intervention-test-of-chain-of-thought-faithfulness.

Chicago (author–date)

Wang, Qianli, Yilong Wang, Dennis Wei, Jingyi Sun, Simon Ostermann, Isabelle Augenstein, Pepa Atanasova, and Nils Feldhus. 2026. "From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness." https://omanscience.com/en/articles/from-concept-alignment-to-causal-grounding-an-intervention-test-of-chain-of-thought-faithfulness.

Harvard

Wang, Q., Wang, Y., Wei, D., Sun, J., Ostermann, S., Augenstein, I., Atanasova, P. and Feldhus, N. (2026) 'From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness', Available at: https://omanscience.com/en/articles/from-concept-alignment-to-causal-grounding-an-intervention-test-of-chain-of-thought-faithfulness.

Vancouver

Wang Q, Wang Y, Wei D, Sun J, Ostermann S, Augenstein I, et al. From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness. https://omanscience.com/en/articles/from-concept-alignment-to-causal-grounding-an-intervention-test-of-chain-of-thought-faithfulness

IEEE

Q. Wang, Y. Wang, D. Wei, J. Sun, S. Ostermann, I. Augenstein, P. Atanasova, and N. Feldhus, "From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness," https://omanscience.com/en/articles/from-concept-alignment-to-causal-grounding-an-intervention-test-of-chain-of-thought-faithfulness.