Abstract
Large language models (LLMs) have been rapidly improving in long-context tasks, powered by Chain-of-Thought (CoT) reasoning. However, the internal mechanisms underlying this improvement remain unclear. We investigate these mechanisms through a needle-in-a-haystack (NIAH) counting task, where an LLM is asked to count the number of records dispersed in a long text. Across twelve model comparison groups, Thinking (or reasoning) improves counting accuracy over Non-thinking, with pronounced gains at larger counts. This motivates our mechanistic analysis, which identifies two contrasting mechanisms: (i) broad retrieval, where Non-thinking models broadly attend to multiple needles; (ii) targeted retrieval, where Thinking models use enumeration in CoT traces to successively retrieve needles. Targeted retrieval concentrates attention on individual needles and is accompanied by more compact internal representations. Moreover, causal intervention analysis suggests that Thinking models use the CoT trace to maintain and update an internal counter as needles are successively retrieved, even without explicit numbering. In small controlled experiments, both retrieval mechanisms and counter states emerge under standard autoregressive training. Together, our results connect long-context retrieval with representation geometry of counting, supporting a state-tracking account of CoT reasoning.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Shan, L. T., Hu, T., Yan, H., & Zhong, Y. (2026). Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting. https://omanscience.com/en/articles/targeted-retrieval-compact-representations-how-cot-reasoning-improves-long-context-counting
MLA 9
Shan, Liang Twist, et al. "Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting." https://omanscience.com/en/articles/targeted-retrieval-compact-representations-how-cot-reasoning-improves-long-context-counting.
Chicago (author–date)
Shan, Liang Twist, Tianyu Hu, Hao Yan, and Yiqiao Zhong. 2026. "Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting." https://omanscience.com/en/articles/targeted-retrieval-compact-representations-how-cot-reasoning-improves-long-context-counting.
Harvard
Shan, L. T., Hu, T., Yan, H. and Zhong, Y. (2026) 'Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting', Available at: https://omanscience.com/en/articles/targeted-retrieval-compact-representations-how-cot-reasoning-improves-long-context-counting.
Vancouver
Shan LT, Hu T, Yan H, Zhong Y. Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting. https://omanscience.com/en/articles/targeted-retrieval-compact-representations-how-cot-reasoning-improves-long-context-counting
IEEE
L. T. Shan, T. Hu, H. Yan, and Y. Zhong, "Targeted Retrieval, Compact Representations: How CoT Reasoning Improves Long-Context Counting," https://omanscience.com/en/articles/targeted-retrieval-compact-representations-how-cot-reasoning-improves-long-context-counting.