الملخص

The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reasoning. Because these traces may reflect post-hoc rationalization rather than genuine reasoning, we first evaluate against native CoT on open-source models and extend to closed-source frontier models including GPT-6 Astra. We find that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines, across competition mathematics, science, and code generation. We then characterize how frontier models structure their intermediate reasoning. Across token efficiency, reasoning-step types, and induced reasoning trees, we identify systematic differences in how models externalize, compress, and organize reasoning. We find that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning. These findings provide a behavioral lens on frontier-model reasoning beyond benchmark scores.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Luo, X., Ren, T., Yu, W., Li, X., Li, Q., & Bjerva, J. (2026). Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models. https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models

MLA 9

Luo, Xiaoyu, et al. "Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models." https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models.

شيكاغو (المؤلف–التاريخ)

Luo, Xiaoyu, Tao Ren, Wenrui Yu, Xiao Li, Qiongxiu Li, and Johannes Bjerva. 2026. "Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models." https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models.

هارفارد

Luo, X., Ren, T., Yu, W., Li, X., Li, Q. and Bjerva, J. (2026) 'Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models', Available at: https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models.

فانكوفر

Luo X, Ren T, Yu W, Li X, Li Q, Bjerva J. Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models. https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models

IEEE

X. Luo, T. Ren, W. Yu, X. Li, Q. Li, and J. Bjerva, "Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models," https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models.