الملخص
The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reasoning. Because these traces may reflect post-hoc rationalization rather than genuine reasoning, we first evaluate against native CoT on open-source models and extend to closed-source frontier models including GPT-6 Astra. We find that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines, across competition mathematics, science, and code generation. We then characterize how frontier models structure their intermediate reasoning. Across token efficiency, reasoning-step types, and induced reasoning trees, we identify systematic differences in how models externalize, compress, and organize reasoning. We find that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning. These findings provide a behavioral lens on frontier-model reasoning beyond benchmark scores.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Luo, X., Ren, T., Yu, W., Li, X., Li, Q., & Bjerva, J. (2026). Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models. https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models
MLA 9
Luo, Xiaoyu, et al. "Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models." https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models.
شيكاغو (المؤلف–التاريخ)
Luo, Xiaoyu, Tao Ren, Wenrui Yu, Xiao Li, Qiongxiu Li, and Johannes Bjerva. 2026. "Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models." https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models.
هارفارد
Luo, X., Ren, T., Yu, W., Li, X., Li, Q. and Bjerva, J. (2026) 'Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models', Available at: https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models.
فانكوفر
Luo X, Ren T, Yu W, Li X, Li Q, Bjerva J. Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models. https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models
IEEE
X. Luo, T. Ren, W. Yu, X. Li, Q. Li, and J. Bjerva, "Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models," https://omanscience.com/ar/articles/capable-yet-parsimonious-extracting-and-characterizing-hidden-chain-of-thought-in-frontier-models.