الملخص
Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evidence, while others correspond to uncertain reasoning decisions. We present Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware. Starting from an initial generated trajectory, we optimize a short hidden-state prefix while routing perception-side visual feedback to image-sensitive tokens and reasoning feedback to high-entropy tokens. Tokens selected by neither route are constrained by an anchor regularizer. Across both perception and reasoning benchmarks on Qwen2.5-VL-7B and InternVL3.5-8B, our method lifts macro accuracy over CoT by +2.57 and +1.51 respectively, and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets. Code is available at https://github.com/Qwen-Applications/TD-LTTS.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Ma, H. X., Liu, Y., Sun, Y., Miao, Y., Zhou, M., Xiao, Y., Chen, L., Li, Z., Ye, H. J., Jiang, X., & Jiang, G. (2026). Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning. https://omanscience.com/ar/articles/token-disentangled-latent-test-time-scaling-for-vision-language-reasoning
MLA 9
Ma, Hao-Xuan, et al. "Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning." https://omanscience.com/ar/articles/token-disentangled-latent-test-time-scaling-for-vision-language-reasoning.
شيكاغو (المؤلف–التاريخ)
Ma, Hao-Xuan, Yihao Liu, Yutao Sun, Yanting Miao, Mengyu Zhou, Yicheng Xiao, Long Chen, Zhenguo Li, Han-Jia Ye, Xiaoxi Jiang, and Guanjun Jiang. 2026. "Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning." https://omanscience.com/ar/articles/token-disentangled-latent-test-time-scaling-for-vision-language-reasoning.
هارفارد
Ma, H. X., Liu, Y., Sun, Y., Miao, Y., Zhou, M., Xiao, Y., Chen, L., Li, Z., Ye, H. J., Jiang, X. and Jiang, G. (2026) 'Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning', Available at: https://omanscience.com/ar/articles/token-disentangled-latent-test-time-scaling-for-vision-language-reasoning.
فانكوفر
Ma HX, Liu Y, Sun Y, Miao Y, Zhou M, Xiao Y, et al. Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning. https://omanscience.com/ar/articles/token-disentangled-latent-test-time-scaling-for-vision-language-reasoning
IEEE
H. X. Ma, Y. Liu, Y. Sun, Y. Miao, M. Zhou, Y. Xiao, L. Chen, Z. Li, H. J. Ye, X. Jiang, and G. Jiang, "Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning," https://omanscience.com/ar/articles/token-disentangled-latent-test-time-scaling-for-vision-language-reasoning.