Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at https://github.com/Zhihan72/TPAE.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhang, Z., & Liao, L. (2026). Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation. https://omanscience.com/en/articles/reinforcing-multimodal-reasoning-via-token-level-perception-grounded-advantage-estimation
MLA 9
Zhang, Zhihan, and Lizi Liao. "Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation." https://omanscience.com/en/articles/reinforcing-multimodal-reasoning-via-token-level-perception-grounded-advantage-estimation.
Chicago (author–date)
Zhang, Zhihan, and Lizi Liao. 2026. "Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation." https://omanscience.com/en/articles/reinforcing-multimodal-reasoning-via-token-level-perception-grounded-advantage-estimation.
Harvard
Zhang, Z. and Liao, L. (2026) 'Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation', Available at: https://omanscience.com/en/articles/reinforcing-multimodal-reasoning-via-token-level-perception-grounded-advantage-estimation.
Vancouver
Zhang Z, Liao L. Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation. https://omanscience.com/en/articles/reinforcing-multimodal-reasoning-via-token-level-perception-grounded-advantage-estimation
IEEE
Z. Zhang, and L. Liao, "Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation," https://omanscience.com/en/articles/reinforcing-multimodal-reasoning-via-token-level-perception-grounded-advantage-estimation.