Abstract
Vision-language models (VLMs) have achieved strong performance in multimodal reasoning, yet they remain prone to generating plausible but incorrect answers. Self-verification offers a practical way to improve answer reliability without relying on external judges, but existing methods typically depend on a single verification criterion or fixed prompt, resulting in incomplete and unstable reliability estimates. We first systematically analyze how verifier capability and prompt design affect verification performance. Our findings show that stronger verifiers provide more reliable judgments, while verification performance is highly sensitive to prompt choice, with no single prompt consistently dominating across tasks. Guided by these findings, we propose \texttt{MOTIVE}, a \textbf{M}ulti-View Self-Verificati\textbf{O}n wi\textbf{T}h Rel\textbf{I}ability-Guided Selecti\textbf{VE} Rethinking framework for reliable multimodal reasoning. \texttt{MOTIVE} evaluates each candidate answer from complementary verification perspectives and learns a correctness-aligned reliability score through correctness-grounded multi-view verification learning. During inference, this score governs an accept-or-rethink decision, allowing reliable answers to be returned directly while uncertain ones trigger history-guided rethinking. Extensive experiments across diverse multimodal benchmarks and VLM backbones demonstrate that \texttt{MOTIVE} consistently outperforms strong self-verification and self-correction baselines. Further results show that reliable verification improves accept-or-rethink decisions and reduces unnecessary reasoning turns, enabling more reliable and efficient self-verification without an external judge.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhu, Z., Zhu, H., Lu, S. Y., Huang, M. Y. C., Lin, Y., Han, W., Chen, T., Wu, M., Yu, H., Jin, G., Liu, L., Sun, B., & Huang, T. (2026). When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models. https://omanscience.com/en/articles/when-to-rethink-learning-multi-perspective-self-verification-for-vision-language-models
MLA 9
Zhu, Ziquan, et al. "When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models." https://omanscience.com/en/articles/when-to-rethink-learning-multi-perspective-self-verification-for-vision-language-models.
Chicago (author–date)
Zhu, Ziquan, Hanruo Zhu, Si-Yuan Lu, Morris Yu-Chao Huang, Yicheng Lin, Wei Han, Tianlong Chen, Mingyuan Wu, Hanchao Yu, Gaojie Jin, Lu Liu, Bo Sun, and Tianjin Huang. 2026. "When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models." https://omanscience.com/en/articles/when-to-rethink-learning-multi-perspective-self-verification-for-vision-language-models.
Harvard
Zhu, Z., Zhu, H., Lu, S. Y., Huang, M. Y. C., Lin, Y., Han, W., Chen, T., Wu, M., Yu, H., Jin, G., Liu, L., Sun, B. and Huang, T. (2026) 'When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models', Available at: https://omanscience.com/en/articles/when-to-rethink-learning-multi-perspective-self-verification-for-vision-language-models.
Vancouver
Zhu Z, Zhu H, Lu SY, Huang MYC, Lin Y, Han W, et al. When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models. https://omanscience.com/en/articles/when-to-rethink-learning-multi-perspective-self-verification-for-vision-language-models
IEEE
Z. Zhu, H. Zhu, S. Y. Lu, M. Y. C. Huang, Y. Lin, W. Han, T. Chen, M. Wu, H. Yu, G. Jin, L. Liu, B. Sun, and T. Huang, "When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models," https://omanscience.com/en/articles/when-to-rethink-learning-multi-perspective-self-verification-for-vision-language-models.