نسخة أولية وصول مفتوح
When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models
Vision-language models (VLMs) have achieved strong performance in multimodal reasoning, yet they remain prone to generating plausible but incorrect answers. Self-verification offers a practical way to improve answer reliability without relying on external judges, but existing methods typically depend on a single verifi …