Abstract
GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for grounded multi-view VQA. Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge. On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75\% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92\% Answer Accuracy and 86.44\% View Accuracy. The final submitted run achieves 88.14\% Joint Accuracy after two additional evaluator-informed post-hoc corrections. Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wang, K., Hu, Y., Cao, R., Liu, H., Li, Z., Xiang, Q., & Cheng, H. (2026). GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA. https://omanscience.com/en/articles/groundsight-at-groundlm-2026-shared-tasks-goldenviewvqa
MLA 9
Wang, Kun, et al. "GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA." https://omanscience.com/en/articles/groundsight-at-groundlm-2026-shared-tasks-goldenviewvqa.
Chicago (author–date)
Wang, Kun, Yupeng Hu, Ruping Cao, Hao Liu, Zhiran Li, Qianlong Xiang, and Harry Cheng. 2026. "GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA." https://omanscience.com/en/articles/groundsight-at-groundlm-2026-shared-tasks-goldenviewvqa.
Harvard
Wang, K., Hu, Y., Cao, R., Liu, H., Li, Z., Xiang, Q. and Cheng, H. (2026) 'GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA', Available at: https://omanscience.com/en/articles/groundsight-at-groundlm-2026-shared-tasks-goldenviewvqa.
Vancouver
Wang K, Hu Y, Cao R, Liu H, Li Z, Xiang Q, et al. GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA. https://omanscience.com/en/articles/groundsight-at-groundlm-2026-shared-tasks-goldenviewvqa
IEEE
K. Wang, Y. Hu, R. Cao, H. Liu, Z. Li, Q. Xiang, and H. Cheng, "GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA," https://omanscience.com/en/articles/groundsight-at-groundlm-2026-shared-tasks-goldenviewvqa.