Abstract
Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluate whether VLMs can move beyond disease recognition toward comparative reasoning and fine-grained spatial grounding. We introduce EyeVQA, a unified visual question answering benchmark for comprehensive evaluation of ophthalmic VLMs. EyeVQA is constructed from 21 available ophthalmic datasets and contains 20,000 question-answer pairs spanning six disease groups and seven question types: Single-Choice, Multi-Select, Variable-Select, True-False, Ranking, Point Location, and Bounding Box. Gold answers are deterministically derived from source-provided diagnoses, severity grades, clinical findings, segmentation masks, bounding boxes, and anatomical landmarks, enabling reproducible evaluation without relying on model-generated annotations. Notably, 44.5% of the questions require reasoning across multiple images, extending evaluation beyond conventional single-image medical VQA. We benchmark fourteen representative general-purpose, scientific, and medically specialized VLMs under a unified zero-shot protocol. The best-performing model only achieves an overall score of 62.8, while substantial gaps remain in spatial grounding and cross-task generalization. These results highlight the limitations of current VLMs in comprehensive ophthalmic visual understanding and establish EyeVQA as a diagnostic benchmark for developing more reliable and spatially grounded ophthalmic multimodal models. The project page is available at https://github.com/PKUTHM/EyeVQA.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Shao, G., Xie, Z., Xing, X., Wang, R., Lan, Z., Qi, Y., Zhang, G., Yang, Y., Li, D., & Tang, H. (2026). EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding. https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding
MLA 9
Shao, Gujie, et al. "EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding." https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding.
Chicago (author–date)
Shao, Gujie, Zixun Xie, Xuechun Xing, Ruixiang Wang, Ziyun Lan, Yanlin Qi, Gangyi Zhang, Yuxin Yang, Dawei Li, and Haiming Tang. 2026. "EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding." https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding.
Harvard
Shao, G., Xie, Z., Xing, X., Wang, R., Lan, Z., Qi, Y., Zhang, G., Yang, Y., Li, D. and Tang, H. (2026) 'EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding', Available at: https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding.
Vancouver
Shao G, Xie Z, Xing X, Wang R, Lan Z, Qi Y, et al. EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding. https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding
IEEE
G. Shao, Z. Xie, X. Xing, R. Wang, Z. Lan, Y. Qi, G. Zhang, Y. Yang, D. Li, and H. Tang, "EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding," https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding.