Abstract

Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluate whether VLMs can move beyond disease recognition toward comparative reasoning and fine-grained spatial grounding. We introduce EyeVQA, a unified visual question answering benchmark for comprehensive evaluation of ophthalmic VLMs. EyeVQA is constructed from 21 available ophthalmic datasets and contains 20,000 question-answer pairs spanning six disease groups and seven question types: Single-Choice, Multi-Select, Variable-Select, True-False, Ranking, Point Location, and Bounding Box. Gold answers are deterministically derived from source-provided diagnoses, severity grades, clinical findings, segmentation masks, bounding boxes, and anatomical landmarks, enabling reproducible evaluation without relying on model-generated annotations. Notably, 44.5% of the questions require reasoning across multiple images, extending evaluation beyond conventional single-image medical VQA. We benchmark fourteen representative general-purpose, scientific, and medically specialized VLMs under a unified zero-shot protocol. The best-performing model only achieves an overall score of 62.8, while substantial gaps remain in spatial grounding and cross-task generalization. These results highlight the limitations of current VLMs in comprehensive ophthalmic visual understanding and establish EyeVQA as a diagnostic benchmark for developing more reliable and spatially grounded ophthalmic multimodal models. The project page is available at https://github.com/PKUTHM/EyeVQA.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Shao, G., Xie, Z., Xing, X., Wang, R., Lan, Z., Qi, Y., Zhang, G., Yang, Y., Li, D., & Tang, H. (2026). EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding. https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding

MLA 9

Shao, Gujie, et al. "EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding." https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding.

Chicago (author–date)

Shao, Gujie, Zixun Xie, Xuechun Xing, Ruixiang Wang, Ziyun Lan, Yanlin Qi, Gangyi Zhang, Yuxin Yang, Dawei Li, and Haiming Tang. 2026. "EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding." https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding.

Harvard

Shao, G., Xie, Z., Xing, X., Wang, R., Lan, Z., Qi, Y., Zhang, G., Yang, Y., Li, D. and Tang, H. (2026) 'EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding', Available at: https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding.

Vancouver

Shao G, Xie Z, Xing X, Wang R, Lan Z, Qi Y, et al. EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding. https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding

IEEE

G. Shao, Z. Xie, X. Xing, R. Wang, Z. Lan, Y. Qi, G. Zhang, Y. Yang, D. Li, and H. Tang, "EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding," https://omanscience.com/en/articles/eyevqa-benchmarking-ophthalmic-vision-language-models-from-recognition-to-spatial-grounding.