Abstract

Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves query-aligned demonstrations using frozen encoders, without task-specific parameter updates. Across eight datasets spanning histopathology, radiology and retinal fundoscopy, RG-ICL improved classification (mean gain 20 percentage points) and visual question answering (VQA) (mean gain 13 percentage points) over no-context and conventional ICL, approaching or exceeding training-based comparators. Which cases were retrieved mattered more than how many: 6 query-aligned cases outperformed up to 32 randomly selected ones, whereas fixed or random cases often reduced accuracy below baseline. For VQA, aligning reference cases with both image content and question intent produced further gains. These findings indicate that for medical image interpretation, curating which reference cases an MLLM sees is a practical alternative to retraining it.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhao, M., Hu, F., Luo, Y., Yang, Y., Cai, J., Zhou, K., Li, M., Liang, P., Du, Y., Shen, L. Q., & Wang, M. (2026). Representation-guided in-context learning for medical image interpretation with multimodal large language models. https://omanscience.com/en/articles/representation-guided-in-context-learning-for-medical-image-interpretation-with-multimodal-large-language-models

MLA 9

Zhao, Minda, et al. "Representation-guided in-context learning for medical image interpretation with multimodal large language models." https://omanscience.com/en/articles/representation-guided-in-context-learning-for-medical-image-interpretation-with-multimodal-large-language-models.

Chicago (author–date)

Zhao, Minda, Fangyu Hu, Yan Luo, Yutong Yang, Jiahui Cai, Kaichen Zhou, Manling Li, Paul Liang, Yilun Du, Lucy Q. Shen, and Mengyu Wang. 2026. "Representation-guided in-context learning for medical image interpretation with multimodal large language models." https://omanscience.com/en/articles/representation-guided-in-context-learning-for-medical-image-interpretation-with-multimodal-large-language-models.

Harvard

Zhao, M., Hu, F., Luo, Y., Yang, Y., Cai, J., Zhou, K., Li, M., Liang, P., Du, Y., Shen, L. Q. and Wang, M. (2026) 'Representation-guided in-context learning for medical image interpretation with multimodal large language models', Available at: https://omanscience.com/en/articles/representation-guided-in-context-learning-for-medical-image-interpretation-with-multimodal-large-language-models.

Vancouver

Zhao M, Hu F, Luo Y, Yang Y, Cai J, Zhou K, et al. Representation-guided in-context learning for medical image interpretation with multimodal large language models. https://omanscience.com/en/articles/representation-guided-in-context-learning-for-medical-image-interpretation-with-multimodal-large-language-models

IEEE

M. Zhao, F. Hu, Y. Luo, Y. Yang, J. Cai, K. Zhou, M. Li, P. Liang, Y. Du, L. Q. Shen, and M. Wang, "Representation-guided in-context learning for medical image interpretation with multimodal large language models," https://omanscience.com/en/articles/representation-guided-in-context-learning-for-medical-image-interpretation-with-multimodal-large-language-models.