Abstract

Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information density, and their utility shifts as training progresses from broad knowledge acquisition to late-stage consolidation. Meanwhile, post-training is often dominated by short-form visual question answering, providing limited supervision for informative and answer-consistent explanations. We introduce InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design. We curate a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing. Our CPT pipeline first adapts the vision encoder, then builds broad medical knowledge, and finally transitions to an evidence-focused data mixture during learning-rate decay. For supervised fine-tuning (SFT), we regenerate visual question-answering responses using answer stability, answer-masked reconstruction, and correctness-constrained selection to produce more informative and answer-consistent supervision. The 4B model is further optimized with reinforcement learning with verifiable rewards (RLVR). Across five medical multimodal benchmarks, InfiMed2-4B achieves 66.73% mean accuracy after RLVR, surpassing the larger Qwen3.5-9B, while InfiMed2-27B reaches 73.72%, the highest among the evaluated open-weight models.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhu, G., Liu, Z., Hou, Z., Wang, P., Sang, Z., Cai, S., Yu, Y., Wang, Y., Gu, Y., Xie, C., Wu, J., & Yang, H. (2026). InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision. https://omanscience.com/en/articles/infimed2-a-generalist-medical-multimodal-foundation-model-from-contextual-evidence-and-stability-aware-supervision

MLA 9

Zhu, Guanghao, et al. "InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision." https://omanscience.com/en/articles/infimed2-a-generalist-medical-multimodal-foundation-model-from-contextual-evidence-and-stability-aware-supervision.

Chicago (author–date)

Zhu, Guanghao, Zeyu Liu, Zhitian Hou, Pengkai Wang, Zhijie Sang, Shuo Cai, Yang Yu, Yuanyi Wang, Yanggan Gu, Congkai Xie, Jianmin Wu, and Hongxia Yang. 2026. "InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision." https://omanscience.com/en/articles/infimed2-a-generalist-medical-multimodal-foundation-model-from-contextual-evidence-and-stability-aware-supervision.

Harvard

Zhu, G., Liu, Z., Hou, Z., Wang, P., Sang, Z., Cai, S., Yu, Y., Wang, Y., Gu, Y., Xie, C., Wu, J. and Yang, H. (2026) 'InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision', Available at: https://omanscience.com/en/articles/infimed2-a-generalist-medical-multimodal-foundation-model-from-contextual-evidence-and-stability-aware-supervision.

Vancouver

Zhu G, Liu Z, Hou Z, Wang P, Sang Z, Cai S, et al. InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision. https://omanscience.com/en/articles/infimed2-a-generalist-medical-multimodal-foundation-model-from-contextual-evidence-and-stability-aware-supervision

IEEE

G. Zhu, Z. Liu, Z. Hou, P. Wang, Z. Sang, S. Cai, Y. Yu, Y. Wang, Y. Gu, C. Xie, J. Wu, and H. Yang, "InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision," https://omanscience.com/en/articles/infimed2-a-generalist-medical-multimodal-foundation-model-from-contextual-evidence-and-stability-aware-supervision.