الملخص
Long-tailed chest X-ray classification requires visual representations that capture both common abnormalities and subtle, infrequent findings. We propose Med-AR-8B and Med-AR-2B, two radiology-native autoregressive vision-language models pretrained with structured reports, abnormality-focused text, and region annotations. We evaluate the transfer of their visual encoders to multi-label classification against contrastive, self-supervised, and supervised pretrained encoders, including Med-CLIP, CheXFound, EVA-Base, ARK, and BioViL-T, using a common ML-Decoder classification head. To assess fine-grained recognition, we also construct LLM-expanded, report-derived label sets for MIMIC-CXR and CheXpert. Across PadChest, MIMIC-CXR, and CheXpert, Med-AR-8B outperforms Med-CLIP in mean AUROC and AUPRC for head, medium, and tail findings. On MIMIC-CXR, it increases tail-label mean AUPRC from 0.1033 to 0.1441. Med-AR-2B achieves the strongest discrimination results on PadChest. Across the broader encoder comparison, a Med-AR variant achieves the highest mean AUROC and AUPRC in every reported prevalence group on each public dataset. Both Med-AR variants also achieve lower excess area under the risk-coverage curve than Med-CLIP on all three public datasets, indicating improved selective-prediction performance under the evaluated protocol. Internal results are metric-dependent, with Med-CLIP retaining advantages in overall and tail AUPRC and in selective prediction. These findings establish Med-AR as a strong pretraining recipe for long-tailed chest X-ray classification on the evaluated public benchmarks and demonstrate the value of assessing discrimination and selective prediction together.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Prabhu, J., Sahil, V, A., Shukla, S., Tadepalli, M., & Putha, P. (2026). Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation. https://omanscience.com/ar/articles/med-ar-autoregressive-vision-language-pretraining-for-long-tailed-chest-x-ray-classification-and-uncertainty-aware-evaluation
MLA 9
Prabhu, Janhavi, et al. "Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation." https://omanscience.com/ar/articles/med-ar-autoregressive-vision-language-pretraining-for-long-tailed-chest-x-ray-classification-and-uncertainty-aware-evaluation.
شيكاغو (المؤلف–التاريخ)
Prabhu, Janhavi, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, and Preetham Putha. 2026. "Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation." https://omanscience.com/ar/articles/med-ar-autoregressive-vision-language-pretraining-for-long-tailed-chest-x-ray-classification-and-uncertainty-aware-evaluation.
هارفارد
Prabhu, J., Sahil, V, A., Shukla, S., Tadepalli, M. and Putha, P. (2026) 'Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation', Available at: https://omanscience.com/ar/articles/med-ar-autoregressive-vision-language-pretraining-for-long-tailed-chest-x-ray-classification-and-uncertainty-aware-evaluation.
فانكوفر
Prabhu J, Sahil, V A, Shukla S, Tadepalli M, Putha P. Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation. https://omanscience.com/ar/articles/med-ar-autoregressive-vision-language-pretraining-for-long-tailed-chest-x-ray-classification-and-uncertainty-aware-evaluation
IEEE
J. Prabhu, Sahil, A. V, S. Shukla, M. Tadepalli, and P. Putha, "Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation," https://omanscience.com/ar/articles/med-ar-autoregressive-vision-language-pretraining-for-long-tailed-chest-x-ray-classification-and-uncertainty-aware-evaluation.