الملخص

We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text. We train on a curated corpus of approximately 550,000 multimodal instruction examples from 70,111 unique CT image inputs, combining standardized reports, abnormality-focused and anatomy-specific questions, multi-turn interactions, and radiologist-authored reasoning from recorded and transcribed expert CT interpretations. Expert annotations provide direct supervision and guide additional report-grounded synthetic reasoning. End-to-end supervised fine-tuning (SFT) is followed by Group Relative Policy Optimization (GRPO), with verifiable rewards over chest and abdominal abnormality sets. The model supports abnormality classification, report generation, and interactive reasoning with reviewable observations, differential diagnoses, and uncertainty. Evaluation spans public CT benchmarks and a held-out NIH cohort. On CT-RATE, NV-Reason-CT achieves a macro-F1 of 0.614 and macro-AUROC of 0.871 without a task-specific classification head; generated reports achieve a report-derived macro-F1 of 0.592. In a preliminary study with expert radiologists, AI-assisted review received favorable confidence ratings and was associated with a 50% reduction in average reported interpretation and reporting time. We release the model and training code to support reproducible research on explainable AI for volumetric medical imaging.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Myronenko, A., Yang, D., Tang, Y., Turkbey, B., Simon, B., Harmon, S., Makwana, R., Aboian, M., Azamat, S., Hamamci, I. E., Er, S., Menze, B., Zhou, Z., Li, W., Edgar, M., He, Y., Guo, P., & Xu, D. (2026). NV-Reason-CT: 3D Visual Language Model for CT Analysis. https://omanscience.com/ar/articles/nv-reason-ct-3d-visual-language-model-for-ct-analysis

MLA 9

Myronenko, Andriy, et al. "NV-Reason-CT: 3D Visual Language Model for CT Analysis." https://omanscience.com/ar/articles/nv-reason-ct-3d-visual-language-model-for-ct-analysis.

شيكاغو (المؤلف–التاريخ)

Myronenko, Andriy, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Zongwei Zhou, Wenxuan Li, Marc Edgar, Yufan He, Pengfei Guo, and Daguang Xu. 2026. "NV-Reason-CT: 3D Visual Language Model for CT Analysis." https://omanscience.com/ar/articles/nv-reason-ct-3d-visual-language-model-for-ct-analysis.

هارفارد

Myronenko, A., Yang, D., Tang, Y., Turkbey, B., Simon, B., Harmon, S., Makwana, R., Aboian, M., Azamat, S., Hamamci, I. E., Er, S., Menze, B., Zhou, Z., Li, W., Edgar, M., He, Y., Guo, P. and Xu, D. (2026) 'NV-Reason-CT: 3D Visual Language Model for CT Analysis', Available at: https://omanscience.com/ar/articles/nv-reason-ct-3d-visual-language-model-for-ct-analysis.

فانكوفر

Myronenko A, Yang D, Tang Y, Turkbey B, Simon B, Harmon S, et al. NV-Reason-CT: 3D Visual Language Model for CT Analysis. https://omanscience.com/ar/articles/nv-reason-ct-3d-visual-language-model-for-ct-analysis

IEEE

A. Myronenko, D. Yang, Y. Tang, B. Turkbey, B. Simon, S. Harmon, R. Makwana, M. Aboian, S. Azamat, I. E. Hamamci, S. Er, B. Menze, Z. Zhou, W. Li, M. Edgar, Y. He, P. Guo, and D. Xu, "NV-Reason-CT: 3D Visual Language Model for CT Analysis," https://omanscience.com/ar/articles/nv-reason-ct-3d-visual-language-model-for-ct-analysis.