Abstract

Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires current diagnostic criteria, standardized acquisition and reporting knowledge, longitudinal monitoring concepts, lesion morphology, and recognition of difficult mimics. We present MS-Exam-Gen, a reproducible framework for constructing and auditing a text-based multiple-choice question (MCQ) benchmark for MS-MRI knowledge; it does not evaluate direct MRI image interpretation. MS-Exam-Gen targets source-grounded criteria, protocols, reporting, and differential diagnosis. The framework combines expert-source indexing, exam-oriented topic induction, evidence-grounded MCQ generation, automated quality audits, a same-family consistency screen, and empirical calibration. From a 66-source corpus indexed into 4,289 retrieval chunks, the pipeline produced a locked 3,058-item candidate benchmark spanning 16 topics and 53 subtopics. Evaluation across 12 primary LLM endpoints yielded 36,696 item-level predictions and separated performance over a 42.8-percentage-point accuracy range (89.7% to 46.9%). Across these endpoints, 25.5% of items were missed by at least four. Post-generation audits showed that refreshed construction reduced measurable answer cues, while option-order testing showed that absolute MCQ scores remain position-sensitive. Generated construction labels remain metadata rather than validated psychometric categories. Because expert adjudication and full option-order counterbalancing remain future work, MS-Exam-Gen is not a clinically certified examination. It should be interpreted as an automatically filtered, source-grounded candidate benchmark and reproducible audit workflow for item-level and topic-specific LLM evaluation.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Basit, A., Hanif, M. A., & Shafique, M. (2026). MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge. https://omanscience.com/en/articles/ms-exam-gen-source-grounded-benchmark-construction-for-evaluating-llms-on-textual-multiple-sclerosis-mri-knowledge

MLA 9

Basit, Abdul, et al. "MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge." https://omanscience.com/en/articles/ms-exam-gen-source-grounded-benchmark-construction-for-evaluating-llms-on-textual-multiple-sclerosis-mri-knowledge.

Chicago (author–date)

Basit, Abdul, Muhammad Abdullah Hanif, and Muhammad Shafique. 2026. "MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge." https://omanscience.com/en/articles/ms-exam-gen-source-grounded-benchmark-construction-for-evaluating-llms-on-textual-multiple-sclerosis-mri-knowledge.

Harvard

Basit, A., Hanif, M. A. and Shafique, M. (2026) 'MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge', Available at: https://omanscience.com/en/articles/ms-exam-gen-source-grounded-benchmark-construction-for-evaluating-llms-on-textual-multiple-sclerosis-mri-knowledge.

Vancouver

Basit A, Hanif MA, Shafique M. MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge. https://omanscience.com/en/articles/ms-exam-gen-source-grounded-benchmark-construction-for-evaluating-llms-on-textual-multiple-sclerosis-mri-knowledge

IEEE

A. Basit, M. A. Hanif, and M. Shafique, "MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge," https://omanscience.com/en/articles/ms-exam-gen-source-grounded-benchmark-construction-for-evaluating-llms-on-textual-multiple-sclerosis-mri-knowledge.