الملخص

Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus's official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Vangala, B. P., Guda, S., Peddi, L., & Vangala, N. (2026). Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification. https://omanscience.com/ar/articles/scores-that-hold-benchmarks-that-leak-measuring-dataset-contamination-in-public-brain-tumor-mri-classification

MLA 9

Vangala, Bhanu Prakash, et al. "Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification." https://omanscience.com/ar/articles/scores-that-hold-benchmarks-that-leak-measuring-dataset-contamination-in-public-brain-tumor-mri-classification.

شيكاغو (المؤلف–التاريخ)

Vangala, Bhanu Prakash, Sowmya Guda, Latha Peddi, and Navya Vangala. 2026. "Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification." https://omanscience.com/ar/articles/scores-that-hold-benchmarks-that-leak-measuring-dataset-contamination-in-public-brain-tumor-mri-classification.

هارفارد

Vangala, B. P., Guda, S., Peddi, L. and Vangala, N. (2026) 'Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification', Available at: https://omanscience.com/ar/articles/scores-that-hold-benchmarks-that-leak-measuring-dataset-contamination-in-public-brain-tumor-mri-classification.

فانكوفر

Vangala BP, Guda S, Peddi L, Vangala N. Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification. https://omanscience.com/ar/articles/scores-that-hold-benchmarks-that-leak-measuring-dataset-contamination-in-public-brain-tumor-mri-classification

IEEE

B. P. Vangala, S. Guda, L. Peddi, and N. Vangala, "Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification," https://omanscience.com/ar/articles/scores-that-hold-benchmarks-that-leak-measuring-dataset-contamination-in-public-brain-tumor-mri-classification.