الملخص

Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Arda, B., Yavuz, A., Gerry, P., Lobentanzer, S., Sarwar, N., Giner-Miguelez, J., Chen, K., Zhang, L., Sachan, M., & Akhtar, M. (2026). CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets. https://omanscience.com/ar/articles/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml-datasets

MLA 9

Arda, Berke, et al. "CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets." https://omanscience.com/ar/articles/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml-datasets.

شيكاغو (المؤلف–التاريخ)

Arda, Berke, Ahmetcan Yavuz, Paul Gerry, Sebastian Lobentanzer, Nobin Sarwar, Joan Giner-Miguelez, Kongtao Chen, Luyao Zhang, Mrinmaya Sachan, and Mubashara Akhtar. 2026. "CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets." https://omanscience.com/ar/articles/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml-datasets.

هارفارد

Arda, B., Yavuz, A., Gerry, P., Lobentanzer, S., Sarwar, N., Giner-Miguelez, J., Chen, K., Zhang, L., Sachan, M. and Akhtar, M. (2026) 'CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets', Available at: https://omanscience.com/ar/articles/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml-datasets.

فانكوفر

Arda B, Yavuz A, Gerry P, Lobentanzer S, Sarwar N, Giner-Miguelez J, et al. CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets. https://omanscience.com/ar/articles/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml-datasets

IEEE

B. Arda, A. Yavuz, P. Gerry, S. Lobentanzer, N. Sarwar, J. Giner-Miguelez, K. Chen, L. Zhang, M. Sachan, and M. Akhtar, "CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets," https://omanscience.com/ar/articles/croissantminer-automated-extraction-and-validation-of-croissant-metadata-for-ml-datasets.