Abstract
Data-driven mechanistic hypotheses are essential to scientific discovery because they explain how underlying processes produce observed phenomena. AI agents and AI scientists increasingly support scientific data analysis. However, their ability to turn empirical findings into mechanistic hypotheses remains insufficiently examined. To address this gap, we introduce MechHypoBench, the first benchmark for evaluating whether AI agents and AI scientists can generate such hypotheses from empirical data. It combines paper-derived mechanisms from 14 scientific fields with real-world datasets containing 17.98 million records. The construction retains the observational complexity of empirical data while providing a specified underlying mechanism. Agents analyze the observations and propose open-form hypotheses. We develop an evaluation framework that assesses open-form mechanistic hypotheses through their consequences under withheld conditions. Experiments with general agents and AI scientists reveal a substantial gap between generated hypotheses and the underlying mechanisms.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Xie, X., Long, Q., Xiao, M., Ju, W., Zhou, Y., Wang, X., & Zhu, H. (2026). From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists. https://omanscience.com/en/articles/from-scientific-observations-to-mechanisms-benchmarking-hypothesis-generation-by-ai-scientists
MLA 9
Xie, Xiaxun, et al. "From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists." https://omanscience.com/en/articles/from-scientific-observations-to-mechanisms-benchmarking-hypothesis-generation-by-ai-scientists.
Chicago (author–date)
Xie, Xiaxun, Qingqing Long, Meng Xiao, Wei Ju, Yuanchun Zhou, Xuezhi Wang, and Hengshu Zhu. 2026. "From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists." https://omanscience.com/en/articles/from-scientific-observations-to-mechanisms-benchmarking-hypothesis-generation-by-ai-scientists.
Harvard
Xie, X., Long, Q., Xiao, M., Ju, W., Zhou, Y., Wang, X. and Zhu, H. (2026) 'From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists', Available at: https://omanscience.com/en/articles/from-scientific-observations-to-mechanisms-benchmarking-hypothesis-generation-by-ai-scientists.
Vancouver
Xie X, Long Q, Xiao M, Ju W, Zhou Y, Wang X, et al. From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists. https://omanscience.com/en/articles/from-scientific-observations-to-mechanisms-benchmarking-hypothesis-generation-by-ai-scientists
IEEE
X. Xie, Q. Long, M. Xiao, W. Ju, Y. Zhou, X. Wang, and H. Zhu, "From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists," https://omanscience.com/en/articles/from-scientific-observations-to-mechanisms-benchmarking-hypothesis-generation-by-ai-scientists.