Abstract

The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers to real scientific data analysis remains untested. We introduce ST-Bench, a benchmark designed to answer two questions: whether MAS outperform single agents on complex scientific data analysis tasks, and if so, by how much and at what additional cost. ST-Bench contains 100 data science tasks adapted from published Earth science studies across hydrology, agriculture, and wetland methane research, expanded into 2,067 queries grounded in additional published studies and validated by domain experts. Using ST-Bench, we evaluate five recent MAS generation methods under two training protocols, against single-agent baselines on the same GPT-5 backbone. Nine of the ten MAS configurations exceed the cheapest single-agent baseline, with the strongest reaching nearly three times its composite score. This gain is primarily attributable to coverage: trained workflows produce realistic numerical metrics on a larger fraction of queries, while the quality of those metrics, conditional on producing realistic output, is comparable to that of the single-agent baseline. The strongest configuration requires approximately four times the single-agent inference time, whereas a more economical workflow captures the majority of the benefit at less than twice the cost. MAS specialization confers measurable benefit on scientific data analysis, but the benefit is conditional rather than universal.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Cheng, Q., Dong, R., Chen, S., Liu, L., Lu, D., Chen, Z., Cheng, W., Xie, Y., Chen, H., Jia, X., & Wang, H. (2026). ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks. https://omanscience.com/en/articles/st-bench-a-spatial-temporal-benchmark-for-multi-agent-system-generation-on-scientific-research-tasks

MLA 9

Cheng, Qi, et al. "ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks." https://omanscience.com/en/articles/st-bench-a-spatial-temporal-benchmark-for-multi-agent-system-generation-on-scientific-research-tasks.

Chicago (author–date)

Cheng, Qi, Rongchao Dong, Shengyu Chen, Licheng Liu, Dan Lu, Zhengzhang Chen, Wei Cheng, Yiqun Xie, Haifeng Chen, Xiaowei Jia, and Haoyu Wang. 2026. "ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks." https://omanscience.com/en/articles/st-bench-a-spatial-temporal-benchmark-for-multi-agent-system-generation-on-scientific-research-tasks.

Harvard

Cheng, Q., Dong, R., Chen, S., Liu, L., Lu, D., Chen, Z., Cheng, W., Xie, Y., Chen, H., Jia, X. and Wang, H. (2026) 'ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks', Available at: https://omanscience.com/en/articles/st-bench-a-spatial-temporal-benchmark-for-multi-agent-system-generation-on-scientific-research-tasks.

Vancouver

Cheng Q, Dong R, Chen S, Liu L, Lu D, Chen Z, et al. ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks. https://omanscience.com/en/articles/st-bench-a-spatial-temporal-benchmark-for-multi-agent-system-generation-on-scientific-research-tasks

IEEE

Q. Cheng, R. Dong, S. Chen, L. Liu, D. Lu, Z. Chen, W. Cheng, Y. Xie, H. Chen, X. Jia, and H. Wang, "ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks," https://omanscience.com/en/articles/st-bench-a-spatial-temporal-benchmark-for-multi-agent-system-generation-on-scientific-research-tasks.