Abstract
Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wu, M., Xie, R., Zhang, R., Li, Y., Fu, T., Chen, B., Yu, K., Chen, X., & Chen, L. (2026). MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows. https://omanscience.com/en/articles/mattoolbench-benchmarking-multimodal-agents-in-real-world-materials-science-workflows
MLA 9
Wu, Mei, et al. "MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows." https://omanscience.com/en/articles/mattoolbench-benchmarking-multimodal-agents-in-real-world-materials-science-workflows.
Chicago (author–date)
Wu, Mei, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, and Lu Chen. 2026. "MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows." https://omanscience.com/en/articles/mattoolbench-benchmarking-multimodal-agents-in-real-world-materials-science-workflows.
Harvard
Wu, M., Xie, R., Zhang, R., Li, Y., Fu, T., Chen, B., Yu, K., Chen, X. and Chen, L. (2026) 'MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows', Available at: https://omanscience.com/en/articles/mattoolbench-benchmarking-multimodal-agents-in-real-world-materials-science-workflows.
Vancouver
Wu M, Xie R, Zhang R, Li Y, Fu T, Chen B, et al. MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows. https://omanscience.com/en/articles/mattoolbench-benchmarking-multimodal-agents-in-real-world-materials-science-workflows
IEEE
M. Wu, R. Xie, R. Zhang, Y. Li, T. Fu, B. Chen, K. Yu, X. Chen, and L. Chen, "MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows," https://omanscience.com/en/articles/mattoolbench-benchmarking-multimodal-agents-in-real-world-materials-science-workflows.