Abstract
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Yang, M., Pan, Y., Piao, J., Song, D., Gong, Y., & Li, Y. (2026). EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making. https://omanscience.com/en/articles/enterprisebench-benchmarking-llm-agents-on-enterprise-level-strategic-reasoning-and-decision-making
MLA 9
Yang, Min, et al. "EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making." https://omanscience.com/en/articles/enterprisebench-benchmarking-llm-agents-on-enterprise-level-strategic-reasoning-and-decision-making.
Chicago (author–date)
Yang, Min, Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong, and Yong Li. 2026. "EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making." https://omanscience.com/en/articles/enterprisebench-benchmarking-llm-agents-on-enterprise-level-strategic-reasoning-and-decision-making.
Harvard
Yang, M., Pan, Y., Piao, J., Song, D., Gong, Y. and Li, Y. (2026) 'EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making', Available at: https://omanscience.com/en/articles/enterprisebench-benchmarking-llm-agents-on-enterprise-level-strategic-reasoning-and-decision-making.
Vancouver
Yang M, Pan Y, Piao J, Song D, Gong Y, Li Y. EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making. https://omanscience.com/en/articles/enterprisebench-benchmarking-llm-agents-on-enterprise-level-strategic-reasoning-and-decision-making
IEEE
M. Yang, Y. Pan, J. Piao, D. Song, Y. Gong, and Y. Li, "EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making," https://omanscience.com/en/articles/enterprisebench-benchmarking-llm-agents-on-enterprise-level-strategic-reasoning-and-decision-making.