EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
EnterpriseOps-Gym is a containerized, resettable enterprise simulation benchmark for evaluating LLM agents on stateful, multi-step planning and tool use across realistic enterprise workflows
Authors
Shiva Krishna Reddy Malay*,1
Shravan Nayak*,1,2,3
Jishnu Sethumadhavan Nair1
Aman Tiwari1
Sathwik Tejaswi Madhusudhan1
Sagar Davasam1
Sridhar Krishna Nemala1
Srinivas Sunkara1
Sai Rajeswar1,2,3
*Equal contribution |
1ServiceNow AI Research |
2Mila β Quebec AI Institute |
3UniversitΓ© de MontrΓ©al
---
## π Introduction
**EnterpriseOps-Gym** evaluates LLM agents on **1,150 expert-curated tasks** across **8 enterprise domains** β Calendar, CSM, Drive, Email, HR, ITSM, Teams, and Hybrid β in a fully interactive, containerized environment.
Unlike static datasets, tasks run against live MCP servers and are evaluated by SQL verifiers that check **final environment state**, not action sequences.
**Key Features:**
- π οΈ **512 tools** across 8 enterprise domains
- ποΈ **164 database tables** with avg 1.7 foreign-key dependencies per table
- π’ **9.15 avg steps** per task (up to 34), with **5.3 avg verification conditions**
- π **89k avg context length** per task
- π Best model achieves only **34.1%** success rate β significant headroom for improvement