AI EvaluationJuly 21, 2026· 8 min read

Why AI Agents Fail on Enterprise Data, Even When the Model Is Powerful

Two new enterprise benchmarks, DAB and EnterpriseClawBench, show frontier models scoring 38% and 0.663 on realistic company data and workflows. Why powerful models aren't automatically reliable agents, and how to evaluate one properly.

Key takeaways

  • On the Data Agent Benchmark's 54 queries across fragmented, real-world enterprise data, the best frontier model reached only 38% pass@1 accuracy
  • EnterpriseClawBench's 852 real-workplace tasks show performance depends on the model, agent framework, tools, cost, runtime, and artifact quality, not one leaderboard score
  • Fragmented systems, inconsistent identifiers, and unstructured text, not model IQ, are where enterprise agents actually fail
  • Evaluate agents end to end on your own data: workflow completion, grounding, escalation, permissions, and auditability

The marketing pitch is familiar: take a powerful frontier model, point it at your systems, and you have an enterprise AI agent. Two recent benchmarks, both built from real enterprise conditions rather than tidy academic datasets, suggest the reality is considerably harder.

The Data Agent Benchmark (DAB), from UC Berkeley's EPIC Data Lab, tests agents on what enterprise data actually looks like: fragmented across multiple heterogeneous database systems, riddled with inconsistent identifiers, with critical information buried in unstructured text. It comprises 54 queries across 12 datasets, 9 domains, and 4 database systems. The result: the best frontier model tested achieved only 38% pass@1 accuracy, and one dataset was never solved correctly by any agent in any trial.

EnterpriseClawBench approaches the problem from the workflow side, constructing 852 reproducible tasks from real workplace agent sessions: reading heterogeneous files, invoking tools, delivering business artifacts. Its strongest reported configuration scored 0.663, and the authors' central conclusion is telling: performance depends on the harness–model combination, artifact delivery, cost, and runtime, not a single leaderboard number.

Neither result says AI agents don't work. Both say something more useful: a powerful general-purpose model does not automatically become a reliable enterprise agent.

Why capable models fail on real company data

The failure modes these benchmarks surface will be familiar to anyone who has run an enterprise data estate:

  • Fragmentation. The answer to one question lives across a CRM, a data warehouse, a document store, and a ticketing system: each with its own schema, access model, and quirks. Benchmarks that test a single clean table miss the hard part: the joins across systems.
  • Inconsistent identifiers. The same employee, asset, or customer appears under different names, IDs, and formats in different systems. Humans resolve this with context; agents frequently don't.
  • Information buried in unstructured text. The decisive fact is often in a ticket comment, a policy PDF, or an email thread, not a queryable field.
  • Long pipelines compound error. An end-to-end task chains many steps: find the right source, transform, join, reason, act. A model that is 90% reliable per step is far less reliable across ten of them. This is why single-turn accuracy scores flatter agents and end-to-end benchmarks humble them.

Why a single leaderboard score misleads buyers

EnterpriseClawBench's design makes a point every buyer should absorb: the same model scores very differently depending on the agent framework, tools, and scaffolding around it. Which means:

  1. "Powered by [frontier model]" tells you almost nothing. The harness (retrieval, tool integrations, workflow design, error recovery), drives outcomes as much as the model does.
  2. Cost and runtime are part of the score. An agent that gets the right answer after ten minutes and dozens of expensive reasoning loops may be worse in production than a faster, cheaper agent with slightly lower raw accuracy.
  3. Artifact quality matters. Enterprise work ends in a deliverable (a resolved ticket, an updated record, a usable document), not a chat reply. Evaluation has to check the deliverable.

How to evaluate an enterprise service-desk agent instead

If leaderboard scores can't answer "will this work for us," what does? Evaluate the way these benchmarks do: end to end, on your reality:

  • Test on your data, not demo data. Fragmented systems, your naming inconsistencies, your knowledge base, including its stale corners.
  • Measure workflow completion, not answer quality. Did the ticket get resolved, the access granted, the record updated, correctly and completely?
  • Check grounding and citations. Can every answer be traced to a source your team can verify? Ungrounded fluency is how errors ship at scale.
  • Test escalation behavior. The most important skill of a production agent is knowing when not to act, handing off cleanly with context instead of guessing.
  • Verify permissions. The agent must respect the asking employee's entitlements. An agent that answers correctly from data the requester shouldn't see has failed a worse test.
  • Demand auditability. Every action traceable: what was asked, what was retrieved, what was done, and under whose approval.
  • Account for cost and latency at your volume. Per-task economics that look fine in a pilot can break at ten thousand tickets a month.

We've packaged these into a practical, printable evaluation framework: the Enterprise AI Agent Readiness Checklist, covering data access, retrieval accuracy, workflow completion, escalation, permissions, and auditability.

How Rezolve.ai builds for this reality

These benchmark results are, in effect, the design brief Rezolve.ai was built against. Sidekick answers employees with grounded, cited responses, the traceability the evaluation above demands, and How Sidekick Thinks shows eight specialized AI agents reasoning over one shared conversation, because reliable enterprise work takes orchestration, not a single model call. Automation created in Agent Studio is approval-gated, audited, and explainable, and DeskIQ grounds automation decisions in your actual ticket history rather than assumptions. The approach holds up in production: roughly 70% of requests resolved before they become tickets, JLL running a program for ~100,000 employees across 80+ countries, and Black Angus Steakhouse cutting after-hours IT dependency from 90% to 10%.

Bottom line

DAB and EnterpriseClawBench put numbers on what practitioners already suspected: raw model power is necessary but nowhere near sufficient for reliable enterprise agents. Context, integrations, permissions, workflow design, and rigorous evaluation are where deployments succeed or fail. Evaluate agents the way these researchers do (end to end, on real conditions), and insist on grounding, governance, and auditability before anything touches your employees.

Download the Enterprise AI Agent Readiness Checklist to run that evaluation, or book a demo to see how a governed, glass-box agent handles your real tickets.

See the agentic service desk in action

Watch Rezolve.ai autonomously resolve real IT and HR tickets: governed, auditable, glass-box.

Book a demo
Neil Dattani
LinkedIn ↗

Get service-desk AI insights in your inbox

Practical guidance on agentic AI for IT and HR support: one email, no spam.

We use these details to respond to you. See our Privacy Policy.