AI EvaluationJuly 21, 2026· 7 min read

From Employee Search to Deep Research: What DRBench Means for Enterprise AI

DRBench, from ServiceNow Research and academic collaborators, tests whether AI agents can combine public web and private company data into cited, accurate research reports, and previews what enterprises will soon demand of every agent.

Key takeaways

  • DRBench evaluates agents on 100 open-ended research tasks across ten business domains, spanning both public web information and private organizational data
  • Agents are scored on insight recall, factual accuracy, and proper citation, and must ignore deliberately planted distractor content
  • Enterprise AI is shifting from FAQ answers to research-shaped work: policy analysis, incident investigation, HR questions, and compliance research
  • The benchmark's success criteria (citations, grounding, permission-aware retrieval, relevance discipline), are the right lens for evaluating any enterprise agent today

Most enterprise AI benchmarks test whether an agent can look something up. A new benchmark from researchers affiliated with ServiceNow, several universities, and Mila asks a harder question: can an agent research, combining public web information with private organizational data into a defensible, cited report?

DRBench evaluates AI agents on complex, open-ended enterprise deep-research tasks. The released benchmark spans 100 tasks across ten business domains (including sales, cybersecurity, and compliance), with a heterogeneous search space of roughly 1,093 files: spreadsheets, PDFs, emails, chat conversations, cloud file systems, and the open web. Tasks are grounded in realistic personas and company contexts, with questions like "What changes should we make to our product roadmap to ensure compliance?": the kind of assignment a real analyst would get, not a fact lookup.

What makes DRBench worth an IT leader's attention is less the leaderboard and more the definition of success it encodes. That definition is a preview of what enterprises will soon demand from every AI agent they deploy.

What DRBench actually measures

Three design choices stand out:

1. Public plus private, together. Real enterprise questions almost never live entirely inside or entirely outside the firewall. A compliance question needs the regulation (public) and your product roadmap, policies, and internal discussions (private). DRBench forces agents to traverse both (across chat logs, file shares, email, and websites), within a realistic containerized enterprise environment.

2. Distractors are part of the test. The benchmark deliberately embeds internal content that is plausible but irrelevant alongside the insights that matter. Agents are scored not only on what they find but on what they correctly ignore. Anyone who has watched a search tool confidently surface a 2019 draft policy knows why this matters.

3. The deliverable is a cited report. Agents are evaluated on insight recall, factual accuracy, and whether findings are properly attributed to sources. An uncited claim, however correct, is not defensible, and DRBench treats attribution as a first-class scoring criterion rather than a nicety.

The researchers' evaluations across open- and closed-source models and research strategies show meaningful headroom: no configuration comes close to solving enterprise deep research. The capability is emerging, not arrived.

From employee search to deep research

DRBench formalizes a shift already underway in what employees ask their service desk. The first generation of enterprise AI answered FAQs: "How do I reset my VPN?" The questions arriving now look more like research assignments:

  • Policy analysis. "Does our travel policy allow this vendor arrangement, and what changed in the latest revision?": requires the current policy, the previous version, and often an external regulation.
  • Incident investigation. "Have we seen this error before, what fixed it, and is it related to last month's change?": spans ticket history, change records, and vendor advisories on the public web.
  • HR questions with context. "How does parental leave interact with my state's requirements?": internal policy plus external law, personalized to the asker.
  • Compliance research. "Which of our systems are affected by this new requirement?": external mandate cross-referenced against internal assets, owners, and configurations.

Each of these has the DRBench shape: multiple internal sources, at least one external source, plausible distractors, and an answer that must be defensible, because someone will act on it.

The requirements this creates for any enterprise agent

If your AI is going to do research-shaped work, four capabilities move from optional to mandatory:

  • Citations on every claim. Employees and auditors must be able to trace each statement to its source. Attribution is what turns an answer into evidence.
  • Factual grounding over fluent synthesis. A well-written report with one fabricated detail is worse than no report; evaluation must check facts against sources, the way DRBench does.
  • Permission-aware retrieval. Research that traverses HR files, finance data, and security records must respect the asker's entitlements at every hop. A correct answer built on data the requester shouldn't see is a governance failure.
  • Relevance discipline. The agent must ignore stale drafts, superseded policies, and adjacent-but-wrong documents, the distractor problem, rather than padding reports with everything it found.

How Rezolve.ai approaches research-shaped questions

This definition of success is one we recognize, because it mirrors how Rezolve.ai is built. Sidekick answers employees on Teams, Slack, email, and voice with grounded, cited responses (attribution is the default, not a premium feature), and How Sidekick Thinks shows eight specialized AI agents reasoning over one shared conversation, the orchestration that multi-source questions require. When research turns into action (a policy exception, an access change), automation built in Agent Studio stays approval-gated, audited, and explainable, and DeskIQ mines your ticket history so recurring research-shaped questions become automated resolutions. It's the glass-box approach that lets organizations like JLL run employee-facing programs for roughly 100,000 people across 80+ countries.

Bottom line

DRBench is a research benchmark, but its real contribution is a standard: enterprise AI agents should be judged on whether they can combine public and private knowledge, ignore distractors, respect permissions, and produce cited, factually accurate output. Today's agents have real headroom against that standard, which makes it exactly the right lens for evaluating vendors now. Ask any AI service desk you're considering how it performs on those four requirements against your own policies and tickets.

Our Enterprise AI Agent Readiness Checklist covers how to run that evaluation, or book a demo to see cited, governed answers on your real content.

See the agentic service desk in action

Watch Rezolve.ai autonomously resolve real IT and HR tickets: governed, auditable, glass-box.

Book a demo
Udaya Bhaskar Reddy
LinkedIn ↗

Get service-desk AI insights in your inbox

Practical guidance on agentic AI for IT and HR support: one email, no spam.

We use these details to respond to you. See our Privacy Policy.