What Is the Draco Benchmark? A Guide to AI Agent Scoring
The Draco benchmark scores deep research AI agents on accuracy and traceability. Discover how it works and why it matters for compliance and diligence teams.

Quick Answer
The DRACO benchmark evaluates whether a deep-research agent can complete complex research tasks with accurate, well-supported, and usable outputs. For enterprise buyers, AI agent scoring is useful only when it connects benchmark performance to source traceability, reviewability, and the ability to defend a decision.
Introduction
A DRACO benchmark score helps distinguish an agent that produces plausible prose from one that can support due diligence, compliance review, or investment research. It evaluates research quality across multi-step tasks rather than rewarding a single fluent answer. That matters when a conclusion must be checked by an analyst and presented to a decision-maker. Shopmonkey closed 64 research jobs in its first 30 days on Grep and cut underwriting research time from hours to minutes per account, beating Gemini head to head, a concrete example of what traceable research delivers beyond a raw benchmark score. Bad research does not merely consume time; it can alter decisions by introducing unsupported claims into a review process.
Key Takeaways:
DRACO evaluates deep-research performance through structured task rubrics.
High scores matter most when outputs remain traceable and reviewable.
Benchmark results should inform, not replace, enterprise validation.

What the DRACO Benchmark Measures in AI Agent Scoring
DRACO is a deep-research evaluation framework for agents asked to investigate, synthesize, and communicate answers to complex questions. It assesses the completed research product, not simply whether an agent can retrieve isolated facts. This makes it more relevant to a deep research AI agent for enterprise than a generic prompt-response test.
From a research prompt to an evaluated deliverable
A DRACO-style task requires the agent to identify relevant evidence, reconcile conflicting information, form a justified conclusion, and present the result coherently. DRACO was developed by researchers from Perplexity and Harvard University and comprises 100 curated tasks spanning ten domains, scored against thousands of expert-defined criteria. In high-stakes settings, AI-generated claims still require rigorous human verification against credible sources before they inform a decision, since the reviewer who relies on unverified information bears responsibility for the error, making human review a governance requirement rather than a cosmetic final step.
Task scope: The question requires investigation, not recall.
Evidence coverage: Relevant sources must inform the conclusion.
Factual accuracy: Claims must match the supporting evidence.
Reasoning quality: The analysis must connect evidence to findings.
Presentation: The result must be usable by a reviewer.
Why rubric-based evaluation is harder to game
Rubric-based scoring forces evaluators to look beyond polished language. DRACO's reported pass rates varied by domain: Perplexity Deep Research reached 89.4% in Law and 82.4% in Academic evaluation, the benchmark's own publisher reported. Those differences show why a single aggregate score should be supplemented with task-level evidence that matches the buyer's own workflow. That domain variance illustrates why a clear benchmark methodology must make quality criteria explicit before a score can be trusted. In a separate Microsoft report, Researcher with Critique improved the aggregated DRACO score by 7.0 points (SEM ± 1.90), or 13.88%, over Perplexity Deep Research using Claude Opus 4.6. The result underscores the value of testing both generation and review steps rather than treating research quality as a single capability.

How DRACO Scores Deep Research Agents
DRACO scoring should be read as an assessment of research execution under a defined task set, not as a universal intelligence ranking. The score becomes meaningful when a buyer can inspect what the benchmark rewards, how answers are judged, and whether the tested work resembles the organization's real risk decisions.
Core dimensions behind a credible research score
Strong performance requires multiple capabilities to hold together: finding evidence, interpreting it correctly, handling uncertainty, and writing conclusions that preserve the trail back to source material. A high aggregate score can hide uneven behavior, so procurement teams should ask for task-level examples, citation review, and failure cases alongside headline results. Buyers evaluating this should also confirm the platform's verified data sources, since a benchmark score is only as trustworthy as the evidence feeding it.
For regulated workflows, the operational test is whether the analyst can reconstruct the agent's path. System operation logs are relevant because they help identify risks, support post-market monitoring, and track how a system behaved. The same principle applies to research: an answer without a reviewable decision trail is difficult to govern.
Generic assistants and purpose-built research agents
Generic assistants and purpose-built agents are not the same product category. Microsoft describes Researcher as an advanced Microsoft 365 Copilot agent for complex, multi-step research, while a research platform is evaluated on the evidence-backed deliverables it produces for a defined high-stakes workflow. Microsoft reports results for a Researcher-with-Critique approach, which is a useful distinction for buyers to test in their own operating model. The comparison below focuses on the distinction a buyer should validate.
Evaluation area | Generic AI assistant | Purpose-built research agent |
|---|---|---|
Primary interaction | Prompt-led assistance | Defined research assignment |
Output review | Review trail varies by product configuration | Can preserve citations and decision trails |
Workflow continuity | Workflow continuity varies by deployment | Can support scheduled or event-triggered work |
Benchmark evidence | Varies by product and task | Can be tested against deep-research tasks |
Governance fit | Depends on deployment controls | Designed around reviewable deliverables |
The key tradeoff is not conversational quality; it is whether the output can become an auditable artifact that survives analyst review, internal challenge, and external scrutiny.
Where Benchmark Performance Meets Enterprise Research Work
Benchmark results are most valuable when they predict better operating behavior in the work that carries real consequences. Agentic AI for institutional onboarding and continuous monitoring requires more than a cited summary: it requires evidence coverage, escalation paths, and repeatable output formats.
Translate benchmark claims into a vendor test plan
Start by selecting a live use case with a known review standard, such as a counterparty assessment or an acquisition diligence brief. Then compare source coverage, unsupported claims, citation usefulness, analyst corrections, and time needed to reach an approved decision. This is the practical form of setting research accuracy standards, because a benchmark score alone cannot reveal whether the agent understands an organization's risk policy.
Grep builds custom agents for high-stakes knowledge work, including due diligence, institutional onboarding, compliance reviews, and continuous monitoring. Its Agents produce citation-backed reports, slide decks, and spreadsheets, while Loops and Monitors run scheduled or event-triggered workflows and ongoing screening for changes such as leadership, website, job-posting, or regulatory developments. A vendor test should separately inspect whether citations support the stated conclusion, whether missing evidence is surfaced, whether an analyst can correct the record, and whether the resulting deliverable preserves the basis for approval. These checks turn benchmark concepts into observable acceptance criteria for a live workflow.
Read Grep's DRACO result as dated evidence
As of April 2026, Grep's benchmark results published by Grep stated that it ranked number one on DRACO, DeepSearchQA, and DeepResearch Bench, with an 18.8-point lead. That is a specific, dated performance claim, not a permanent category ranking, and buyers should test it against their own research corpus and approval process.
The claim is relevant because a deep-research platform performance matters when teams must convert evidence into outputs a board or regulator can inspect. Grep also states that its platform has been in regulated production since 2023, with SOC 2 and GDPR as trust considerations, no model training on customer data, scoped least-privilege credentials, exportable decision trails, configurable retention, and delete-on-request controls. Teams can also review Grep's pricing to evaluate deployment options alongside benchmark performance.

Conclusion
DRACO gives enterprise teams a structured way to examine whether an agent can conduct research rather than merely generate text. Use it to narrow a vendor field, then run a controlled evaluation using real work, known sources, and human review criteria. For teams moving high-stakes research beyond generic assistants, evaluate whether any platform can provide custom agents producing traceable, auditable, and defensible deliverables for due diligence, compliance oversight, or continuous monitoring. The decisive measure remains whether each finding can be checked before it influences a material decision.
Ready to assess research workflows against a higher trust bar? Explore Grep AI for high-stakes research work.
Frequently Asked Questions (FAQs)
How are AI agent benchmarks like DRACO scored?
AI agent benchmarks like DRACO are scored by evaluating completed research tasks against defined criteria such as factual accuracy, evidence coverage, analytical quality, citation quality, and presentation, rather than by judging whether an answer merely sounds fluent.
How does the DRACO benchmark evaluate deep research agents?
The DRACO benchmark evaluates deep research agents by testing whether they can investigate complex questions, synthesize evidence into supported conclusions, and produce usable research outputs that can be examined against a structured evaluation rubric.
What is the difference between generic AI and custom research agents?
The difference between generic AI and custom research agents is that generic AI primarily responds to prompts, while custom research agents can be configured around defined high-stakes assignments, evidence requirements, deliverable formats, and repeatable review workflows.
How does Grep provide traceable and defensible AI outputs?
Grep provides traceable and defensible AI outputs through citation-backed reports, slide decks, and spreadsheets, plus exportable decision trails that let reviewers inspect the evidence and reasoning behind findings used in high-stakes work.
What makes an AI agent suitable for regulatory compliance?
An AI agent is suitable for regulatory compliance when it supports controlled access, preserves a reviewable trail, follows defined policies, handles evidence carefully, and enables human experts to validate findings before they affect regulated decisions.
Why is auditability critical for AI in financial services?
Auditability is critical for AI in financial services because institutions must explain how a finding was reached, identify the information used, correct errors, and demonstrate that material decisions received appropriate human review and governance.
About the Author
Miguel Rios-Berrios is Founder and CTO of GREP.ai, with experience leading engineering and data science teams in fintech and enterprise software. His work focuses on AI agents, distributed systems, and building traceable systems for compliance and other high-stakes business operations. Connect on LinkedIn.