benchmark

#1 on Every Major Benchmark

GREP AI leads DRACO (78.6%), DeepSearchQA (84.5%), and DeepResearch Bench (56.27) — PhD-level research tasks graded by domain experts.

29 SOURCE-BACKED CLAIMS

Each claim has a visible scope, source, as-of date, and review deadline.

Browse the evidence register

DRACO Benchmark

  1. GREP AI78.6%
  2. Perplexity DR (Opus 4.6)70.5%
  3. Claude Opus 4.659.8%
  4. Gemini Deep Research59%
  5. OpenAI Deep Research (o3)52.1%

DeepSearchQA

  1. GREP AI84.5%
  2. Perplexity Deep Research81.9%
  3. Moonshot K2.577.1%
  4. Anthropic Opus 4.576.1%
  5. Parallel Ultra2x72.6%

DeepResearch Bench

  1. GREP AI56.27
  2. Cellcog Max56.13
  3. nvidia-aiq55.95
  4. Cellcog55.31
  5. CMCC-DeepInsight55.24

#1 on Every Major Benchmark

Evaluated on DRACO, DeepSearchQA, and DeepResearch Bench — PhD-level research tasks graded by domain experts.

DRACO Benchmark

100 open-ended research questions across 10 domains, judged by Gemini-2.5-Pro against 3,934 weighted rubric criteria. Grep leads all four evaluation axes and wins 9 of 10 domains.

GREP AI wins 9 of 10 domains

DeepSearchQA

896 multi-step research questions across 17 subject domains. Judge: Gemini 2.5 Flash. Grep achieves 84.5% FC with perfect scores in Linguistics, Biology, and Arts & Entertainment.

14 of 17 categories exceed 80% FC

DeepResearch Bench

100 PhD-level research questions (50 Chinese, 50 English), judged by Gemini-2.5-Pro. A score above 50 means the system outperformed the human expert. Grep leads the field of 34 systems.

Methodology

Grep orchestrates specialised sub-agents — each responsible for search, synthesis, verification, and citation — then merges their outputs into a single, coherent research report.

All reasoning and synthesis steps are powered by Claude Opus 4.6, giving Grep best-in-class analytical depth, nuanced judgement, and instruction following.

Experience #1 Ranked Research

See why Grep outperforms OpenAI, Google, Perplexity, and every specialised research platform on PhD-level tasks.

Evaluation axes

01Factual Accuracy

75.4%

+7.5pp vs Perplexity

View supporting evidence

02Breadth & Depth

80.3%

+7.2pp

View supporting evidence

03Presentation

93.3%

+3.0pp

View supporting evidence

04Citation

79.1%

+14.5pp

View supporting evidence

Methodology

Methodology

01 / catalog

Multi-Agent Architecture

benchmark-methodology

1 records

  • Grep orchestrates specialised sub-agents for search, synthesis, verification, and citation.

02 / catalog

Claude Opus 4.6 Backbone

benchmark-methodology

1 records

  • All reasoning and synthesis steps use the stated backbone.

Evidence register

Every claim used on this page, with its scope, date, review deadline, and supporting source.

29 evidence claims
  1. #1 on Every Major Benchmark

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  2. Experience #1 Ranked Research

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  3. Claude Opus 4.6 Backbone

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  4. 100 open-ended research questions across 10 domains, judged by Gemini-2.5-Pro against 3,934 weighted rubric criteria. Grep leads all four evaluation axes and wins 9 of 10 domains.

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  5. GREP AI wins 9 of 10 domains

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  6. 896 multi-step research questions across 17 subject domains. Judge: Gemini 2.5 Flash. Grep achieves 84.5% FC with perfect scores in Linguistics, Biology, and Arts & Entertainment.

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  7. 14 of 17 categories exceed 80% FC

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  8. 100 PhD-level research questions (50 Chinese, 50 English), judged by Gemini-2.5-Pro. A score above 50 means the system outperformed the human expert. Grep leads the field of 34 systems.

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  9. All reasoning and synthesis steps are powered by Claude Opus 4.6, giving Grep best-in-class analytical depth, nuanced judgement, and instruction following.

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  10. See why Grep outperforms OpenAI, Google, Perplexity, and every specialised research platform on PhD-level tasks.

    Scope:
    benchmark-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  11. DRACO Benchmark Grep: 78.6

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  12. DRACO Benchmark Perplexity DR (Opus 4.6): 70.5

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  13. DRACO Benchmark Claude Opus 4.6: 59.8

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  14. DRACO Benchmark Gemini Deep Research: 59

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  15. DRACO Benchmark OpenAI Deep Research (o3): 52.1

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  16. DeepSearchQA Grep: 84.5

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  17. DeepSearchQA Perplexity Deep Research: 81.9

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  18. DeepSearchQA Moonshot K2.5: 77.1

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  19. DeepSearchQA Anthropic Opus 4.5: 76.1

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  20. DeepSearchQA Parallel Ultra2x: 72.6

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  21. DeepResearch Bench Grep: 56.27

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  22. DeepResearch Bench Cellcog Max: 56.13

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  23. DeepResearch Bench nvidia-aiq: 55.95

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  24. DeepResearch Bench Cellcog: 55.31

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  25. DeepResearch Bench CMCC-DeepInsight: 55.24

    Scope:
    benchmark-leaderboard-score
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  26. Factual Accuracy: 75.4% (+7.5pp vs Perplexity)

    Scope:
    benchmark-evaluation-axis
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  27. Breadth & Depth: 80.3% (+7.2pp)

    Scope:
    benchmark-evaluation-axis
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  28. Presentation: 93.3% (+3.0pp)

    Scope:
    benchmark-evaluation-axis
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
  29. Citation: 79.1% (+14.5pp)

    Scope:
    benchmark-evaluation-axis
    As of:
    Review by:
    Snapshot:
    src/components/marketing/BenchmarkPage.tsx
Evidence current as of