benchmark
#1 on Every Major Benchmark
GREP AI leads DRACO (78.6%), DeepSearchQA (84.5%), and DeepResearch Bench (56.27) — PhD-level research tasks graded by domain experts.
Each claim has a visible scope, source, as-of date, and review deadline.
Browse the evidence registerDRACO Benchmark
- GREP AI78.6%
- Perplexity DR (Opus 4.6)70.5%
- Claude Opus 4.659.8%
- Gemini Deep Research59%
- OpenAI Deep Research (o3)52.1%
DeepSearchQA
- GREP AI84.5%
- Perplexity Deep Research81.9%
- Moonshot K2.577.1%
- Anthropic Opus 4.576.1%
- Parallel Ultra2x72.6%
DeepResearch Bench
- GREP AI56.27
- Cellcog Max56.13
- nvidia-aiq55.95
- Cellcog55.31
- CMCC-DeepInsight55.24
#1 on Every Major Benchmark
Evaluated on DRACO, DeepSearchQA, and DeepResearch Bench — PhD-level research tasks graded by domain experts.
DRACO Benchmark
100 open-ended research questions across 10 domains, judged by Gemini-2.5-Pro against 3,934 weighted rubric criteria. Grep leads all four evaluation axes and wins 9 of 10 domains.
GREP AI wins 9 of 10 domains
DeepSearchQA
896 multi-step research questions across 17 subject domains. Judge: Gemini 2.5 Flash. Grep achieves 84.5% FC with perfect scores in Linguistics, Biology, and Arts & Entertainment.
14 of 17 categories exceed 80% FC
DeepResearch Bench
100 PhD-level research questions (50 Chinese, 50 English), judged by Gemini-2.5-Pro. A score above 50 means the system outperformed the human expert. Grep leads the field of 34 systems.
Methodology
Grep orchestrates specialised sub-agents — each responsible for search, synthesis, verification, and citation — then merges their outputs into a single, coherent research report.
All reasoning and synthesis steps are powered by Claude Opus 4.6, giving Grep best-in-class analytical depth, nuanced judgement, and instruction following.
Experience #1 Ranked Research
See why Grep outperforms OpenAI, Google, Perplexity, and every specialised research platform on PhD-level tasks.
Evaluation axes
Methodology
Methodology
01 / catalog
Multi-Agent Architecture
benchmark-methodology
1 records
Grep orchestrates specialised sub-agents for search, synthesis, verification, and citation.
02 / catalog
Claude Opus 4.6 Backbone
benchmark-methodology
1 records
All reasoning and synthesis steps use the stated backbone.
Evidence register
Every claim used on this page, with its scope, date, review deadline, and supporting source.
29 evidence claims
#1 on Every Major Benchmark
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
Experience #1 Ranked Research
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
Claude Opus 4.6 Backbone
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
100 open-ended research questions across 10 domains, judged by Gemini-2.5-Pro against 3,934 weighted rubric criteria. Grep leads all four evaluation axes and wins 9 of 10 domains.
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
GREP AI wins 9 of 10 domains
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
896 multi-step research questions across 17 subject domains. Judge: Gemini 2.5 Flash. Grep achieves 84.5% FC with perfect scores in Linguistics, Biology, and Arts & Entertainment.
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
14 of 17 categories exceed 80% FC
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
100 PhD-level research questions (50 Chinese, 50 English), judged by Gemini-2.5-Pro. A score above 50 means the system outperformed the human expert. Grep leads the field of 34 systems.
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
All reasoning and synthesis steps are powered by Claude Opus 4.6, giving Grep best-in-class analytical depth, nuanced judgement, and instruction following.
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
See why Grep outperforms OpenAI, Google, Perplexity, and every specialised research platform on PhD-level tasks.
- Scope:
- benchmark-page-snapshot
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DRACO Benchmark Grep: 78.6
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DRACO Benchmark Perplexity DR (Opus 4.6): 70.5
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DRACO Benchmark Claude Opus 4.6: 59.8
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DRACO Benchmark Gemini Deep Research: 59
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DRACO Benchmark OpenAI Deep Research (o3): 52.1
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepSearchQA Grep: 84.5
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepSearchQA Perplexity Deep Research: 81.9
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepSearchQA Moonshot K2.5: 77.1
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepSearchQA Anthropic Opus 4.5: 76.1
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepSearchQA Parallel Ultra2x: 72.6
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepResearch Bench Grep: 56.27
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepResearch Bench Cellcog Max: 56.13
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepResearch Bench nvidia-aiq: 55.95
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepResearch Bench Cellcog: 55.31
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
DeepResearch Bench CMCC-DeepInsight: 55.24
- Scope:
- benchmark-leaderboard-score
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
Factual Accuracy: 75.4% (+7.5pp vs Perplexity)
- Scope:
- benchmark-evaluation-axis
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
Breadth & Depth: 80.3% (+7.2pp)
- Scope:
- benchmark-evaluation-axis
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
Presentation: 93.3% (+3.0pp)
- Scope:
- benchmark-evaluation-axis
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx
Citation: 79.1% (+14.5pp)
- Scope:
- benchmark-evaluation-axis
- As of:
- Review by:
- Snapshot:
- src/components/marketing/BenchmarkPage.tsx