performance

Verified Performance Benchmarks

Transparent, verified performance benchmarks for Grep AI. 94% research accuracy, 15-minute turnaround, 50+ sources per report. Updated quarterly.

24 SOURCE-BACKED CLAIMS

Each claim has a visible scope, source, as-of date, and review deadline.

Browse the evidence register

Verified Performance Benchmarks

Performance & Accuracy

We believe in transparency about our capabilities. These benchmarks are derived from systematic testing across thousands of research tasks, verified against ground-truth datasets, and updated quarterly.

How We Measure Performance

Our benchmarking process is designed to be rigorous, reproducible, and transparent.

We maintain a curated dataset of entities with known risk profiles, verified by human experts. This serves as the baseline for measuring accuracy.

GREP AI research is run against the ground-truth dataset without any prior knowledge of expected results. Outputs are compared against known findings.

We measure accuracy across multiple dimensions: entity identification, risk factor detection, source citation accuracy, and completeness of coverage.

Benchmarks are re-run quarterly with updated datasets and methodology. Results are published transparently, including any declines.

Benchmark results directly inform engineering priorities. When we identify gaps, we address them in subsequent releases.

Data Sources & Coverage

Explore the comprehensive data sources that power Grep's research capabilities.

Verify Our Claims Yourself

Don't take our benchmarks at face value. Run a research report on an entity you already know well and compare GREP AI's findings against your existing intelligence.

Performance metrics

01Research Accuracy

94%

Verified against ground-truth datasets

View supporting evidence

02Average Completion

15 min

Deep research report turnaround

View supporting evidence

03Sources per Report

50+

Average databases checked per research

View supporting evidence

04Reports Generated

10,000+

Across all expert modes

View supporting evidence

How We Measure Performance

  1. 1

    Ground-Truth Dataset

    We maintain a curated dataset of entities with known risk profiles, verified by human experts. This serves as the baseline for measuring accuracy.

  2. 2

    Blind Testing

    GREP AI research is run against the ground-truth dataset without any prior knowledge of expected results. Outputs are compared against known findings.

  3. 3

    Multi-Dimensional Scoring

    We measure accuracy across multiple dimensions: entity identification, risk factor detection, source citation accuracy, and completeness of coverage.

  4. 4

    Quarterly Review

    Benchmarks are re-run quarterly with updated datasets and methodology. Results are published transparently, including any declines.

  5. 5

    Continuous Improvement

    Benchmark results directly inform engineering priorities. When we identify gaps, we address them in subsequent releases.

Data Sources & Coverage

01 / catalog

Accuracy Metrics

performance-fact

4 records

02 / catalog

Speed Metrics

performance-fact

4 records

03 / catalog

Coverage Metrics

performance-fact

4 records

04 / catalog

Reliability Metrics

performance-fact

4 records

Frequently Asked Questions

How do you define 'accuracy'?

We measure accuracy as the percentage of known risk factors correctly identified in our ground-truth test dataset. This includes sanctions matches, litigation findings, adverse media, and corporate registry data. Our 94% accuracy rate reflects the overall detection rate across all risk categories.

Why not 100% accuracy?

No research system — human or AI — achieves 100% accuracy. The remaining 6% typically involves edge cases: entities with extremely common names requiring disambiguation, findings only available in non-digital records, or very recent events not yet indexed by our sources. We're transparent about our limitations.

How do these benchmarks compare to manual research?

Independent comparisons show that a typical manual research process using 5-10 databases achieves approximately 60-75% coverage of the findings GREP AI produces from 50+ databases. Human researchers may catch nuances that AI misses, but miss far more through limited source coverage.

Are these benchmarks independently verified?

Our benchmarking methodology is documented and reproducible. We welcome third-party verification and have shared our methodology with enterprise customers for independent validation.

How often are benchmarks updated?

We re-run our complete benchmark suite quarterly with updated ground-truth datasets. Results are published on this page within two weeks of testing completion. We report all results, including any performance declines.

Evidence register

Every claim used on this page, with its scope, date, review deadline, and supporting source.

24 evidence claims
  1. Verified Performance Benchmarks

    Scope:
    performance-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  2. We believe in transparency about our capabilities. These benchmarks are derived from systematic testing across thousands of research tasks, verified against ground-truth datasets, and updated quarterly.

    Scope:
    performance-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  3. We maintain a curated dataset of entities with known risk profiles, verified by human experts. This serves as the baseline for measuring accuracy.

    Scope:
    performance-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  4. Explore the comprehensive data sources that power Grep's research capabilities.

    Scope:
    performance-page-snapshot
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  5. Research Accuracy: 94%

    Scope:
    performance-stat
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  6. Average Completion: 15 min

    Scope:
    performance-stat
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  7. Sources per Report: 50+

    Scope:
    performance-stat
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  8. Reports Generated: 10,000+

    Scope:
    performance-stat
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  9. Entity Identification: 97% accuracy in correctly identifying and disambiguating target entities

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  10. Risk Factor Detection: 94% detection rate for known risk factors (sanctions, litigation, adverse media)

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  11. Citation Accuracy: 99%+ of cited sources link to valid, verifiable primary documents

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  12. False Positive Rate: Less than 5% false positive rate on sanctions and adverse media screening

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  13. Ultra Fast Mode: Average completion in 2-3 minutes with core risk screening

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  14. Standard Mode: Average completion in 8-12 minutes with comprehensive coverage

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  15. Deep Research Mode: Average completion in 12-18 minutes with exhaustive analysis

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  16. API Response Time: Sub-second API acknowledgment with webhook delivery on completion

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  17. Jurisdictional Coverage: Corporate registry data available for 200+ countries and territories

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  18. Sanctions List Coverage: 50+ sanctions, watchlist, and PEP databases screened per query

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  19. Media Coverage: Global news archive spanning 10+ years with multi-language support

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  20. Court Record Coverage: Federal and major state court records across the United States

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  21. System Uptime: 99.9% uptime over the trailing 12 months

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  22. Data Freshness: Sanctions data updated within hours of official publication

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  23. Report Consistency: Same query produces consistent results — no methodology variance

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
  24. Source Availability: 98%+ source availability with automatic failover for maintenance windows

    Scope:
    performance-fact
    As of:
    Review by:
    Snapshot:
    src/components/marketing/trust/PerformanceBenchmarks.tsx
Evidence current as of