DataAIHub
DataAIHubNews · Research · Tools · Learning
RAGragcorrectiveretrievalfactuality

CRAG

Comprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.

Last reviewed: 23 July 2026

What it measures

How RAG systems behave when retrieval is incomplete, noisy, or conflicting — not only when the right passage is easy to find.

Input / Output

Input: queries over corpora with controlled retrieval conditions. Output: answers + citations; scored for correctness and groundedness.

Evaluation methodology

Multi-scenario evaluation (accurate retrieval, incomplete, incorrect). Report scenario-level and aggregate scores.

Metadata

Task
Realistic / corrective RAG
Domain
Industry QA / factuality
Modality
Text
Input type
Query + retrieval conditions
Output type
Answer (+ optional citations)
Evaluation type
Scenario-based automatic scoring
Primary metric
Accuracy / score by scenario
Secondary metrics
Dynamic vs static, Incomplete retrieval score
Paper
CRAG — Comprehensive RAG Benchmark
GitHub
facebookresearch/CRAG
Dataset
facebookresearch/CRAG
Leaderboard
Official live leaderboard (external)

Live leaderboard

Rankings change as new evaluations are published. View current results on the official leaderboard.

View live leaderboard →

Datasets

GitHub repositories

Related research

Related guides

Related tools

Related models

Companies

Related rankings

Explore more benchmarks

All benchmarks →