DataAIHub
DataAIHubNews · Research · Tools · Learning
RAGragretrievalgroundednessevaluation

RAGBench

Benchmark suite for retrieval-augmented generation systems — evaluating retrieval quality, groundedness, and answer correctness on realistic RAG workloads.

Last reviewed: 23 July 2026

What it measures

End-to-end RAG quality: whether retrieved context is relevant and whether generated answers are faithful and correct.

Input / Output

Input: user query + corpus. Output: retrieved docs + generated answer. Metrics include retrieval recall, faithfulness, answer correctness.

Evaluation methodology

Separate retrieval and generation evaluation where possible. Use automated judges and/or gold annotations. Compare against naive RAG baselines.

Metadata

Task
RAG evaluation
Domain
Retrieval-augmented QA
Modality
Text
Input type
Query + document corpus
Output type
Retrieved passages + answer
Evaluation type
Retrieval + generation metrics
Primary metric
Answer correctness / faithfulness
Secondary metrics
Recall@k, Context precision, Groundedness
Paper
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
GitHub
rungalileo/ragbench
Dataset
RAGBench (Galileo)
Leaderboard
Official live leaderboard (external)

Live leaderboard

Rankings change as new evaluations are published. View current results on the official leaderboard.

View live leaderboard →

Datasets

GitHub repositories

Related research

Related guides

Related tools

Related models

Companies

Related rankings

Explore more benchmarks

All benchmarks →