Interactive Lab
Retrieval Evaluation Lab
Measure retrieval quality with explicit relevance judgments using Recall@K, MRR, and nDCG@K — and see why more retrieval machinery does not automatically mean better results.
GuidedNo API keys in Lab
Measure → Compare → Explain
This Lab replays measured retrieval evaluation from the Cookbook — no API keys required here. Compare measured retrieval results, calculate the metrics, and see why a more advanced pipeline can still regress on a particular query.
Example question
Pick an example question, then run the measured evaluation walkthrough.
Evaluation walk
Choose an example question and run to inspect how retrieval quality is measured against frozen relevance judgments — not generation quality.
You will see: Dataset → Judgments → Retrieve → Results → Recall → RR/MRR → nDCG → Compare → Failures → Summary
Previous Lab
← Query TransformationContinue Learning
Deeper explanations in the DataAIHub knowledge graph.
Implementation
Run the same evaluation locally from the public Cookbook.
Related Tools
Canonical tools often used with retrieval evaluation stacks.
Lab ID retrieval-evaluation · guided / precomputed