AI Benchmarks
Evaluation suites for models, agents, and RAG systems. Every benchmark links to guides, papers, tools, and related rankings.
13 benchmarks · Research feed · Benchmarks guide
Agents
GAIA
AgentsGeneral AI Assistants benchmark: real-world questions requiring tool use, multi-step reasoning, and web/file interaction — hard for both humans and agents.
WebArena
AgentsRealistic web environment benchmark where agents must complete tasks on self-hosted sites (e.g. e-commerce, CMS, forums) via browser actions.
Chat
Coding
General
Holistic
Knowledge
Long Context
Multimodal
RAG
RAGBench
RAGBenchmark suite for retrieval-augmented generation systems — evaluating retrieval quality, groundedness, and answer correctness on realistic RAG workloads.
CRAG
RAGComprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.