Agentsagentstoolsmultistepgeneral
GAIA
General AI Assistants benchmark: real-world questions requiring tool use, multi-step reasoning, and web/file interaction — hard for both humans and agents.
Last reviewed: 23 July 2026
What it measures
General assistant competence: planning, tool use, and factual correctness on questions that are conceptually simple for humans but operationally hard for agents.
Input / Output
Input: leveled questions (1–3). Output: short final answers. Metric: exact-match / validated correctness.
Evaluation methodology
Level-based difficulty. Agents may use tools (search, code, files). Compare against human baselines.
Metadata
- Task
- General AI assistant / tool-using agents
- Domain
- Real-world multi-step questions
- Modality
- Text (+ tools / files / web)
- Input type
- Leveled natural-language questions
- Output type
- Short final answer
- Evaluation type
- Exact match / validated correctness
- Primary metric
- Accuracy by level
- Secondary metrics
- Level 1 / 2 / 3, Human baseline
- Paper
- GAIA: a benchmark for General AI Assistants
- GitHub
- GAIA dataset (Hugging Face)
- Dataset
- GAIA on Hugging Face
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- WebArenaAgentsRealistic web environment benchmark where agents must complete tasks on self-hosted sites (e.g. e-commerce, CMS, forums) via browser actions.
- SWE-BenchCodingEvaluates systems on real GitHub issues: given a repository and issue, produce a patch that passes the project’s tests.
- LiveBenchGeneralContamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
- Arena HardChatChallenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.
- CRAGRAGComprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.
- GPQAScienceGraduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.