DataAIHub
DataAIHubNews · Research · Tools · Learning
Agentsagentstoolsmultistepgeneral

GAIA

General AI Assistants benchmark: real-world questions requiring tool use, multi-step reasoning, and web/file interaction — hard for both humans and agents.

Last reviewed: 23 July 2026

What it measures

General assistant competence: planning, tool use, and factual correctness on questions that are conceptually simple for humans but operationally hard for agents.

Input / Output

Input: leveled questions (1–3). Output: short final answers. Metric: exact-match / validated correctness.

Evaluation methodology

Level-based difficulty. Agents may use tools (search, code, files). Compare against human baselines.

Metadata

Task
General AI assistant / tool-using agents
Domain
Real-world multi-step questions
Modality
Text (+ tools / files / web)
Input type
Leveled natural-language questions
Output type
Short final answer
Evaluation type
Exact match / validated correctness
Primary metric
Accuracy by level
Secondary metrics
Level 1 / 2 / 3, Human baseline
Paper
GAIA: a benchmark for General AI Assistants
GitHub
GAIA dataset (Hugging Face)
Dataset
GAIA on Hugging Face
Leaderboard
Official live leaderboard (external)

Live leaderboard

Rankings change as new evaluations are published. View current results on the official leaderboard.

View live leaderboard →

Datasets

GitHub repositories

Related research

Related guides

Related tools

Related models

Companies

Related rankings

Explore more benchmarks

All benchmarks →