DataAIHub
DataAIHubNews · Research · Tools · Learning
Sciencesciencegraduatehardreasoning

GPQA

Graduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.

Last reviewed: 23 July 2026

What it measures

Expert-level scientific reasoning in biology, physics, and chemistry — resistance to shallow retrieval and memorization.

Input / Output

Input: hard multiple-choice science questions. Output: answer choice; primary metric is accuracy (often Diamond subset).

Evaluation methodology

Zero-shot / few-shot evaluation. Diamond subset is the hardest. Compare against human expert baselines.

Metadata

Task
Graduate science QA
Domain
Biology, physics, chemistry
Modality
Text
Input type
Multiple-choice question
Output type
Answer choice
Evaluation type
Automatic accuracy
Primary metric
GPQA Diamond accuracy
Secondary metrics
GPQA Main accuracy, Human expert baseline
Paper
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
GitHub
idavidrein/gpqa
Dataset
GPQA on Hugging Face
Leaderboard
Official live leaderboard (external)

Live leaderboard

Rankings change as new evaluations are published. View current results on the official leaderboard.

View live leaderboard →

Datasets

GitHub repositories

Related research

Related guides

Related tools

Related models

Companies

Related rankings

Explore more benchmarks

All benchmarks →