Sciencesciencegraduatehardreasoning
GPQA
Graduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.
Last reviewed: 23 July 2026
What it measures
Expert-level scientific reasoning in biology, physics, and chemistry — resistance to shallow retrieval and memorization.
Input / Output
Input: hard multiple-choice science questions. Output: answer choice; primary metric is accuracy (often Diamond subset).
Evaluation methodology
Zero-shot / few-shot evaluation. Diamond subset is the hardest. Compare against human expert baselines.
Metadata
- Task
- Graduate science QA
- Domain
- Biology, physics, chemistry
- Modality
- Text
- Input type
- Multiple-choice question
- Output type
- Answer choice
- Evaluation type
- Automatic accuracy
- Primary metric
- GPQA Diamond accuracy
- Secondary metrics
- GPQA Main accuracy, Human expert baseline
- Paper
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- GitHub
- idavidrein/gpqa
- Dataset
- GPQA on Hugging Face
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- MMLUKnowledgeMassive Multitask Language Understanding evaluates broad knowledge and problem-solving across 57 subjects spanning STEM, humanities, and social sciences.
- Arena HardChatChallenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.
- LiveBenchGeneralContamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
- HELMHolisticHolistic Evaluation of Language Models (Stanford CRFM): multi-metric, multi-scenario evaluation emphasizing transparency and taxonomy of use cases.
- CRAGRAGComprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.
- LongBenchLong ContextBilingual long-context benchmark covering single/multi-document QA, summarization, few-shot learning, synthetic tasks, and code across long inputs.