Knowledgeknowledgemultitaskllmexam
MMLU
Massive Multitask Language Understanding evaluates broad knowledge and problem-solving across 57 subjects spanning STEM, humanities, and social sciences.
Last reviewed: 23 July 2026
What it measures
Factual knowledge and reasoning under a multiple-choice exam format. It is a proxy for general capability rather than a single skill.
Input / Output
Input: multiple-choice questions (4 options) across 57 subjects. Output: selected answer letter; scored as accuracy.
Evaluation methodology
Zero-shot or few-shot prompting. Accuracy averaged across subjects (macro-average). Variants include MMLU-Pro and MMLU-Redux for harder / cleaned sets.
Metadata
- Task
- Multiple-choice knowledge QA
- Domain
- General knowledge / academic subjects
- Modality
- Text
- Input type
- Question + 4 choices
- Output type
- Answer letter
- Evaluation type
- Automatic accuracy
- Primary metric
- Accuracy (macro-average)
- Secondary metrics
- Per-subject accuracy, MMLU-Pro accuracy
- Paper
- Measuring Massive Multitask Language Understanding
- GitHub
- hendrycks/test
- Dataset
- MMLU test set (hendrycks/test)
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- GPQAScienceGraduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.
- LiveBenchGeneralContamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
- HELMHolisticHolistic Evaluation of Language Models (Stanford CRFM): multi-metric, multi-scenario evaluation emphasizing transparency and taxonomy of use cases.
- Arena HardChatChallenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.
- LongBenchLong ContextBilingual long-context benchmark covering single/multi-document QA, summarization, few-shot learning, synthetic tasks, and code across long inputs.
- CRAGRAGComprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.