DataAIHub
DataAIHubNews · Research · Tools · Learning
Knowledgeknowledgemultitaskllmexam

MMLU

Massive Multitask Language Understanding evaluates broad knowledge and problem-solving across 57 subjects spanning STEM, humanities, and social sciences.

Last reviewed: 23 July 2026

What it measures

Factual knowledge and reasoning under a multiple-choice exam format. It is a proxy for general capability rather than a single skill.

Input / Output

Input: multiple-choice questions (4 options) across 57 subjects. Output: selected answer letter; scored as accuracy.

Evaluation methodology

Zero-shot or few-shot prompting. Accuracy averaged across subjects (macro-average). Variants include MMLU-Pro and MMLU-Redux for harder / cleaned sets.

Metadata

Task
Multiple-choice knowledge QA
Domain
General knowledge / academic subjects
Modality
Text
Input type
Question + 4 choices
Output type
Answer letter
Evaluation type
Automatic accuracy
Primary metric
Accuracy (macro-average)
Secondary metrics
Per-subject accuracy, MMLU-Pro accuracy
Paper
Measuring Massive Multitask Language Understanding
GitHub
hendrycks/test
Dataset
MMLU test set (hendrycks/test)
Leaderboard
Official live leaderboard (external)

Live leaderboard

Rankings change as new evaluations are published. View current results on the official leaderboard.

View live leaderboard →

Datasets

GitHub repositories

Related research

Related guides

Related tools

Related models

Companies

Related rankings

Explore more benchmarks

All benchmarks →