DataAIHub
DataAIHubNews · Research · Tools · Learning
Multimodalmultimodalvisioncollegereasoning

MMMU

Massive Multi-discipline Multimodal Understanding: college-level problems combining images and text across six disciplines.

Last reviewed: 23 July 2026

What it measures

Expert-level multimodal reasoning — diagrams, charts, and figures with subject knowledge.

Input / Output

Input: image(s) + question (often multiple choice). Output: answer; accuracy by subject and overall.

Evaluation methodology

Zero-shot / few-shot VLM evaluation. Subject-level and aggregate accuracy. Compare open vs proprietary VLMs.

Metadata

Task
Multimodal college-level QA
Domain
Multi-discipline academic
Modality
Vision + text
Input type
Image(s) + question
Output type
Answer / choice
Evaluation type
Automatic accuracy
Primary metric
Overall accuracy
Secondary metrics
Per-discipline accuracy, Perception vs reasoning split
Paper
MMMU: A Massive Multi-discipline Multimodal Understanding Benchmark
GitHub
MMMU-Benchmark/MMMU
Dataset
MMMU Benchmark
Leaderboard
Official live leaderboard (external)

Live leaderboard

Rankings change as new evaluations are published. View current results on the official leaderboard.

View live leaderboard →

Datasets

GitHub repositories

Related research

Related guides

Related tools

Related models

Companies

Related rankings

Explore more benchmarks

All benchmarks →