Multimodalmultimodalvisioncollegereasoning
MMMU
Massive Multi-discipline Multimodal Understanding: college-level problems combining images and text across six disciplines.
Last reviewed: 23 July 2026
What it measures
Expert-level multimodal reasoning — diagrams, charts, and figures with subject knowledge.
Input / Output
Input: image(s) + question (often multiple choice). Output: answer; accuracy by subject and overall.
Evaluation methodology
Zero-shot / few-shot VLM evaluation. Subject-level and aggregate accuracy. Compare open vs proprietary VLMs.
Metadata
- Task
- Multimodal college-level QA
- Domain
- Multi-discipline academic
- Modality
- Vision + text
- Input type
- Image(s) + question
- Output type
- Answer / choice
- Evaluation type
- Automatic accuracy
- Primary metric
- Overall accuracy
- Secondary metrics
- Per-discipline accuracy, Perception vs reasoning split
- Paper
- MMMU: A Massive Multi-discipline Multimodal Understanding Benchmark
- GitHub
- MMMU-Benchmark/MMMU
- Dataset
- MMMU Benchmark
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- MMLUKnowledgeMassive Multitask Language Understanding evaluates broad knowledge and problem-solving across 57 subjects spanning STEM, humanities, and social sciences.
- GPQAScienceGraduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.
- LongBenchLong ContextBilingual long-context benchmark covering single/multi-document QA, summarization, few-shot learning, synthetic tasks, and code across long inputs.
- LiveBenchGeneralContamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
- Arena HardChatChallenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.
- CRAGRAGComprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.