Generalcontamination-freelivellmupdated
LiveBench
Contamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
Last reviewed: 23 July 2026
What it measures
General capability with reduced train-set contamination risk by releasing new questions over time.
Input / Output
Input: tasks across multiple categories with objective ground truth. Output: task-specific answers; scored by category then aggregate.
Evaluation methodology
Objective scoring (not LLM-as-judge for core tasks). Monthly refreshes. Category scores + overall LiveBench score.
Metadata
- Task
- Multi-category capability
- Domain
- Math, coding, reasoning, language, IF, data analysis
- Modality
- Text
- Input type
- Task prompts with ground truth
- Output type
- Task-specific answers
- Evaluation type
- Objective automatic scoring
- Primary metric
- LiveBench overall score
- Secondary metrics
- Category scores, Monthly delta
- Paper
- LiveBench: A Challenging, Contamination-Free LLM Benchmark
- GitHub
- LiveBench/LiveBench
- Dataset
- LiveBench releases
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- Arena HardChatChallenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.
- MMLUKnowledgeMassive Multitask Language Understanding evaluates broad knowledge and problem-solving across 57 subjects spanning STEM, humanities, and social sciences.
- HumanEvalCodingClassic function-level Python code generation benchmark: models complete function bodies from docstrings and signatures.
- GPQAScienceGraduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.
- HELMHolisticHolistic Evaluation of Language Models (Stanford CRFM): multi-metric, multi-scenario evaluation emphasizing transparency and taxonomy of use cases.
- LongBenchLong ContextBilingual long-context benchmark covering single/multi-document QA, summarization, few-shot learning, synthetic tasks, and code across long inputs.