Chatchatpreferencearenaauto-eval
Arena Hard
Challenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.
Last reviewed: 23 July 2026
What it measures
Instruction-following and response quality on difficult conversational prompts, often via LLM-as-judge vs a baseline.
Input / Output
Input: hard user prompts. Output: model responses ranked or judged against a reference model.
Evaluation methodology
Automatic pairwise judging (LLM judge) on Arena-Hard prompt set. Correlates with Chatbot Arena rankings.
Metadata
- Task
- Hard chat preference
- Domain
- Instruction following / dialogue
- Modality
- Text
- Input type
- User prompt
- Output type
- Assistant response
- Evaluation type
- LLM-as-judge / pairwise
- Primary metric
- Win rate vs baseline
- Secondary metrics
- Style control, Arena correlation
- Paper
- From Crowdsourced Data to High-Quality Benchmarks (Arena-Hard)
- GitHub
- lmarena/arena-hard-auto
- Dataset
- Arena-Hard-Auto
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- LiveBenchGeneralContamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
- MMLUKnowledgeMassive Multitask Language Understanding evaluates broad knowledge and problem-solving across 57 subjects spanning STEM, humanities, and social sciences.
- GPQAScienceGraduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.
- HELMHolisticHolistic Evaluation of Language Models (Stanford CRFM): multi-metric, multi-scenario evaluation emphasizing transparency and taxonomy of use cases.
- CRAGRAGComprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.
- LongBenchLong ContextBilingual long-context benchmark covering single/multi-document QA, summarization, few-shot learning, synthetic tasks, and code across long inputs.