DataAIHub
DataAIHubNews · Research · Tools · Learning
Chatchatpreferencearenaauto-eval

Arena Hard

Challenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.

Last reviewed: 23 July 2026

What it measures

Instruction-following and response quality on difficult conversational prompts, often via LLM-as-judge vs a baseline.

Input / Output

Input: hard user prompts. Output: model responses ranked or judged against a reference model.

Evaluation methodology

Automatic pairwise judging (LLM judge) on Arena-Hard prompt set. Correlates with Chatbot Arena rankings.

Metadata

Task
Hard chat preference
Domain
Instruction following / dialogue
Modality
Text
Input type
User prompt
Output type
Assistant response
Evaluation type
LLM-as-judge / pairwise
Primary metric
Win rate vs baseline
Secondary metrics
Style control, Arena correlation
Paper
From Crowdsourced Data to High-Quality Benchmarks (Arena-Hard)
GitHub
lmarena/arena-hard-auto
Dataset
Arena-Hard-Auto
Leaderboard
Official live leaderboard (external)

Live leaderboard

Rankings change as new evaluations are published. View current results on the official leaderboard.

View live leaderboard →

Datasets

GitHub repositories

Related research

Related guides

Related tools

Related models

Companies

Related rankings

Explore more benchmarks

All benchmarks →