Holisticholisticevaluationtransparencyllm
HELM
Holistic Evaluation of Language Models (Stanford CRFM): multi-metric, multi-scenario evaluation emphasizing transparency and taxonomy of use cases.
Last reviewed: 23 July 2026
What it measures
Broad model behavior: accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency across scenarios — not a single score.
Input / Output
Input: many scenario datasets (QA, info retrieval, toxicity, etc.). Output: per-metric reports and scenario dashboards.
Evaluation methodology
Standardized scenarios + metrics with published runs. Emphasizes documenting assumptions and limitations alongside numbers.
Metadata
- Task
- Holistic multi-metric evaluation
- Domain
- Cross-scenario LLM behavior
- Modality
- Text
- Input type
- Scenario-specific prompts
- Output type
- Scenario responses + metric suite
- Evaluation type
- Multi-metric automatic + taxonomy
- Primary metric
- Scenario-dependent (multi-metric)
- Secondary metrics
- Calibration, Robustness, Fairness, Toxicity, Efficiency
- Paper
- Holistic Evaluation of Language Models
- GitHub
- stanford-crfm/helm
- Dataset
- HELM scenarios
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- MMLUKnowledgeMassive Multitask Language Understanding evaluates broad knowledge and problem-solving across 57 subjects spanning STEM, humanities, and social sciences.
- LiveBenchGeneralContamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
- Arena HardChatChallenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.
- GPQAScienceGraduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.
- LongBenchLong ContextBilingual long-context benchmark covering single/multi-document QA, summarization, few-shot learning, synthetic tasks, and code across long inputs.
- RAGBenchRAGBenchmark suite for retrieval-augmented generation systems — evaluating retrieval quality, groundedness, and answer correctness on realistic RAG workloads.