Long Contextlong-contextretrievalbilingualllm
LongBench
Bilingual long-context benchmark covering single/multi-document QA, summarization, few-shot learning, synthetic tasks, and code across long inputs.
Last reviewed: 23 July 2026
What it measures
Ability to use information distributed across long contexts (typically 5k–15k+ tokens depending on task).
Input / Output
Input: long documents + task prompt. Output: answers/summaries; task-specific metrics (F1, ROUGE, accuracy).
Evaluation methodology
Per-task metrics then average. Often compared with Needle-in-a-Haystack and RULER for long-context stress tests.
Metadata
- Task
- Long-context understanding
- Domain
- QA, summarization, few-shot, code
- Modality
- Text
- Input type
- Long document(s) + prompt
- Output type
- Answer / summary
- Evaluation type
- Task-specific automatic metrics
- Primary metric
- Average task score
- Secondary metrics
- F1, ROUGE, Accuracy
- Paper
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- GitHub
- THUDM/LongBench
- Dataset
- LongBench
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- RAGBenchRAGBenchmark suite for retrieval-augmented generation systems — evaluating retrieval quality, groundedness, and answer correctness on realistic RAG workloads.
- CRAGRAGComprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.
- HELMHolisticHolistic Evaluation of Language Models (Stanford CRFM): multi-metric, multi-scenario evaluation emphasizing transparency and taxonomy of use cases.
- MMLUKnowledgeMassive Multitask Language Understanding evaluates broad knowledge and problem-solving across 57 subjects spanning STEM, humanities, and social sciences.
- LiveBenchGeneralContamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
- Arena HardChatChallenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.