DataAIHub
DataAIHubNews · Research · Tools · Learning
Long Contextlong-contextretrievalbilingualllm

LongBench

Bilingual long-context benchmark covering single/multi-document QA, summarization, few-shot learning, synthetic tasks, and code across long inputs.

Last reviewed: 23 July 2026

What it measures

Ability to use information distributed across long contexts (typically 5k–15k+ tokens depending on task).

Input / Output

Input: long documents + task prompt. Output: answers/summaries; task-specific metrics (F1, ROUGE, accuracy).

Evaluation methodology

Per-task metrics then average. Often compared with Needle-in-a-Haystack and RULER for long-context stress tests.

Metadata

Task
Long-context understanding
Domain
QA, summarization, few-shot, code
Modality
Text
Input type
Long document(s) + prompt
Output type
Answer / summary
Evaluation type
Task-specific automatic metrics
Primary metric
Average task score
Secondary metrics
F1, ROUGE, Accuracy
Paper
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
GitHub
THUDM/LongBench
Dataset
LongBench
Leaderboard
Official live leaderboard (external)

Live leaderboard

Rankings change as new evaluations are published. View current results on the official leaderboard.

View live leaderboard →

Datasets

GitHub repositories

Related research

Related guides

Related tools

Related models

Companies

Related rankings

Explore more benchmarks

All benchmarks →