Agentsagentswebbrowserenvironments
WebArena
Realistic web environment benchmark where agents must complete tasks on self-hosted sites (e.g. e-commerce, CMS, forums) via browser actions.
Last reviewed: 23 July 2026
What it measures
Web agency: navigation, form filling, multi-page planning, and success against functional task goals.
Input / Output
Input: natural language task + live site environment. Output: action sequence; success = task completion rate.
Evaluation methodology
Deterministic environment checks for task success. Often paired with VisualWebArena / WorkArena-style suites.
Metadata
- Task
- Browser / web agent tasks
- Domain
- Web automation
- Modality
- Text + web UI
- Input type
- Natural language task + site
- Output type
- Browser action sequence
- Evaluation type
- Functional success rate
- Primary metric
- Task success rate
- Secondary metrics
- Partial credit, Trajectory length
- Paper
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- GitHub
- web-arena-x/webarena
- Dataset
- WebArena
- Leaderboard
- Official live leaderboard (external)
Live leaderboard
Rankings change as new evaluations are published. View current results on the official leaderboard.
View live leaderboard →Datasets
GitHub repositories
Related research
Related guides
Related tools
Related models
Companies
Related rankings
Explore more benchmarks
- GAIAAgentsGeneral AI Assistants benchmark: real-world questions requiring tool use, multi-step reasoning, and web/file interaction — hard for both humans and agents.
- SWE-BenchCodingEvaluates systems on real GitHub issues: given a repository and issue, produce a patch that passes the project’s tests.
- LiveBenchGeneralContamination-free LLM benchmark with frequently updated questions across math, coding, reasoning, language, instruction following, and data analysis.
- Arena HardChatChallenging Chatbot Arena prompts packaged for automatic evaluation — a harder preference-style benchmark derived from real user battles.
- CRAGRAGComprehensive RAG benchmark from Meta stressing retrieval failures, outdated knowledge, and answer faithfulness across realistic industry scenarios.
- GPQAScienceGraduate-Level Google-Proof Q&A: PhD-level science questions designed so that non-experts with web search still struggle, testing deep domain reasoning.