| MMLU-Pro (and MMLU) |
Multiple-choice questions across dozens of academic and professional subjects, testing broad knowledge and reasoning. MMLU-Pro expands the original MMLU to ten answer options and adds harder, reasoning-focused questions to better separate strong models. |
The original 2020 MMLU is effectively saturated and contains documented label errors and ambiguous items (prompting fixes such as MMLU-Redux). MMLU-Pro reduces but does not eliminate contamination, and frontier models are now approaching roughly 90%, so its remaining discriminative headroom is shrinking. |
Paper |
| GPQA (Diamond) |
448 graduate-level multiple-choice questions in biology, physics, and chemistry written by domain experts; the 'Diamond' subset is the hardest ~198 items. Designed to be 'Google-proof' so skilled non-experts with web access score low (~34%) while PhD-level experts score much higher. |
The small size makes scores noisy and sensitive to a handful of items. It is a public set with canary strings but not immune to contamination or to models pattern-matching similar training data, and top reasoning models now score in the 80-90%+ range, approaching saturation. |
Paper |
| SWE-bench Verified |
Real GitHub issues from popular Python repositories where a model must produce a code patch that passes the project's hidden tests. Original SWE-bench has 2,294 instances; SWE-bench Verified is a 500-instance human-filtered subset (built with OpenAI) confirmed to be solvable and clearly specified. |
Verified is a curated subset, not a random sample, so it is easier than full SWE-bench and skews toward well-specified fixes. Public repositories raise contamination risk, results depend heavily on the agent scaffold wrapped around the model, and it covers only Python. OpenAI, which helped build it, has stopped reporting it, citing an audit that found flawed tests rejecting correct solutions plus training-data contamination, and points to SWE-bench Pro instead. |
Site
Paper |
| SWE-bench Pro |
1,865 long-horizon software engineering problems from 41 actively maintained repositories spanning business applications, B2B services, and developer tools, often requiring multi-file patches that could take a professional engineer hours to days. Built by Scale AI as a harder, contamination-resistant successor to SWE-bench, it is split into a public set (11 repositories), a held-out set (12 repositories), and a commercial set drawn from 18 proprietary startup codebases; all tasks are human-verified. |
Only the public set can be inspected and run by outsiders, so held-out and commercial results depend on Scale AI's own evaluation, and reported numbers do not always say which split they come from. The public set can still leak into training data over time (Anthropic reports screening a subset of SWE-bench Pro problems for signs of memorization), and as with other agentic coding benchmarks, scores shift with the agent scaffold wrapped around the model. |
Site
Paper |
| FrontierCode |
Whether a model's code is mergeable, not merely test-passing: would the repository's own maintainer accept this pull request? Built by Cognition with the maintainers of 36 flagship open-source repositories, who each spent more than 40 hours per task drawing on their own codebases. Grading uses an ensemble of unit tests, custom rubrics, and new verifiers, scoring correctness, test quality, scope discipline, style, and adherence to codebase standards. The current version is FrontierCode 1.1, which tightened the methodology around what counts as legitimate internet use during a task. |
Cognition does not release the tasks publicly, to limit contamination, so outsiders cannot inspect or reproduce them; it instead opens the evaluation to model creators, and its leaderboard carries a Main table it runs itself plus an Extended table that includes provider self-reports, so check which table a number comes from. Rubric-based grading introduces some subjectivity that pure test-passing benchmarks avoid, and as with every agentic coding benchmark, a score reflects a model-plus-harness combination rather than the model alone. |
Site
|
| DeepSWE |
113 original long-horizon software engineering tasks spanning 91 active open-source projects across five programming languages, run by Datacurve. Unlike benchmarks built from existing issues and pull requests, the tasks are written from scratch and were never contributed back upstream, so their reference solutions stay out of the public record that model training scrapes. Each task is graded by a hand-written verifier that checks the requested functionality and accepts any implementation providing it, rather than reusing a project's own test suite. Datacurve reports that its prompts are shorter than comparable benchmarks while describing solutions that need about 5.5 times more code. The current version is DeepSWE v1.1, which grades an agent's committed diff by applying it in a clean, isolated container and running the tests there, following SWE-bench's approach so an agent cannot game its own environment. |
All models are run on the same harness (mini-swe-agent) for comparability, so the leaderboard shows how a model performs under one common scaffold rather than under its vendor's own agent, and numbers are not interchangeable with vendor-reported figures from other harnesses. The task set is small at 113 items, so scores are noisy and sensitive to individual tasks, and leaderboard coverage is partial - not every model and configuration is published. Contamination resistance rests on the tasks staying unpublished, which is a property of the benchmark's operation rather than something an outsider can verify. |
Site
Paper |
| HumanEval |
164 hand-written Python programming problems where the model generates a function from a docstring and is graded by unit tests (pass@k). It was an early standard for functional code generation. |
Largely legacy and saturated: frontier models cluster around 96-98% pass@1, so it no longer separates leading systems, and its problems are widely believed to be present in training data (contamination). It is now used mainly as a minimum-capability bar, with SWE-bench and live coding benchmarks preferred for frontier comparison. |
Paper |
| MATH and AIME (competition math) |
MATH is 12,500 competition mathematics problems with step-by-step solutions across difficulty levels. Reasoning models are increasingly evaluated on AIME (American Invitational Mathematics Examination) sets such as AIME 2024 and 2025 - 15 integer-answer problems per exam - as a harder, contamination-fresh test. |
MATH is largely saturated for frontier reasoning models. AIME sets are tiny (about 30 problems for a given year), so differences of a few percentage points are within noise, and once a year's problems become public they can leak into training data, which is why the evaluation year matters. |
Paper |
| ARC-AGI, ARC-AGI-2, and ARC-AGI-3 |
Abstract reasoning tests built to resist memorization. ARC-AGI and ARC-AGI-2 (released 2025) are colored-grid puzzles: from a few input-output examples the model must infer a transformation rule and apply it to a new grid, with version 2 emphasizing symbolic interpretation, compositional reasoning, and contextual rule application. ARC-AGI-3 (launched March 2026) is the first fully interactive version: hundreds of handcrafted turn-based game environments with no stated rules or goals, where an agent must explore, work out how each environment works and what winning looks like, and carry that learning across harder levels; a 100% score means beating every game as efficiently as humans. |
Deliberately resists memorization, so it is hard to contaminate, but that also makes it niche and unrepresentative of typical language or coding work. Results are highly sensitive to the evaluation harness and compute budget: ARC Prize verified GPT-6 Astra at 62.7% on ARC-AGI-3 with its standard harness but 99.9% with a provider adapter harness that preserves the model's reasoning state between requests (September 2026). Differing rule sets (public vs. semi-private sets, harness, cost limits) make cross-report comparisons difficult, and ARC Prize says saturating the benchmark would not by itself prove general SI (AGI). |
Site
Paper |
| Chatbot Arena (LMArena / Arena) |
Crowdsourced human-preference ranking: users chat with two anonymous models side by side and vote for the better response, aggregated into an Elo-style leaderboard. Run by the team formerly known as LMSYS Chatbot Arena / LMArena, now operating as Arena (lmarena.ai redirects to arena.ai). |
It measures human preference, not correctness, and has a documented style/verbosity bias toward longer, well-formatted answers (a separate Style Control ranking partially corrects this). Votes come from non-expert, self-selected users, and rankings can be affected by which models are served and how prompts are distributed. |
Site
Paper |
| Humanity's Last Exam (HLE) |
About 2,500 expert-written questions across dozens of subjects, in multiple-choice and short-answer form, each with an unambiguous verifiable answer that cannot be found by quick web retrieval. Built by over 1,000 contributors to probe the frontier of academic knowledge. |
Intended to be extremely hard; as of 2026 frontier models still score well below the roughly 90% expert baseline, and reported numbers vary widely with model, reasoning effort, and tool access. Models tend to be overconfident and poorly calibrated on it, and as a public set it faces the usual future-contamination risk. |
Paper |
| OSWorld and OSWorld 2.0 |
369 real computer tasks executed in a live operating-system environment (files, office apps, browsers, multi-app workflows), graded by execution-based checks on whether the resulting state is correct. A standard test of GUI/computer-use agents; WebArena is a companion benchmark focused specifically on browser tasks. OSWorld 2.0 (released June 2026) is a successor built from 108 long-horizon computer-use workflows that take a human a median of about 1.6 hours each. |
Execution-based grading can be brittle (environment setup, app versions, flaky UI). On the original OSWorld the best agents solved about 12% at release versus about 72% for humans, and scores have since risen sharply, but its tasks are comparatively short, so high numbers overstate long-horizon reliability. OSWorld 2.0 targets that gap: its tasks average about 318 tool calls versus about 30 on the original, and in its launch paper the best agent (Claude Opus 4.8) fully completed only 20.6% of tasks. Results are reported both as strict full completion and as partial credit, and the task set is versioned by release (Anthropic reports scores on the August 2026 release), so check the scoring mode and release before comparing numbers. |
Site
Paper |
| RULER (long context) |
Synthetic long-context benchmark that extends needle-in-a-haystack with 13 tasks in four categories (multi-key/multi-value retrieval, multi-hop variable tracing, aggregation, and question answering) at controlled lengths up to 128K+ tokens, measuring effective rather than advertised context length. |
Because it is synthetic, it may not reflect messy real-world long-document reasoning, and the exact tasks and lengths are configurable, so setups differ between reports. Its central finding - that models degrade well before their claimed context window - continues to hold, which is the reason to use it over simple needle tests. |
Paper |
| MMMU (multimodal) |
About 11,500 college-level multimodal multiple-choice questions across six disciplines and 30 subjects, pairing text with charts, diagrams, tables, chemical structures, medical images and more, testing perception plus subject-specific reasoning in vision-language models. |
The multiple-choice format allows some guessing and text-only shortcutting; a harder MMMU-Pro variant was created to address this. It is a public set (contamination risk), and its answers come from academic materials, so it rewards textbook knowledge more than genuinely novel visual reasoning. |
Site
Paper |
| tau-bench / tau2-bench / tau3-bench (tool-agent-user) |
Simulates multi-turn customer-service conversations in which an agent must use domain APIs (e.g., retail, airline) while following written policies and interacting with a simulated user. It reports a pass^k reliability metric over repeated attempts; tau2-bench extends it to dual-control settings where the user can also take actions, and tau3-bench (March 2026) adds a banking domain that requires retrieving knowledge from documents, plus full-duplex voice evaluation. |
It is domain-specific (a handful of business domains), so it does not capture general tool use, and the simulated user is itself an LLM, which can introduce its own errors. In the original 2024 paper even strong function-calling agents such as GPT-4o succeeded on under 50% of tasks with low run-to-run consistency; top text scores on Sierra's official tau2-bench leaderboard have since reached the high 80s, while the tau3-bench banking knowledge-retrieval domain tops out around 55%, so check which version and domain a reported number covers. |
Site
Paper |
| Terminal-Bench |
Hand-crafted, human-verified tasks executed inside containerized command-line (terminal) environments, spanning software engineering, machine learning, security, system administration, and data science. Each task pairs a natural-language instruction with a Docker environment and a programmatic test suite, scoring an agent on end-to-end task-resolution success rather than isolated code snippets. Terminal-Bench 2.0 comprised 89 tasks; Terminal-Bench 4.0 is the current leaderboard version, which recalibrated per-task time, CPU, and memory budgets, fixed 19 tasks, removed 8 that the latest model generation had saturated, and set a flat 8-hour agent timeout on every task. It is hosted by Stanford, Harbor, and the Laude Institute. |
Execution-based grading in containers can be brittle (environment setup, tool and package versions), and the hand-crafted task set is small, so scores are noisy and sensitive to individual items. Results depend heavily on the agent harness wrapped around the model and on the effort setting used - Anthropic, for instance, reports its own Terminal-Bench 4.0 figures at its highest effort level - and the benchmark is versioned rapidly (1.0, 2.0, 2.1, 3.0, and now 4.0), so the version matters when comparing reported numbers. A companion Terminal-Bench-Science track covers agentic scientific research separately. |
Site
|
| GDPval |
Whether models can produce the actual deliverables of paid professional work. Built by OpenAI, it spans 1,320 tasks covering the majority of US Bureau of Labor Statistics work activities for 44 occupations across the nine sectors contributing most to US GDP. Tasks were written by industry professionals averaging 14 years of experience, and unlike text-only benchmarks they come with reference files and context and expect deliverables such as documents, slides, diagrams, spreadsheets, and multimedia. OpenAI open-sourced a 220-task gold subset plus an automated grader. |
Grading compares model deliverables against expert work rather than checking a single correct answer, so results depend on the grader and on how much context and scaffolding a model is given - OpenAI reports that reasoning effort, task context, and scaffolding each move scores independently. Only the 220-task gold subset is public, so full-set numbers rest on OpenAI's own runs, and third-party variants are not interchangeable with it: Artificial Analysis runs GDPval-AA (v2.1 at time of writing) over those 220 tasks and reports an Elo from blind pairwise LLM-judge comparisons rather than a percentage, which is the figure several labs now cite. |
Site
Paper |
| Vals Index |
A single aggregate score from Vals AI intended to reflect the potential economic impact of SI models, combining eight agentic benchmarks across four professional sectors and weighting each sector by its share of US GDP from Bureau of Economic Analysis value-added-by-industry data. The sectors and their weights are finance (8.0), via Finance Agent v2 and an Excel Modeling Benchmark; coding (5.6), via Terminal-Bench 4.0, Vibe Code Bench, and Code Migration; legal (1.2), via Legal Research Bench and Harvey's Legal Agent Benchmark; and tax (0.5), via Tax Agent Bench. Benchmarks are averaged within a sector, then the sectors are combined on those weights. |
Vals AI itself describes the index as a vast oversimplification of how SI might affect the economy. Six of the eight component benchmarks are private, so outsiders cannot reproduce or audit most of the score. The four sectors sum to roughly 15% of US GDP, so it is a weighted read on a few professional domains rather than an economy-wide measure, and it says nothing about adoption, deployment cost, or the share of a job a model can actually take on. Vals also notes that some Anthropic scores include fallback results after provider refusals, which would lower those scores slightly if counted as failures without changing the rankings. |
Site
|