What this page covers
This page gathers, category by category, the scores of major LLMs on the most closely watched benchmarks available today. Every row in every table carries its source, the date the value was measured, and a badge showing whether the number comes from the model's vendor or from an independent measurement. Last global review: July 28, 2026.
No number is copied into the text of this page: they all live in the tables, which stay the single source of truth. The prose explains what each family of benchmarks actually measures and where it breaks down.
Vendor-reported or independent: the distinction that matters
Before reading the tables
A score published by a model's own vendor, in a system card or an announcement post, hasn't been checked by anyone else. A score published by a third party that doesn't sell any of the compared models (Artificial Analysis, Arena, a benchmark's official maintainer) was measured under conditions the vendor doesn't control. Both have their place, but they don't carry the same weight.
Every row carries a Vendor or Independent badge, along with the publisher's name. The same model can appear twice on the same benchmark, once under each status: that's not redundancy, it's the most useful piece of information on this page. Before comparing two rows, check they carry the same badge first.
Coding and agentic
These benchmarks have the model solve real development tasks: fixing a bug in a repository, getting a test suite to pass, running commands in a terminal. The harness wrapped around the model, meaning the agent that orchestrates its tool calls, shapes the result as much as the model itself does. The same model tested with two different harnesses won't produce the same score: the notes column records the agent used (agent=) whenever the source documents it.
The official SWE-bench Verified leaderboard hasn't taken a new submission since February 2026. The most recent scores in this category therefore come only from vendor cards, not from a verified submission on that specific leaderboard.
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| Claude Fable 5agent=Claude Code · ci=±1.2 · submitted=2026-06-07 | Anthropic | Proprietary | 83.8 % | IndependentTerminal-Bench | tbench.ai Open the source of the Claude Fable 5 score in a new tab | |
| GPT-5.5agent=Codex · ci=±1.1 · submitted=2026-05-01 | OpenAI | Proprietary | 83.1 % | IndependentTerminal-Bench | tbench.ai Open the source of the GPT-5.5 score in a new tab | |
| Claude Opus 4.8agent=Claude Code · ci=±1.3 · submitted=2026-07-09 | Anthropic | Proprietary | 78.9 % | IndependentTerminal-Bench | tbench.ai Open the source of the Claude Opus 4.8 score in a new tab | |
| GPT-5.6 Terraagent=Codex · ci=±1.3 · submitted=2026-07-11 | OpenAI | Proprietary | 78.4 % | IndependentTerminal-Bench | tbench.ai Open the source of the GPT-5.6 Terra score in a new tab | |
| Muse Spark 1.1agent=mini-SWE-agent · ci=±1.2 · submitted=2026-07-09 | Meta | Not documented | 76.2 % | IndependentTerminal-Bench | tbench.ai Open the source of the Muse Spark 1.1 score in a new tab | |
| GPT-5.6 Lunaagent=Codex · ci=±1.3 · submitted=2026-07-11 | OpenAI | Proprietary | 75.7 % | IndependentTerminal-Bench | tbench.ai Open the source of the GPT-5.6 Luna score in a new tab | |
| Claude Sonnet 5agent=Claude Code · ci=±1.6 · submitted=2026-07-09 | Anthropic | Proprietary | 74.6 % | IndependentTerminal-Bench | tbench.ai Open the source of the Claude Sonnet 5 score in a new tab | |
| Gemini 3 Proagent=Terminus 2 · ci=±1.3 · submitted=2026-05-01 | Proprietary | 73.9 % | IndependentTerminal-Bench | tbench.ai Open the source of the Gemini 3 Pro score in a new tab | ||
| Claude Opus 4.7agent=Claude Code · ci=±1.4 · submitted=2026-05-01 | Anthropic | Proprietary | 68.9 % | IndependentTerminal-Bench | tbench.ai Open the source of the Claude Opus 4.7 score in a new tab | |
| Gemini 3.1 Proagent=Gemini CLI · ci=±1.7 · submitted=2026-05-05 | Proprietary | 65.8 % | IndependentTerminal-Bench | tbench.ai Open the source of the Gemini 3.1 Pro score in a new tab |
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| GPT-5.6 Soleffort=xhigh · agent=Terminus 2 · runs=3 | OpenAI | Proprietary | 89.5 % | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the GPT-5.6 Sol score in a new tab | |
| Claude Opus 5effort=max · mode=adaptive · agent=Terminus 2 · runs=3 | Anthropic | Proprietary | 89.1 % | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Claude Opus 5 score in a new tab |
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| Claude Opus 5effort=max · runs=5 · n=500 | Anthropic | Proprietary | 96 % | VendorAnthropic | www-cdn.anthropic.com Open the source of the Claude Opus 5 score in a new tab | |
| Claude Opus 4.5effort=high · agent=mini-SWE-agent | Anthropic | Proprietary | 76.8 % | IndependentSWE-bench | swebench.com Open the source of the Claude Opus 4.5 score in a new tab | |
| Gemini 3 Flasheffort=high · agent=mini-SWE-agent | Proprietary | 75.8 % | IndependentSWE-bench | swebench.com Open the source of the Gemini 3 Flash score in a new tab | ||
| Claude Opus 4.6agent=mini-SWE-agent | Anthropic | Proprietary | 75.6 % | IndependentSWE-bench | swebench.com Open the source of the Claude Opus 4.6 score in a new tab | |
| GPT-5.2 Codexagent=mini-SWE-agent | OpenAI | Proprietary | 72.8 % | IndependentSWE-bench | swebench.com Open the source of the GPT-5.2 Codex score in a new tab | |
| Claude Sonnet 4.5effort=high · agent=mini-SWE-agent | Anthropic | Proprietary | 71.4 % | IndependentSWE-bench | swebench.com Open the source of the Claude Sonnet 4.5 score in a new tab | |
| Gemini 3 Proagent=mini-SWE-agent | Proprietary | 69.6 % | IndependentSWE-bench | swebench.com Open the source of the Gemini 3 Pro score in a new tab |
Reasoning
GPQA Diamond and Humanity's Last Exam test expert-level knowledge on closed-ended questions, with no access to external tools. GPQA Diamond is approaching saturation: the most recent models sit within a single point of each other, which no longer separates them meaningfully. Humanity's Last Exam stays more discriminating, but its score depends heavily on the scope used (text-only or the full set) and on the model used to grade open-ended answers, two parameters the notes column records when known.
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| Claude Fable 5tools=off | Anthropic | Proprietary | 56.5 % | VendorAnthropic | www-cdn.anthropic.com Open the source of the Claude Fable 5 score in a new tab | |
| Claude Opus 5tools=off · effort=max · runs=5 | Anthropic | Proprietary | 56.3 % | VendorAnthropic | www-cdn.anthropic.com Open the source of the Claude Opus 5 score in a new tab | |
| Claude Fable 5effort=max · fallback=Claude Opus 4.8 · subset=text-only · n=2158 | Anthropic | Proprietary | 53.3 % | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Claude Fable 5 score in a new tab | |
| Claude Opus 5effort=max · mode=adaptive · subset=text-only · n=2158 | Anthropic | Proprietary | 52.6 % | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Claude Opus 5 score in a new tab | |
| Claude Opus 4.8tools=off | Anthropic | Proprietary | 49.8 % | VendorAnthropic | www-cdn.anthropic.com Open the source of the Claude Opus 4.8 score in a new tab | |
| GPT-5.6 Soltools=off · dataset=full | OpenAI | Proprietary | 44.5 % | VendorMoonshot AI | huggingface.co Open the source of the GPT-5.6 Sol score in a new tab | |
| Kimi K3tools=off · dataset=full · effort=max | Moonshot AI | Kimi K3 License | 43.5 % | VendorMoonshot AI | huggingface.co Open the source of the Kimi K3 score in a new tab | |
| GPT-5.5tools=off · dataset=full | OpenAI | Proprietary | 41.4 % | VendorMoonshot AI | huggingface.co Open the source of the GPT-5.5 score in a new tab | |
| GLM-5.2tools=off · subset=text-only | Z.ai | MIT | 40.5 % | VendorZ.ai | huggingface.co Open the source of the GLM-5.2 score in a new tab |
Agents and tools
GDPval-AA, OSWorld, and BrowseComp measure capabilities closer to real-world use: completing a professional task graded by human preference for GDPval-AA, operating a graphical computer interface for OSWorld, browsing and searching the web to track down a precise piece of information for BrowseComp. These agentic benchmarks are more sensitive to safety refusals and automatic fallbacks to another model than a single-answer test is: a refusal or a fallback counts as a failure in the final average, same as an outright model error.
Long context
AA-LCR and OpenAI MRCR test a model's ability to retrieve and reason over information buried in a very long context window, rather than reasoning ability in the abstract. A high score here says nothing about the same model's reasoning quality on a short context, and vice versa. Few vendors publish this kind of number across their entire lineup, which is why these tables run shorter than the other categories.
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| GPT-5.2 Codexeffort=xhigh | OpenAI | Proprietary | 75.7 % | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the GPT-5.2 Codex score in a new tab | |
| GPT-5effort=high | OpenAI | Proprietary | 75.6 % | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the GPT-5 score in a new tab | |
| GPT-5.1effort=high | OpenAI | Proprietary | 75 % | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the GPT-5.1 score in a new tab |
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | Proprietary | 91.5 % | VendorOpenAI | openai.com Open the source of the GPT-5.6 Sol score in a new tab | |
| GPT-5.6 Terra | OpenAI | Proprietary | 89.6 % | VendorOpenAI | openai.com Open the source of the GPT-5.6 Terra score in a new tab | |
| GPT-5.5 | OpenAI | Proprietary | 81.5 % | VendorOpenAI | openai.com Open the source of the GPT-5.5 score in a new tab | |
| GPT-5.6 Luna | OpenAI | Proprietary | 41.3 % | VendorOpenAI | openai.com Open the source of the GPT-5.6 Luna score in a new tab |
Human preference
Text Arena and WebDev Arena rank models by an Elo score built from blind human preference votes, not by automated grading. The confidence interval and the vote count matter more here than the displayed rank: on these leaderboards, several consecutive models often overlap statistically, and a recently added model can show a flattering rank on a still-small vote count. The notes column records the confidence interval (ci=) and the vote count (votes=) whenever the source publishes them.
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| Claude Fable 5ci=±6 | Anthropic | Proprietary | 1,508 | IndependentArena | arena.ai Open the source of the Claude Fable 5 score in a new tab | |
| Claude Opus 4.6mode=thinking · ci=±4 | Anthropic | Proprietary | 1,505 | IndependentArena | arena.ai Open the source of the Claude Opus 4.6 score in a new tab | |
| Claude Opus 4.7mode=thinking · ci=±4 | Anthropic | Proprietary | 1,502 | IndependentArena | arena.ai Open the source of the Claude Opus 4.7 score in a new tab | |
| Claude Opus 5effort=max · ci=±12 | Anthropic | Proprietary | 1,495 | IndependentArena | arena.ai Open the source of the Claude Opus 5 score in a new tab | |
| Muse Spark 1.1ci=±7 | Meta | Not documented | 1,491 | IndependentArena | arena.ai Open the source of the Muse Spark 1.1 score in a new tab | |
| Muse Sparkci=±6 | Meta | Not documented | 1,488 | IndependentArena | arena.ai Open the source of the Muse Spark score in a new tab | |
| Gemini 3.1 Proci=±3 | Proprietary | 1,486 | IndependentArena | arena.ai Open the source of the Gemini 3.1 Pro score in a new tab |
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| Claude Opus 5effort=max · votes=686 | Anthropic | Proprietary | 1,725 | IndependentArena | arena.ai Open the source of the Claude Opus 5 score in a new tab | |
| Kimi K3effort=max · votes=3777 | Moonshot AI | Kimi K3 License | 1,682 | IndependentArena | arena.ai Open the source of the Kimi K3 score in a new tab | |
| Claude Fable 5votes=5801 | Anthropic | Proprietary | 1,629 | IndependentArena | arena.ai Open the source of the Claude Fable 5 score in a new tab | |
| GPT-5.6 Soleffort=xhigh · votes=5340 | OpenAI | Proprietary | 1,623 | IndependentArena | arena.ai Open the source of the GPT-5.6 Sol score in a new tab | |
| GLM-5.2effort=max · votes=5779 | Z.ai | MIT | 1,587 | IndependentArena | arena.ai Open the source of the GLM-5.2 score in a new tab | |
| Claude Opus 4.8mode=thinking · votes=8321 | Anthropic | Proprietary | 1,568 | IndependentArena | arena.ai Open the source of the Claude Opus 4.8 score in a new tab | |
| Claude Opus 4.7votes=11138 | Anthropic | Proprietary | 1,560 | IndependentArena | arena.ai Open the source of the Claude Opus 4.7 score in a new tab |
Composite indices
The Artificial Analysis Intelligence Index aggregates several benchmarks into a single score, weighted according to a published methodology. A composite index simplifies comparison but hides the gaps between categories: two models with a similar index can have very different strength profiles depending on which sub-score you look at. It gives an order of magnitude, not a verdict.
| Model | Vendor | License | Score | Status | Source | Recorded on |
|---|---|---|---|---|---|---|
| Claude Opus 5effort=max | Anthropic | Proprietary | 61 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Claude Opus 5 score in a new tab | |
| Claude Fable 5fallback=Claude Opus 4.8 | Anthropic | Proprietary | 60 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Claude Fable 5 score in a new tab | |
| GPT-5.6 Soleffort=max | OpenAI | Proprietary | 59 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the GPT-5.6 Sol score in a new tab | |
| Kimi K3 | Moonshot AI | Kimi K3 License | 57 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Kimi K3 score in a new tab | |
| GPT-5.6 Terraeffort=max | OpenAI | Proprietary | 55 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the GPT-5.6 Terra score in a new tab | |
| Claude Sonnet 5effort=max | Anthropic | Proprietary | 53 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Claude Sonnet 5 score in a new tab | |
| GLM-5.2effort=max | Z.ai | MIT | 51 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the GLM-5.2 score in a new tab | |
| GPT-5.6 Lunaeffort=max | OpenAI | Proprietary | 51 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the GPT-5.6 Luna score in a new tab | |
| Muse Spark 1.1effort=xhigh | Meta | Not documented | 51 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Muse Spark 1.1 score in a new tab | |
| Gemini 3.5 Flash | Proprietary | 50 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Gemini 3.5 Flash score in a new tab | ||
| Gemini 3.6 Flash | Proprietary | 50 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Gemini 3.6 Flash score in a new tab | ||
| Gemini 3.1 Pro | Proprietary | 46 | IndependentArtificial Analysis | artificialanalysis.ai Open the source of the Gemini 3.1 Pro score in a new tab |
Methodology and limits
This page doesn't cover every pitfall in reading a score: the effect of reasoning effort level, test set contamination, contradictions between an announcement's text and its own table. The article How to Read an LLM Benchmark Without Getting Fooled walks through nine of these pitfalls using verified cases. The scores on this page are refreshed periodically following the editorial team's update procedure; the last global review date appears at the top of the page.
Next steps
- How to Read an LLM Benchmark Without Getting Fooled: The Survival Guide: the methodological pitfalls behind the numbers on this page
- Kimi K3: the first open model in the 3T class, and what actually changes: a concrete case where several of these benchmarks come into play
- Opus 5, GPT-5.6, Kimi K3: three launches, three rankings that contradict each other: a comparative read of the most disputed scores on this page
- The real costs of Claude Code: putting these scores in perspective against what they actually cost to run