Skip to main content

LLM Benchmarks: Reference Tables by Category

Terminal-Bench, SWE-bench, GPQA Diamond, Arena: scores for major LLMs organized by category, each with its source, vendor-reported or independent status, and measurement date.

  • Reference
  • Tooling
Published

What this page covers

This page gathers, category by category, the scores of major LLMs on the most closely watched benchmarks available today. Every row in every table carries its source, the date the value was measured, and a badge showing whether the number comes from the model's vendor or from an independent measurement. Last global review: July 28, 2026.

No number is copied into the text of this page: they all live in the tables, which stay the single source of truth. The prose explains what each family of benchmarks actually measures and where it breaks down.

Vendor-reported or independent: the distinction that matters

Every row carries a Vendor or Independent badge, along with the publisher's name. The same model can appear twice on the same benchmark, once under each status: that's not redundancy, it's the most useful piece of information on this page. Before comparing two rows, check they carry the same badge first.

Coding and agentic

These benchmarks have the model solve real development tasks: fixing a bug in a repository, getting a test suite to pass, running commands in a terminal. The harness wrapped around the model, meaning the agent that orchestrates its tool calls, shapes the result as much as the model itself does. The same model tested with two different harnesses won't produce the same score: the notes column records the agent used (agent=) whenever the source documents it.

The official SWE-bench Verified leaderboard hasn't taken a new submission since February 2026. The most recent scores in this category therefore come only from vendor cards, not from a verified submission on that specific leaderboard.

Terminal-Bench 2.1
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Fable 5agent=Claude Code · ci=±1.2 · submitted=2026-06-07AnthropicProprietary83.8 %IndependentTerminal-Benchtbench.ai Open the source of the Claude Fable 5 score in a new tab
GPT-5.5agent=Codex · ci=±1.1 · submitted=2026-05-01OpenAIProprietary83.1 %IndependentTerminal-Benchtbench.ai Open the source of the GPT-5.5 score in a new tab
Claude Opus 4.8agent=Claude Code · ci=±1.3 · submitted=2026-07-09AnthropicProprietary78.9 %IndependentTerminal-Benchtbench.ai Open the source of the Claude Opus 4.8 score in a new tab
GPT-5.6 Terraagent=Codex · ci=±1.3 · submitted=2026-07-11OpenAIProprietary78.4 %IndependentTerminal-Benchtbench.ai Open the source of the GPT-5.6 Terra score in a new tab
Muse Spark 1.1agent=mini-SWE-agent · ci=±1.2 · submitted=2026-07-09MetaNot documented76.2 %IndependentTerminal-Benchtbench.ai Open the source of the Muse Spark 1.1 score in a new tab
GPT-5.6 Lunaagent=Codex · ci=±1.3 · submitted=2026-07-11OpenAIProprietary75.7 %IndependentTerminal-Benchtbench.ai Open the source of the GPT-5.6 Luna score in a new tab
Claude Sonnet 5agent=Claude Code · ci=±1.6 · submitted=2026-07-09AnthropicProprietary74.6 %IndependentTerminal-Benchtbench.ai Open the source of the Claude Sonnet 5 score in a new tab
Gemini 3 Proagent=Terminus 2 · ci=±1.3 · submitted=2026-05-01GoogleProprietary73.9 %IndependentTerminal-Benchtbench.ai Open the source of the Gemini 3 Pro score in a new tab
Claude Opus 4.7agent=Claude Code · ci=±1.4 · submitted=2026-05-01AnthropicProprietary68.9 %IndependentTerminal-Benchtbench.ai Open the source of the Claude Opus 4.7 score in a new tab
Gemini 3.1 Proagent=Gemini CLI · ci=±1.7 · submitted=2026-05-05GoogleProprietary65.8 %IndependentTerminal-Benchtbench.ai Open the source of the Gemini 3.1 Pro score in a new tab
Leaderboard maintained by Terminal-Bench (tbench.ai). View the official leaderboard
Terminal-Bench 2.1, replayed by Artificial Analysis
ModelVendorLicenseScoreStatusSourceRecorded on
GPT-5.6 Soleffort=xhigh · agent=Terminus 2 · runs=3OpenAIProprietary89.5 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the GPT-5.6 Sol score in a new tab
Claude Opus 5effort=max · mode=adaptive · agent=Terminus 2 · runs=3AnthropicProprietary89.1 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the Claude Opus 5 score in a new tab
Leaderboard maintained by Artificial Analysis. View the official leaderboard
SWE-bench Verified
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Opus 5effort=max · runs=5 · n=500AnthropicProprietary96 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 5 score in a new tab
Claude Opus 4.5effort=high · agent=mini-SWE-agentAnthropicProprietary76.8 %IndependentSWE-benchswebench.com Open the source of the Claude Opus 4.5 score in a new tab
Gemini 3 Flasheffort=high · agent=mini-SWE-agentGoogleProprietary75.8 %IndependentSWE-benchswebench.com Open the source of the Gemini 3 Flash score in a new tab
Claude Opus 4.6agent=mini-SWE-agentAnthropicProprietary75.6 %IndependentSWE-benchswebench.com Open the source of the Claude Opus 4.6 score in a new tab
GPT-5.2 Codexagent=mini-SWE-agentOpenAIProprietary72.8 %IndependentSWE-benchswebench.com Open the source of the GPT-5.2 Codex score in a new tab
Claude Sonnet 4.5effort=high · agent=mini-SWE-agentAnthropicProprietary71.4 %IndependentSWE-benchswebench.com Open the source of the Claude Sonnet 4.5 score in a new tab
Gemini 3 Proagent=mini-SWE-agentGoogleProprietary69.6 %IndependentSWE-benchswebench.com Open the source of the Gemini 3 Pro score in a new tab
Leaderboard maintained by SWE-bench team. View the official leaderboard
SWE-bench Pro
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Mythos 5AnthropicProprietary80.3 %VendorOpenAIopenai.com Open the source of the Claude Mythos 5 score in a new tab
Claude Fable 5effort=maxAnthropicProprietary80 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Fable 5 score in a new tab
Claude Opus 5effort=max · runs=5AnthropicProprietary79.2 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 5 score in a new tab
Claude Mythos PreviewAnthropicProprietary77.8 %VendorOpenAIopenai.com Open the source of the Claude Mythos Preview score in a new tab
Claude Opus 4.8AnthropicProprietary69.2 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 4.8 score in a new tab
GPT-5.6 SolOpenAIProprietary64.6 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Sol score in a new tab
GPT-5.6 TerraOpenAIProprietary63.4 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Terra score in a new tab
GPT-5.6 LunaOpenAIProprietary62.7 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Luna score in a new tab
GLM-5.2effort=maxZ.aiMIT62.1 %VendorZ.aihuggingface.co Open the source of the GLM-5.2 score in a new tab
Muse Spark 1.1split=public · ci=±3.10MetaNot documented61.5 %IndependentScale AIlabs.scale.com Open the source of the Muse Spark 1.1 score in a new tab
GPT-5.5OpenAIProprietary59.4 %VendorOpenAIopenai.com Open the source of the GPT-5.5 score in a new tab
GPT-5.4split=public · effort=xhigh · ci=±3.56OpenAIProprietary59.1 %IndependentScale AIlabs.scale.com Open the source of the GPT-5.4 score in a new tab
Muse Sparksplit=public · ci=±3.60MetaNot documented55 %IndependentScale AIlabs.scale.com Open the source of the Muse Spark score in a new tab
Gemini 3.1 ProGoogleProprietary54.2 %VendorOpenAIopenai.com Open the source of the Gemini 3.1 Pro score in a new tab
Claude Opus 4.6split=public · mode=thinking · ci=±3.61AnthropicProprietary51.9 %IndependentScale AIlabs.scale.com Open the source of the Claude Opus 4.6 score in a new tab
Gemini 3.1 Prosplit=public · mode=thinking · ci=±3.60GoogleProprietary46.1 %IndependentScale AIlabs.scale.com Open the source of the Gemini 3.1 Pro score in a new tab
Leaderboard maintained by Scale AI. View the official leaderboard

Reasoning

GPQA Diamond and Humanity's Last Exam test expert-level knowledge on closed-ended questions, with no access to external tools. GPQA Diamond is approaching saturation: the most recent models sit within a single point of each other, which no longer separates them meaningfully. Humanity's Last Exam stays more discriminating, but its score depends heavily on the scope used (text-only or the full set) and on the model used to grade open-ended answers, two parameters the notes column records when known.

GPQA Diamond
ModelVendorLicenseScoreStatusSourceRecorded on
GPT-5.6 SolOpenAIProprietary94.6 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Sol score in a new tab
Gemini 3.1 ProGoogleProprietary94.3 %VendorOpenAIopenai.com Open the source of the Gemini 3.1 Pro score in a new tab
Gemini 3.1 ProGoogleProprietary94.1 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the Gemini 3.1 Pro score in a new tab
GPT-5.6 Soleffort=maxOpenAIProprietary94.1 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the GPT-5.6 Sol score in a new tab
Claude Opus 5effort=high · mode=adaptiveAnthropicProprietary93.7 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the Claude Opus 5 score in a new tab
GPT-5.5OpenAIProprietary93.6 %VendorOpenAIopenai.com Open the source of the GPT-5.5 score in a new tab
Kimi K3effort=maxMoonshot AIKimi K3 License93.5 %VendorMoonshot AIhuggingface.co Open the source of the Kimi K3 score in a new tab
GPT-5.6 TerraOpenAIProprietary92.9 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Terra score in a new tab
Claude Fable 5AnthropicProprietary92.6 %VendorOpenAIopenai.com Open the source of the Claude Fable 5 score in a new tab
GPT-5.6 LunaOpenAIProprietary92.3 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Luna score in a new tab
Claude Opus 4.8AnthropicProprietary92 %VendorOpenAIopenai.com Open the source of the Claude Opus 4.8 score in a new tab
GLM-5.2effort=maxZ.aiMIT91.2 %VendorMoonshot AIhuggingface.co Open the source of the GLM-5.2 score in a new tab
Leaderboard maintained by Artificial Analysis. View the official leaderboard
Humanity's Last Exam
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Fable 5tools=offAnthropicProprietary56.5 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Fable 5 score in a new tab
Claude Opus 5tools=off · effort=max · runs=5AnthropicProprietary56.3 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 5 score in a new tab
Claude Fable 5effort=max · fallback=Claude Opus 4.8 · subset=text-only · n=2158AnthropicProprietary53.3 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the Claude Fable 5 score in a new tab
Claude Opus 5effort=max · mode=adaptive · subset=text-only · n=2158AnthropicProprietary52.6 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the Claude Opus 5 score in a new tab
Claude Opus 4.8tools=offAnthropicProprietary49.8 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 4.8 score in a new tab
GPT-5.6 Soltools=off · dataset=fullOpenAIProprietary44.5 %VendorMoonshot AIhuggingface.co Open the source of the GPT-5.6 Sol score in a new tab
Kimi K3tools=off · dataset=full · effort=maxMoonshot AIKimi K3 License43.5 %VendorMoonshot AIhuggingface.co Open the source of the Kimi K3 score in a new tab
GPT-5.5tools=off · dataset=fullOpenAIProprietary41.4 %VendorMoonshot AIhuggingface.co Open the source of the GPT-5.5 score in a new tab
GLM-5.2tools=off · subset=text-onlyZ.aiMIT40.5 %VendorZ.aihuggingface.co Open the source of the GLM-5.2 score in a new tab
Leaderboard maintained by Center for AI Safety, Scale AI. View the official leaderboard

Agents and tools

GDPval-AA, OSWorld, and BrowseComp measure capabilities closer to real-world use: completing a professional task graded by human preference for GDPval-AA, operating a graphical computer interface for OSWorld, browsing and searching the web to track down a precise piece of information for BrowseComp. These agentic benchmarks are more sensitive to safety refusals and automatic fallbacks to another model than a single-answer test is: a refusal or a fallback counts as a failure in the final average, same as an outright model error.

GDPval-AA v2
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Opus 5effort=maxAnthropicProprietary1,861VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 5 score in a new tab
GPT-5.6 SolOpenAIProprietary1,747.8VendorOpenAIopenai.com Open the source of the GPT-5.6 Sol score in a new tab
Claude Fable 5AnthropicProprietary1,747VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Fable 5 score in a new tab
Kimi K3effort=maxMoonshot AIKimi K3 License1,686VendorMoonshot AIhuggingface.co Open the source of the Kimi K3 score in a new tab
Claude Opus 4.8AnthropicProprietary1,593VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 4.8 score in a new tab
GPT-5.6 TerraOpenAIProprietary1,593VendorOpenAIopenai.com Open the source of the GPT-5.6 Terra score in a new tab
GPT-5.6 LunaOpenAIProprietary1,591.8VendorOpenAIopenai.com Open the source of the GPT-5.6 Luna score in a new tab
GLM-5.2effort=maxZ.aiMIT1,510VendorMoonshot AIhuggingface.co Open the source of the GLM-5.2 score in a new tab
GPT-5.5OpenAIProprietary1,493.7VendorOpenAIopenai.com Open the source of the GPT-5.5 score in a new tab
Gemini 3.5 FlashGoogleProprietary1,348.8VendorOpenAIopenai.com Open the source of the Gemini 3.5 Flash score in a new tab
Gemini 3.1 ProGoogleProprietary962.3VendorOpenAIopenai.com Open the source of the Gemini 3.1 Pro score in a new tab
Leaderboard maintained by Artificial Analysis. View the official leaderboard
OSWorld 2.0
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Opus 5effort=maxAnthropicProprietary70.6 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 5 score in a new tab
Claude Fable 5AnthropicProprietary66.1 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Fable 5 score in a new tab
GPT-5.6 SolOpenAIProprietary62.6 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Sol score in a new tab
Kimi K3effort=maxMoonshot AIKimi K3 License58.3 %VendorMoonshot AIhuggingface.co Open the source of the Kimi K3 score in a new tab
Claude Opus 4.8AnthropicProprietary55.7 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 4.8 score in a new tab
GPT-5.6 TerraOpenAIProprietary50.2 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Terra score in a new tab
GPT-5.5OpenAIProprietary47.5 %VendorOpenAIopenai.com Open the source of the GPT-5.5 score in a new tab
GPT-5.6 LunaOpenAIProprietary45.6 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Luna score in a new tab
Leaderboard maintained by OSWorld team, University of Hong Kong. View the official leaderboard
BrowseComp
ModelVendorLicenseScoreStatusSourceRecorded on
Kimi K3effort=max · compaction=300KMoonshot AIKimi K3 License91.2 %VendorMoonshot AIhuggingface.co Open the source of the Kimi K3 score in a new tab
Claude Opus 5effort=maxAnthropicProprietary90.8 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 5 score in a new tab
GPT-5.6 SolOpenAIProprietary90.4 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Sol score in a new tab
GPT-5.6 TerraOpenAIProprietary87.5 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Terra score in a new tab
Claude Fable 5AnthropicProprietary87.4 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Fable 5 score in a new tab
Gemini 3.1 ProGoogleProprietary85.9 %VendorOpenAIopenai.com Open the source of the Gemini 3.1 Pro score in a new tab
GPT-5.5OpenAIProprietary84.4 %VendorOpenAIopenai.com Open the source of the GPT-5.5 score in a new tab
Claude Opus 4.8AnthropicProprietary84.3 %VendorAnthropicwww-cdn.anthropic.com Open the source of the Claude Opus 4.8 score in a new tab
GPT-5.6 LunaOpenAIProprietary83.3 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Luna score in a new tab
Leaderboard maintained by OpenAI. View the official leaderboard

Long context

AA-LCR and OpenAI MRCR test a model's ability to retrieve and reason over information buried in a very long context window, rather than reasoning ability in the abstract. A high score here says nothing about the same model's reasoning quality on a short context, and vice versa. Few vendors publish this kind of number across their entire lineup, which is why these tables run shorter than the other categories.

AA-LCR: long-context reasoning
ModelVendorLicenseScoreStatusSourceRecorded on
GPT-5.2 Codexeffort=xhighOpenAIProprietary75.7 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the GPT-5.2 Codex score in a new tab
GPT-5effort=highOpenAIProprietary75.6 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the GPT-5 score in a new tab
GPT-5.1effort=highOpenAIProprietary75 %IndependentArtificial Analysisartificialanalysis.ai Open the source of the GPT-5.1 score in a new tab
Leaderboard maintained by Artificial Analysis. View the official leaderboard
OpenAI MRCR v2 (8-needle, 256K-512K)
ModelVendorLicenseScoreStatusSourceRecorded on
GPT-5.6 SolOpenAIProprietary91.5 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Sol score in a new tab
GPT-5.6 TerraOpenAIProprietary89.6 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Terra score in a new tab
GPT-5.5OpenAIProprietary81.5 %VendorOpenAIopenai.com Open the source of the GPT-5.5 score in a new tab
GPT-5.6 LunaOpenAIProprietary41.3 %VendorOpenAIopenai.com Open the source of the GPT-5.6 Luna score in a new tab
Leaderboard maintained by OpenAI. View the official leaderboard

Human preference

Text Arena and WebDev Arena rank models by an Elo score built from blind human preference votes, not by automated grading. The confidence interval and the vote count matter more here than the displayed rank: on these leaderboards, several consecutive models often overlap statistically, and a recently added model can show a flattering rank on a still-small vote count. The notes column records the confidence interval (ci=) and the vote count (votes=) whenever the source publishes them.

Text Arena
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Fable 5ci=±6AnthropicProprietary1,508IndependentArenaarena.ai Open the source of the Claude Fable 5 score in a new tab
Claude Opus 4.6mode=thinking · ci=±4AnthropicProprietary1,505IndependentArenaarena.ai Open the source of the Claude Opus 4.6 score in a new tab
Claude Opus 4.7mode=thinking · ci=±4AnthropicProprietary1,502IndependentArenaarena.ai Open the source of the Claude Opus 4.7 score in a new tab
Claude Opus 5effort=max · ci=±12AnthropicProprietary1,495IndependentArenaarena.ai Open the source of the Claude Opus 5 score in a new tab
Muse Spark 1.1ci=±7MetaNot documented1,491IndependentArenaarena.ai Open the source of the Muse Spark 1.1 score in a new tab
Muse Sparkci=±6MetaNot documented1,488IndependentArenaarena.ai Open the source of the Muse Spark score in a new tab
Gemini 3.1 Proci=±3GoogleProprietary1,486IndependentArenaarena.ai Open the source of the Gemini 3.1 Pro score in a new tab
Leaderboard maintained by Arena. View the official leaderboard
WebDev Arena
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Opus 5effort=max · votes=686AnthropicProprietary1,725IndependentArenaarena.ai Open the source of the Claude Opus 5 score in a new tab
Kimi K3effort=max · votes=3777Moonshot AIKimi K3 License1,682IndependentArenaarena.ai Open the source of the Kimi K3 score in a new tab
Claude Fable 5votes=5801AnthropicProprietary1,629IndependentArenaarena.ai Open the source of the Claude Fable 5 score in a new tab
GPT-5.6 Soleffort=xhigh · votes=5340OpenAIProprietary1,623IndependentArenaarena.ai Open the source of the GPT-5.6 Sol score in a new tab
GLM-5.2effort=max · votes=5779Z.aiMIT1,587IndependentArenaarena.ai Open the source of the GLM-5.2 score in a new tab
Claude Opus 4.8mode=thinking · votes=8321AnthropicProprietary1,568IndependentArenaarena.ai Open the source of the Claude Opus 4.8 score in a new tab
Claude Opus 4.7votes=11138AnthropicProprietary1,560IndependentArenaarena.ai Open the source of the Claude Opus 4.7 score in a new tab
Leaderboard maintained by Arena. View the official leaderboard

Composite indices

The Artificial Analysis Intelligence Index aggregates several benchmarks into a single score, weighted according to a published methodology. A composite index simplifies comparison but hides the gaps between categories: two models with a similar index can have very different strength profiles depending on which sub-score you look at. It gives an order of magnitude, not a verdict.

Artificial Analysis Intelligence Index v4.1
ModelVendorLicenseScoreStatusSourceRecorded on
Claude Opus 5effort=maxAnthropicProprietary61IndependentArtificial Analysisartificialanalysis.ai Open the source of the Claude Opus 5 score in a new tab
Claude Fable 5fallback=Claude Opus 4.8AnthropicProprietary60IndependentArtificial Analysisartificialanalysis.ai Open the source of the Claude Fable 5 score in a new tab
GPT-5.6 Soleffort=maxOpenAIProprietary59IndependentArtificial Analysisartificialanalysis.ai Open the source of the GPT-5.6 Sol score in a new tab
Kimi K3Moonshot AIKimi K3 License57IndependentArtificial Analysisartificialanalysis.ai Open the source of the Kimi K3 score in a new tab
GPT-5.6 Terraeffort=maxOpenAIProprietary55IndependentArtificial Analysisartificialanalysis.ai Open the source of the GPT-5.6 Terra score in a new tab
Claude Sonnet 5effort=maxAnthropicProprietary53IndependentArtificial Analysisartificialanalysis.ai Open the source of the Claude Sonnet 5 score in a new tab
GLM-5.2effort=maxZ.aiMIT51IndependentArtificial Analysisartificialanalysis.ai Open the source of the GLM-5.2 score in a new tab
GPT-5.6 Lunaeffort=maxOpenAIProprietary51IndependentArtificial Analysisartificialanalysis.ai Open the source of the GPT-5.6 Luna score in a new tab
Muse Spark 1.1effort=xhighMetaNot documented51IndependentArtificial Analysisartificialanalysis.ai Open the source of the Muse Spark 1.1 score in a new tab
Gemini 3.5 FlashGoogleProprietary50IndependentArtificial Analysisartificialanalysis.ai Open the source of the Gemini 3.5 Flash score in a new tab
Gemini 3.6 FlashGoogleProprietary50IndependentArtificial Analysisartificialanalysis.ai Open the source of the Gemini 3.6 Flash score in a new tab
Gemini 3.1 ProGoogleProprietary46IndependentArtificial Analysisartificialanalysis.ai Open the source of the Gemini 3.1 Pro score in a new tab
Leaderboard maintained by Artificial Analysis. View the official leaderboard

Methodology and limits

This page doesn't cover every pitfall in reading a score: the effect of reasoning effort level, test set contamination, contradictions between an announcement's text and its own table. The article How to Read an LLM Benchmark Without Getting Fooled walks through nine of these pitfalls using verified cases. The scores on this page are refreshed periodically following the editorial team's update procedure; the last global review date appears at the top of the page.

Next steps