Anthropic
Claude Opus 5.5
Hard reasoning, long agentic work, anything that has to be right
Anthropic's flagship. Leads the index with room to spare and is also the top model on LMArena.
Published evals HLE 61.4% · SciCode 66.9% · AA-LCR 84.7%
AZ Labs rankings · Updated 3 October 2026
Ranked purely on independent benchmarks: the Artificial Analysis Intelligence Index, 10 evaluations covering agents, coding, science and knowledge. No vendor claims, no paid placement, no votes.
The ranking. 100% Artificial Analysis Intelligence Index.
Anthropic
Hard reasoning, long agentic work, anything that has to be right
Anthropic's flagship. Leads the index with room to spare and is also the top model on LMArena.
Published evals HLE 61.4% · SciCode 66.9% · AA-LCR 84.7%
Anthropic
Scores 56 on the Intelligence Index.
Published evals HLE 55.0% · SciCode 61.0% · AA-LCR 82.7%
Anthropic
Writing, nuance and long-form work
Anthropic's premium writing-focused model. Near the top on benchmarks, and among the most expensive per token.
Published evals HLE 59.1% · SciCode 63.1% · AA-LCR 85.3% · Coding Index 81.6
OpenAI's frontier reasoning model. OpenAI's strongest model, neck and neck with Fable 5.1 on the index.
Published evals HLE 54.7% · SciCode 56.5% · AA-LCR 80.7% · Coding Index 76.9
Scores 52.6 on the Intelligence Index.
Published evals HLE 57.1% · SciCode 61.8% · AA-LCR 79.7%
Scores 51.8 on the Intelligence Index.
Published evals HLE 52.9% · SciCode 54.2% · AA-LCR 83.0%
Near-frontier answers at high speed. Meta's best. One of the fastest models near the top of the board, at a low per-token price.
Published evals HLE 48.7% · SciCode 58.8% · AA-LCR 83.0% · Coding Index 75.8
GPT-6 reasoning at a lower price. The mid-tier GPT-6 model: a few points behind Astra at a fraction of the price.
Published evals HLE 47.9% · SciCode 57.6% · AA-LCR 83.7%
The strongest model in the xAI stack. xAI's newest flagship. LMArena voters rate it well below its benchmark score.
Published evals HLE 43.1% · SciCode 57.4% · AA-LCR 76.7%
Absurd value for an open model. The top open-weights model on the index, MIT-licensed, and priced below every other model in the top ten.
Published evals HLE 49.4% · SciCode 60.9% · AA-LCR 86.3%
Picked from the benchmark top ten.
Claude Sonnet 5.5
#2 at $4
Highest-ranked model with a blended API price under $5 per million tokens.
Muse Spark 1.3
162 tok/s
Highest median output speed in this top ten.
MiMo-V2.6-Pro
$0.54 / 1M
Cheapest blended API price (3:1 input to output) in this top ten.
MiMo-V2.6-Pro
#10 overall
Highest-ranked model you can download and self-host.
The ranking uses the horizontal axis only. The vertical axis shows how LMArena voters rate the same models, so you can see where the two disagree. Hover a point for details.
The next five current models on the index.
| Rank | Model | AA index | Arena | Price / 1M | Note |
|---|---|---|---|---|---|
| #11 | Qwen3.8 MaxAlibaba | 45.4 | 1479 | $3 | Alibaba's flagship. Strong on benchmarks but one of the slower models in the list. |
| #12 | GLM-5.3Z AI | 44.8 | 1480 | $2.15 | Z.ai's flagship, MIT-licensed, with no weak spot on either leaderboard. |
| #13 | Grok 4.6SpaceXAI | 44.3 | 1453 | $3 | The previous Grok flagship, still close behind its successor on the index. |
| #14 | Step 5 PreviewStepFun | 43.7 | — | $1.43 | StepFun's preview release. A competitive score at a low price; not on LMArena yet. |
| #15 | Kimi K3Kimi | 43.6 | 1488 | $6 | Moonshot's open-weights flagship. Voters rank it above every other open model, but it is slow. |
Scores, prices and speeds: Artificial Analysis Data API via AIMI, fetched 2026-10-03. Arena scores: LMArena, 2026-09-25. 42 current models tracked.
Every current model Artificial Analysis publishes, pulled from its Data API into AIMI daily and on every new release. 42 are tracked here; superseded versions are excluded.
Most models ship several reasoning settings (low, high, max). Each model is scored at its strongest listed setting, and we show which setting that was.
The published score is used as-is. No normalising, no weighting of our own, no blending with other sources.
Scores come to one decimal, so ties are rare; an exact tie ranks the newer release first. Votes, price and speed never change the order.
Weighted average of 10 independent evaluations run by Artificial Analysis: agents 30%, coding 20%, scientific reasoning 20%, general 30%. Pulled from the Artificial Analysis Data API into AIMI daily and whenever a new model appears. This is the only input to the ranking.
Elo-style score from 8.5M+ blind head-to-head votes, leaderboard dated 25 Sep 2026. Shown for context only; it does not affect the ranking.
The 10 evaluations behind every score on this page, with the weight each carries in index v4.3.2. Private sets are held back by Artificial Analysis so models cannot be trained on them. Full methodology
| Category | Evaluation | Weight | What it tests | Size | Scoring |
|---|---|---|---|---|---|
| Agents | AA-Briefcase v1.1Private | 15% | Agentic knowledge work that ends in real file deliverables | 91 tasks, 4 scenarios | Elo from pairwise comparisons of task success, analysis and presentation |
| Agents | GDPval-AA v2.1 | 10% | Economically valuable professional tasks with file outputs | 220 tasks | Pairwise Elo by a judge panel |
| Agents | AutomationBench-AAPrivate | 5% | SaaS workflow automation through REST API tools | 657 tasks | Task completion; zero credit if a guardrail is violated |
| Coding | Terminal-Bench 4.0 | 10% | Real tasks executed in a terminal | 66 tasks × 3 runs | Test suite pass/fail, pass@1 |
| Coding | SciCode | 10% | Scientific Python that must pass every unit test | 288 subproblems × 3 runs | Code execution, pass@1 |
| General | AA-OmnisciencePrivate | 15% | Knowledge accuracy and how often the model makes things up | 6,000 questions | Accuracy (10%) plus 1 − hallucination rate (5%) |
| General | GDP.pdf | 10% | Answers grounded in long PDF documents across 10 domains | 100 tasks × 5 runs | All-pass headline and mean pass rate |
| General | AA-LCR v1.1 | 5% | Long-context reasoning over very large inputs | 100 questions × 3 runs | LLM equality checker, pass@1 |
| Scientific reasoning | Humanity's Last Exam | 10% | Expert-level questions across academic fields | 2,158 questions | LLM equality checker, pass@1 |
| Scientific reasoning | CritPtPrivate | 10% | Research-level physics problems | 70 problems × 5 runs | Official grading server, pass@1 |
On our 3 October 2026 snapshot, Claude Opus 5.5 is the best model overall. It scores 57.6 on the Artificial Analysis Intelligence Index v4.3.2, ahead of Claude Sonnet 5.5 on 56.0.
Only the Artificial Analysis Intelligence Index v4.3.2. It is a weighted average of 10 independent evaluations: agents 30% (AA-Briefcase, GDPval-AA, AutomationBench-AA), coding 20% (Terminal-Bench 4.0, SciCode), general 30% (AA-Omniscience, GDP.pdf, AA-LCR) and scientific reasoning 20% (Humanity's Last Exam, CritPt). We take each model at its best reasoning setting and sort by that score.
The Artificial Analysis Data API. Our model catalogue, AIMI, pulls it daily and as soon as it spots a new model release, keeps the raw response as evidence, records every score with its date, and publishes the result straight to this page. No scores are typed in by hand.
Benchmarks measure whether a model gets hard problems right. Blind votes show whether people prefer its answers. The two often disagree: Gemini 3.8 Flash is popular with voters but ranks #19 on benchmarks. So votes are shown as context and kept out of the ranking.
MiMo-V2.6-Pro, at #10 with 46.3 on the index.
Automatically. AIMI checks Artificial Analysis once a day, and within about half an hour of spotting a new model release. When anything changes, this page updates on its own. The date at the top shows when the data was last fetched.