Frontier models · snapshot 23 September 2026
Which AI model is best at what?
The most capable AI models right now, ranked on one independent benchmark and then split by task, because no single model wins everything. Every number links to where it came from.
Best at each job
Overall intelligence
Claude Opus 5.5
Highest Artificial Analysis Intelligence Index score at 58, five points clear of GPT-6 Astra and Claude Fable 5.1 (both 53).
Independent measurement
Agentic coding
Claude Opus 5.5
Terminal-Bench 4.0: 66.4% against 57.9% for GPT-6 Astra and 55.8% for Fable 5.1. Before Opus 5.5 arrived, Artificial Analysis had GPT-6 Astra and Fable 5.1 tied at 62 on its Coding Agent Index, with Astra doing it at about 60% of the cost.
Vendor-reported benchmark
Office and knowledge work
Claude Opus 5.5
GDPval-AA v2.1, which scores real professional deliverables: 1846 Elo against 1735 for Fable 5.1 and 1542 for GPT-6 Astra.
Vendor-reported benchmark
Business automation
GPT-6 Astra
AutomationBench: 41.4%, narrowly ahead of Opus 5.5 at 40.0%. One of the few results in Anthropic’s own table that Anthropic does not win.
Vendor-reported benchmark
Speed and price
Gemini 3.8 Flash
286.9 output tokens per second at $0.75 / $3.75 per million tokens — around five times faster than the top three and a fraction of their price, for a lower index score (41).
Independent measurement
Best model you can download
MiMo-V2.6-Pro
Highest-scoring open-weights model at 46, ahead of GLM-5.3 (45) and Kimi K3 (44). At about one trillion parameters it needs a server, not a desktop PC.
Independent measurement
Overall ranking
Artificial Analysis Intelligence Index v4.3.2 combines ten evaluations — coding, reasoning, science, knowledge work and tool use — which Artificial Analysis runs itself. Each model is shown at its highest-effort setting. Prices are US dollars per million input / output tokens.
| # | Model | Released | Index | Price | Speed | Weights |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5Anthropic | 22 Sep 2026 | 58 | $4 / $20 | Not yet measured | Hosted only |
| 2 | Claude Fable 5.1Anthropic | 1 Sep 2026 | 53 | $10 / $50 | 65.7 tokens/s | Hosted only |
| 3 | GPT-6 AstraOpenAI | 3 Sep 2026 | 53 | $10 / $50 | 58.1 tokens/s | Hosted only |
| 4 | MiMo-V2.6-ProXiaomi | — | 46 | — | — | Downloadable |
| 5 | GLM-5.3Z AI | — | 45 | — | — | Downloadable |
| 6 | Grok 4.6xAI | 12 Aug 2026 | 44 | $2 / $6 | 58.6 tokens/s | Hosted only |
| 7 | Kimi K3Moonshot AI (Kimi) | — | 44 | — | — | Downloadable |
| 8 | Gemini 3.8 FlashGoogle | 2 Sep 2026 | 41 | $0.75 / $3.75 | 286.9 tokens/s | Hosted only |
“—” means the figure is not published on the source page, not that it is zero. Scores from different index versions are not comparable, so this table uses only v4.3.2.
Head to head: Anthropic’s launch figures
From Anthropic’s Opus 5.5 announcement. These are vendor-reported: the company releasing the model chose the tests and ran them. Treat them as a claim to check against independent results, not as a verdict.
| Benchmark | Opus 5.5 | Fable 5.1 | GPT-6 Astra | Opus 5 |
|---|---|---|---|---|
| Terminal-Bench 4.0Agentic coding in a terminal | 66.4% | 55.8% | 57.9% | 52.3% |
| FrontierCode v1.1Hard software tasks | 54.4% | 50.3% | 53.3% | 48.0% |
| GDPval-AA v2.1Professional deliverables (Elo) | 1846 | 1735 | 1542 | 1708 |
| AutomationBenchBusiness workflows | 40.0% | 31.4% | 41.4% | 26.9% |
| Humanity’s Last ExamExpert reasoning, with tools | 67.7% | 65.6% | 57.2% | 63.6% |
| OSWorld 2.0Operating a computer | 81.8% | 80.7% | Not reported | 74.0% |
What this means if you run AI on your own PC
None of the top three can be downloaded; they only run on their makers’ servers. The best open-weights models are close behind, but at 750 billion to 2.8 trillion parameters they need data-centre hardware. What runs on a home graphics card is smaller models from the same families — and whether a given one fits depends on your video memory, not on this ranking.
Sources
All figures retrieved 23 September 2026. Leaderboards change as models are added and re-tested.