Best AI models for coding.
Which models write working code: independent coding evaluations combined with blind votes on web apps the models built themselves. As of 23 Sep 2026, Claude Fable 5.1 leads for coding, ahead of GPT-6 Astra and Claude Opus 5, out of 42 ranked models. The highest-ranked open-weights model is Kimi K3 (5th). The cheapest model in the top ten is Gemini 3.7 Flash, at A$2.11 per million tokens.
| MODEL | IN AUSTRALIA | |||||
|---|---|---|---|---|---|---|
| 01Claude Fable 5.1ANTHROPIC | 97.1 | 81.6 | 1755 | 91.4% | A$28.08 | GLOBAL |
| 02GPT-6 AstraOPENAI | 92.6 | 76.9 | 1793 | 88.4% | A$28.08 | GLOBAL |
| 03Claude Opus 5ANTHROPIC | 86.6 | 78.0 | 1691 | 89.1% | A$14.04 | IN AU |
| 04Qwen3.8 MaxALIBABA | 82.3 | 76.2 | 1671 | 88.8% | A$4.21 | NOT IN AU |
| 05Kimi K3MOONSHOT AI | 81.3 | 76.2 | 1658 | 85.0% | A$8.42 | GLOBAL |
| 06Muse Spark 1.3META | 80.6 | 75.8 | 1657 | 84.3% | A$2.81 | NOT IN AU |
| 07GPT-5.6 SolOPENAI | 80.1 | 77.4 | 1617 | 88.0% | A$11.23 | GLOBAL |
| 08Claude Fable 5ANTHROPIC | 79.4 | 76.5 | 1627 | 84.6% | A$28.08 | GLOBAL |
| 09Grok 4.6XAI | 79.1 | 76.8 | 1616 | 88.4% | A$4.21 | GLOBAL |
| 10Gemini 3.7 FlashGOOGLE | 76.4 | 76.1 | 1595 | 85.8% | A$2.11 | GLOBAL |
| 11GLM-5.3Z.AI | 76.2 | 74.8 | 1620 | 83.9% | A$3.02 | SELF-HOST |
| 12Gemini 3.8 FlashGOOGLE | 75.8 | 76.3 | 1583 | 87.6% | A$2.11 | GLOBAL |
| 13Qwen3.8 Flash NextALIBABA | 74.7 | 73.1 | 1636 | 86.1% | A$0.32 | SELF-HOST |
| 14GPT-5.6 TerraOPENAI | 71.5 | 76.7 | 1518 | 88.0% | A$6.32 | GLOBAL |
| 15Claude Opus 4.8ANTHROPIC | 70.6 | 74.3 | 1556 | 84.6% | A$14.04 | IN AU |
| 16GLM-5.3 FlashZ.AI | 70.4 | 71.5 | 1612 | 84.3% | A$0.33 | SELF-HOST |
| 17Claude Opus 4.7ANTHROPIC | 69.6 | 73.6 | 1558 | 83.1% | A$14.04 | IN AU |
| 18GPT-5.5OPENAI | 68.1 | 74.9 | 1510 | 84.3% | A$15.79 | GLOBAL |
| 19Grok 4.5XAI | 67.3 | 72.4 | 1552 | 81.6% | A$4.21 | NOT IN AU |
| 20Muse Spark 1.2META | 65.7 | 72.2 | 1534 | 80.1% | A$2.81 | NOT IN AU |
| 21GLM-5.2Z.AI | 65.1 | 68.8 | 1598 | 77.9% | A$3.02 | SELF-HOST |
| 22Claude Sonnet 5ANTHROPIC | 64.8 | 71.5 | 1538 | 80.5% | A$5.62 | IN AU |
| 23Muse Spark 1.1META | 64.8 | 71.3 | 1542 | 77.9% | A$2.81 | |
| 24DeepSeek V4 FlashDEEPSEEK | 64.2 | 69.1 | 1580 | 78.7% | A$0.93 | GLOBAL |
| 25DeepSeek V4 ProDEEPSEEK | 63.8 | 68.8 | 1580 | 78.7% | A$2.78 | GLOBAL |
| 26Qwen3.8 27BALIBABA | 63.5 | 68.1 | 1591 | 79.8% | A$1.58 | SELF-HOST |
| 27GPT-5.6 LunaOPENAI | 63.3 | 71.4 | 1520 | 80.9% | A$0.63 | GLOBAL |
| 28Gemini 3.6 FlashGOOGLE | 61.1 | 69.2 | 1537 | 77.5% | A$2.11 | GLOBAL |
| 29Gemini 3.5 FlashGOOGLE | 59.8 | 70.1 | 1500 | 78.7% | A$4.74 | IN AU |
| 30GPT-5.4OPENAI | 58.5 | 71.1 | 1463 | 78.3% | A$7.90 | IN AU |
| 31Qwen3.7 MaxALIBABA | 54.6 | 66.0 | 1517 | 74.5% | A$5.26 | |
| 32Gemini 3.1 ProGOOGLE | 53.7 | 68.8 | 1447 | 73.8% | A$6.32 | GLOBAL |
| 33Claude Sonnet 4.6ANTHROPIC | 50.1 | 63.0 | 1521 | 71.2% | A$8.42 | IN AU |
| 34Kimi K2.6MOONSHOT AI | 47.4 | 61.8 | 1509 | 65.9% | A$2.40 | SELF-HOST |
| 35Hy3TENCENT | 42.7 | 58.8 | 1509 | 64.4% | A$0.35 | SELF-HOST |
| 36MiMo V2.5 ProXIAOMI | 42.4 | 60.2 | 1477 | 65.2% | A$0.76 | SELF-HOST |
| 37MiniMax M3MINIMAX | 40.6 | 58.6 | 1485 | 65.2% | A$0.74 | SELF-HOST |
| 38Muse SparkMETAPARTIAL | 39.8 | 58.6 | 62.2% | |||
| 39GPT-5.4 miniOPENAI | 30.0 | 56.1 | 1397 | 59.2% | A$2.37 | |
| 40Gemini 3.5 Flash-LiteGOOGLE | 23.1 | 49.3 | 1447 | 53.6% | A$1.19 | GLOBAL |
| 41Gemma 4 31BGOOGLE | 7.5 | 43.4 | 1364 | 43.4% | A$0.29 | SELF-HOST |
| 42Mistral Medium 3.5MISTRAL AI | 5.5 | 46.9 | 1265 | 50.6% | A$4.21 | SELF-HOST |
What goes into the coding score.
The Coding Index covers agentic terminal work and scientific programming; the WebDev arena adds how people rate finished, running apps. Terminal-Bench is shown alongside because it is the closest test to how coding agents are actually used, but it is already inside the Coding Index, so it isn’t counted twice.
Artificial Analysis’s composite of its coding evaluations, including Terminal-Bench and SciCode.
Blind votes on working web apps each model builds from the same prompt.
Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →
Choosing a model for coding.
What is the best AI model for coding right now?
As of 23 Sep 2026, Claude Fable 5.1 ranks first, followed by GPT-6 Astra and Claude Opus 5. The ranking combines AA Coding Index and LMArena WebDev and is refreshed weekly.
What is the best open-weights model for coding?
Kimi K3 from Moonshot AI is the highest-ranked open-weights model (5th of 42). Open weights can be run on your own infrastructure, including in an Australian region.
Which top coding model is cheapest?
Of the top ten, Gemini 3.7 Flash is cheapest at A$2.11 per million tokens (blended), against A$28.08 for Claude Fable 5.1. Prices are list prices converted at A$1 = US$0.7123, excluding GST.
Which of these models can keep data in Australia?
Of the top ten, Claude Opus 5 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.
How is this ranking calculated?
The Coding Index covers agentic terminal work and scientific programming; the WebDev arena adds how people rate finished, running apps. Terminal-Bench is shown alongside because it is the closest test to how coding agents are actually used, but it is already inside the Coding Index, so it isn’t counted twice. Each input is rescaled 0–100 across the models we track, then weighted: AA Coding Index 60% and LMArena WebDev 40%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →