Best AI models for coding.

Which models write working code: independent coding evaluations combined with blind votes on web apps the models built themselves. As of 23 Sep 2026, Claude Fable 5.1 leads for coding, ahead of GPT-6 Astra and Claude Opus 5, out of 42 ranked models. The highest-ranked open-weights model is Kimi K3 (5th). The cheapest model in the top ten is Gemini 3.7 Flash, at A$2.11 per million tokens.

42 MODELS RANKEDA$1 = US$0.7123UPDATED 23 SEP 2026
CODING RANKING42 MODELS · 23 SEP 2026
42 MODELS
Best AI models for coding, ranked, as of 23 Sep 2026
MODELIN AUSTRALIA
01Claude Fable 5.1ANTHROPIC97.181.6175591.4%A$28.08GLOBAL
02GPT-6 AstraOPENAI92.676.9179388.4%A$28.08GLOBAL
03Claude Opus 5ANTHROPIC86.678.0169189.1%A$14.04IN AU
04Qwen3.8 MaxALIBABA82.376.2167188.8%A$4.21NOT IN AU
05Kimi K3MOONSHOT AI81.376.2165885.0%A$8.42GLOBAL
06Muse Spark 1.3META80.675.8165784.3%A$2.81NOT IN AU
07GPT-5.6 SolOPENAI80.177.4161788.0%A$11.23GLOBAL
08Claude Fable 5ANTHROPIC79.476.5162784.6%A$28.08GLOBAL
09Grok 4.6XAI79.176.8161688.4%A$4.21GLOBAL
10Gemini 3.7 FlashGOOGLE76.476.1159585.8%A$2.11GLOBAL
11GLM-5.3Z.AI76.274.8162083.9%A$3.02SELF-HOST
12Gemini 3.8 FlashGOOGLE75.876.3158387.6%A$2.11GLOBAL
13Qwen3.8 Flash NextALIBABA74.773.1163686.1%A$0.32SELF-HOST
14GPT-5.6 TerraOPENAI71.576.7151888.0%A$6.32GLOBAL
15Claude Opus 4.8ANTHROPIC70.674.3155684.6%A$14.04IN AU
16GLM-5.3 FlashZ.AI70.471.5161284.3%A$0.33SELF-HOST
17Claude Opus 4.7ANTHROPIC69.673.6155883.1%A$14.04IN AU
18GPT-5.5OPENAI68.174.9151084.3%A$15.79GLOBAL
19Grok 4.5XAI67.372.4155281.6%A$4.21NOT IN AU
20Muse Spark 1.2META65.772.2153480.1%A$2.81NOT IN AU
21GLM-5.2Z.AI65.168.8159877.9%A$3.02SELF-HOST
22Claude Sonnet 5ANTHROPIC64.871.5153880.5%A$5.62IN AU
23Muse Spark 1.1META64.871.3154277.9%A$2.81
24DeepSeek V4 FlashDEEPSEEK64.269.1158078.7%A$0.93GLOBAL
25DeepSeek V4 ProDEEPSEEK63.868.8158078.7%A$2.78GLOBAL
26Qwen3.8 27BALIBABA63.568.1159179.8%A$1.58SELF-HOST
27GPT-5.6 LunaOPENAI63.371.4152080.9%A$0.63GLOBAL
28Gemini 3.6 FlashGOOGLE61.169.2153777.5%A$2.11GLOBAL
29Gemini 3.5 FlashGOOGLE59.870.1150078.7%A$4.74IN AU
30GPT-5.4OPENAI58.571.1146378.3%A$7.90IN AU
31Qwen3.7 MaxALIBABA54.666.0151774.5%A$5.26
32Gemini 3.1 ProGOOGLE53.768.8144773.8%A$6.32GLOBAL
33Claude Sonnet 4.6ANTHROPIC50.163.0152171.2%A$8.42IN AU
34Kimi K2.6MOONSHOT AI47.461.8150965.9%A$2.40SELF-HOST
35Hy3TENCENT42.758.8150964.4%A$0.35SELF-HOST
36MiMo V2.5 ProXIAOMI42.460.2147765.2%A$0.76SELF-HOST
37MiniMax M3MINIMAX40.658.6148565.2%A$0.74SELF-HOST
38Muse SparkMETAPARTIAL39.858.662.2%
39GPT-5.4 miniOPENAI30.056.1139759.2%A$2.37
40Gemini 3.5 Flash-LiteGOOGLE23.149.3144753.6%A$1.19GLOBAL
41Gemma 4 31BGOOGLE7.543.4136443.4%A$0.29SELF-HOST
42Mistral Medium 3.5MISTRAL AI5.546.9126550.6%A$4.21SELF-HOST
HOW THIS RANKING WORKS

What goes into the coding score.

The Coding Index covers agentic terminal work and scientific programming; the WebDev arena adds how people rate finished, running apps. Terminal-Bench is shown alongside because it is the closest test to how coding agents are actually used, but it is already inside the Coding Index, so it isn’t counted twice.

60%AA Coding Index

Artificial Analysis’s composite of its coding evaluations, including Terminal-Bench and SciCode.

40%LMArena WebDev

Blind votes on working web apps each model builds from the same prompt.

Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →

QUESTIONS

Choosing a model for coding.

As of 23 Sep 2026, Claude Fable 5.1 ranks first, followed by GPT-6 Astra and Claude Opus 5. The ranking combines AA Coding Index and LMArena WebDev and is refreshed weekly.

Kimi K3 from Moonshot AI is the highest-ranked open-weights model (5th of 42). Open weights can be run on your own infrastructure, including in an Australian region.

Of the top ten, Gemini 3.7 Flash is cheapest at A$2.11 per million tokens (blended), against A$28.08 for Claude Fable 5.1. Prices are list prices converted at A$1 = US$0.7123, excluding GST.

Of the top ten, Claude Opus 5 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.

The Coding Index covers agentic terminal work and scientific programming; the WebDev arena adds how people rate finished, running apps. Terminal-Bench is shown alongside because it is the closest test to how coding agents are actually used, but it is already inside the Coding Index, so it isn’t counted twice. Each input is rescaled 0–100 across the models we track, then weighted: AA Coding Index 60% and LMArena WebDev 40%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?