AI model benchmarks and LLM leaderboard.

51 models from 14 labs, ranked on Artificial Analysis’s independent tests and LMArena’s blind votes, with list prices converted to Australian dollars and where each model can run in Australia. Snapshot: 23 Sep 2026. Use the results to shortlist models, then test them on your own work.

51 MODELS RANKEDA$1 = US$0.7123UPDATED 23 SEP 2026
OVERALL RANKING51 MODELS · 23 SEP 2026
51 MODELS
LLM leaderboard, ranked, as of 23 Sep 2026
MODELIN AUSTRALIA
01Claude Opus 5.5ANTHROPICPARTIAL100.057.6A$11.23IN AU
02Claude Fable 5ANTHROPIC90.849.61506A$28.08#8GLOBAL
03Claude Fable 5.1ANTHROPIC90.653.41499A$28.08#1GLOBAL
04Claude Opus 5ANTHROPIC84.150.81493A$14.04#3IN AU
05Muse Spark 1.3META81.148.11493A$2.81#6NOT IN AU
06Claude Opus 4.7ANTHROPIC78.040.71502A$14.04#17IN AU
07GPT-6 AstraOPENAI78.052.71480A$28.08#2GLOBAL
08GPT-6 SolOPENAIPARTIAL76.747.5A$5.62GLOBAL
09Muse Spark 1.2META75.439.61500A$2.81#20NOT IN AU
10Grok 4.7XAIPARTIAL74.246.4A$4.21NOT IN AU
11MiMo V2.6 ProXIAOMIPARTIAL74.046.3A$0.76NOT IN AU
12GPT-5.6 SolOPENAI73.847.01484A$11.23#7GLOBAL
13Gemini 3.8 FlashGOOGLE72.840.91493A$2.11#12GLOBAL
14GLM-5.3Z.AI71.044.81483A$3.02#11SELF-HOST
15Kimi K3MOONSHOT AI70.743.61485A$8.42#5GLOBAL
16Qwen3.8 MaxALIBABA70.145.41481A$4.21#4NOT IN AU
17Claude Opus 4.6ANTHROPIC69.731.91505A$14.04IN AU
18Gemini 3.7 FlashGOOGLE68.739.11490A$2.11#10GLOBAL
19Step 5STEPFUNPARTIAL68.043.7A$2.00
20Claude Opus 4.8ANTHROPIC66.541.81481A$14.04#15IN AU
21Muse Spark 1.1META64.233.71493A$2.81#23
22GPT-5.5OPENAI62.838.41482A$15.79#18GLOBAL
23GLM-5.3 FlashZ.AI62.741.81475A$0.33#16SELF-HOST
24GPT-5.4OPENAI60.139.01476A$7.90#30IN AU
25Qwen3.8 Flash NextALIBABAPARTIAL59.039.8A$0.32#13SELF-HOST
26Muse SparkMETA58.631.31488#38
27DeepSeek V4.1 FlashDEEPSEEKPARTIAL58.339.5A$0.74SELF-HOST
28GPT-5.6 TerraOPENAI57.142.11466A$6.32#14GLOBAL
29Gemini 3.6 FlashGOOGLE56.634.01480A$2.11#28GLOBAL
30Gemini 3.1 ProGOOGLE56.029.71487A$6.32#32GLOBAL
31Grok 4.5XAI54.738.81468A$4.21#19NOT IN AU
32Gemini 3.5 FlashGOOGLE53.632.61478A$4.74#29IN AU
33Gemini 3 ProGOOGLE53.228.01486A$6.32
34GPT-6 LunaOPENAIPARTIAL53.237.3A$0.28GLOBAL
35Grok 4.6XAI53.144.31456A$4.21#9GLOBAL
36GLM-5.2Z.AI51.333.71472A$3.02#21SELF-HOST
37Claude Sonnet 5ANTHROPIC49.638.21461A$5.62#22IN AU
38DeepSeek V4 ProDEEPSEEK48.536.01463A$2.78#25GLOBAL
39Claude Sonnet 4.6ANTHROPIC47.530.11473A$8.42#33IN AU
40Qwen3.7 MaxALIBABA47.229.51473A$5.26#31
41DeepSeek V4 FlashDEEPSEEKPARTIAL46.334.3A$0.93#24GLOBAL
42GPT-5.6 LunaOPENAI43.137.31453A$0.63#27GLOBAL
43MiMo V2.5 ProXIAOMI39.526.01467A$0.76#36SELF-HOST
44Kimi K2.6MOONSHOT AI36.227.01460A$2.40#34SELF-HOST
45Hy3TENCENT31.425.31456A$0.35#35SELF-HOST
46Qwen3.8 27BALIBABA29.333.71437A$1.58#26SELF-HOST
47Gemini 3.5 Flash-LiteGOOGLE27.922.21456A$1.19#40GLOBAL
48MiniMax M3MINIMAX26.729.21441A$0.74#37SELF-HOST
49GPT-5.4 miniOPENAI25.124.11448A$2.37#39
50Gemma 4 31BGOOGLE21.119.01451A$0.29#41SELF-HOST
51Mistral Medium 3.5MISTRAL AI0.014.21426A$4.21#42SELF-HOST

As of 23 Sep 2026, Claude Opus 5.5 leads, ahead of Claude Fable 5 and Claude Fable 5.1, out of 51 ranked models. The highest-ranked open-weights model is GLM-5.3 (14th). The cheapest model in the top ten is Muse Spark 1.3, at A$2.81 per million tokens.

DATA SOVEREIGNTY

Where you can run them in Australia.

9 of the 51 models we track can run with requests processed in Australia on at least one major cloud platform — mostly Claude, through AWS Bedrock in Sydney and Melbourne. The rest are reached through global routing, or self-hosted if the weights are open.

AUSTRALIAN AVAILABILITY
HOW THIS RANKING WORKS

What goes into the overall score.

Half the score is Artificial Analysis’s Intelligence Index (independent tests), half is LMArena’s text rating (which answers people prefer in blind votes). Scores from tests and votes don’t always agree, so using both gives a steadier ranking than either alone. New models often appear in one source a week or two before the other; they are ranked on what’s available and marked partial.

50%AA Intelligence Index

Artificial Analysis’s composite of its independent evaluations: reasoning, knowledge, maths, coding, long context and agent tasks.

50%LMArena Text

People compare two anonymous answers to the same prompt and vote for the better one. Style-controlled, so longer or more heavily formatted answers don’t win on looks alone.

Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →

QUESTIONS

What people ask about AI model rankings.

As of 23 Sep 2026, Claude Opus 5.5 ranks first, followed by Claude Fable 5 and Claude Fable 5.1. The ranking combines AA Intelligence Index and LMArena Text and is refreshed weekly.

GLM-5.3 from Z.ai is the highest-ranked open-weights model (14th of 51). Open weights can be run on your own infrastructure, including in an Australian region.

Of the top ten, Muse Spark 1.3 is cheapest at A$2.81 per million tokens (blended), against A$28.08 for Claude Fable 5. Prices are list prices converted at A$1 = US$0.7123, excluding GST.

Of the top ten, Claude Opus 5.5, Claude Opus 5 and Claude Opus 4.7 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.

Half the score is Artificial Analysis’s Intelligence Index (independent tests), half is LMArena’s text rating (which answers people prefer in blind votes). Scores from tests and votes don’t always agree, so using both gives a steadier ranking than either alone. New models often appear in one source a week or two before the other; they are ranked on what’s available and marked partial. Each input is rescaled 0–100 across the models we track, then weighted: AA Intelligence Index 50% and LMArena Text 50%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?