AI model benchmarks and LLM leaderboard.
51 models from 14 labs, ranked on Artificial Analysis’s independent tests and LMArena’s blind votes, with list prices converted to Australian dollars and where each model can run in Australia. Snapshot: 23 Sep 2026. Use the results to shortlist models, then test them on your own work.
| MODEL | IN AUSTRALIA | |||||
|---|---|---|---|---|---|---|
| 01Claude Opus 5.5ANTHROPICPARTIAL | 100.0 | 57.6 | A$11.23 | IN AU | ||
| 02Claude Fable 5ANTHROPIC | 90.8 | 49.6 | 1506 | A$28.08 | #8 | GLOBAL |
| 03Claude Fable 5.1ANTHROPIC | 90.6 | 53.4 | 1499 | A$28.08 | #1 | GLOBAL |
| 04Claude Opus 5ANTHROPIC | 84.1 | 50.8 | 1493 | A$14.04 | #3 | IN AU |
| 05Muse Spark 1.3META | 81.1 | 48.1 | 1493 | A$2.81 | #6 | NOT IN AU |
| 06Claude Opus 4.7ANTHROPIC | 78.0 | 40.7 | 1502 | A$14.04 | #17 | IN AU |
| 07GPT-6 AstraOPENAI | 78.0 | 52.7 | 1480 | A$28.08 | #2 | GLOBAL |
| 08GPT-6 SolOPENAIPARTIAL | 76.7 | 47.5 | A$5.62 | GLOBAL | ||
| 09Muse Spark 1.2META | 75.4 | 39.6 | 1500 | A$2.81 | #20 | NOT IN AU |
| 10Grok 4.7XAIPARTIAL | 74.2 | 46.4 | A$4.21 | NOT IN AU | ||
| 11MiMo V2.6 ProXIAOMIPARTIAL | 74.0 | 46.3 | A$0.76 | NOT IN AU | ||
| 12GPT-5.6 SolOPENAI | 73.8 | 47.0 | 1484 | A$11.23 | #7 | GLOBAL |
| 13Gemini 3.8 FlashGOOGLE | 72.8 | 40.9 | 1493 | A$2.11 | #12 | GLOBAL |
| 14GLM-5.3Z.AI | 71.0 | 44.8 | 1483 | A$3.02 | #11 | SELF-HOST |
| 15Kimi K3MOONSHOT AI | 70.7 | 43.6 | 1485 | A$8.42 | #5 | GLOBAL |
| 16Qwen3.8 MaxALIBABA | 70.1 | 45.4 | 1481 | A$4.21 | #4 | NOT IN AU |
| 17Claude Opus 4.6ANTHROPIC | 69.7 | 31.9 | 1505 | A$14.04 | IN AU | |
| 18Gemini 3.7 FlashGOOGLE | 68.7 | 39.1 | 1490 | A$2.11 | #10 | GLOBAL |
| 19Step 5STEPFUNPARTIAL | 68.0 | 43.7 | A$2.00 | |||
| 20Claude Opus 4.8ANTHROPIC | 66.5 | 41.8 | 1481 | A$14.04 | #15 | IN AU |
| 21Muse Spark 1.1META | 64.2 | 33.7 | 1493 | A$2.81 | #23 | |
| 22GPT-5.5OPENAI | 62.8 | 38.4 | 1482 | A$15.79 | #18 | GLOBAL |
| 23GLM-5.3 FlashZ.AI | 62.7 | 41.8 | 1475 | A$0.33 | #16 | SELF-HOST |
| 24GPT-5.4OPENAI | 60.1 | 39.0 | 1476 | A$7.90 | #30 | IN AU |
| 25Qwen3.8 Flash NextALIBABAPARTIAL | 59.0 | 39.8 | A$0.32 | #13 | SELF-HOST | |
| 26Muse SparkMETA | 58.6 | 31.3 | 1488 | #38 | ||
| 27DeepSeek V4.1 FlashDEEPSEEKPARTIAL | 58.3 | 39.5 | A$0.74 | SELF-HOST | ||
| 28GPT-5.6 TerraOPENAI | 57.1 | 42.1 | 1466 | A$6.32 | #14 | GLOBAL |
| 29Gemini 3.6 FlashGOOGLE | 56.6 | 34.0 | 1480 | A$2.11 | #28 | GLOBAL |
| 30Gemini 3.1 ProGOOGLE | 56.0 | 29.7 | 1487 | A$6.32 | #32 | GLOBAL |
| 31Grok 4.5XAI | 54.7 | 38.8 | 1468 | A$4.21 | #19 | NOT IN AU |
| 32Gemini 3.5 FlashGOOGLE | 53.6 | 32.6 | 1478 | A$4.74 | #29 | IN AU |
| 33Gemini 3 ProGOOGLE | 53.2 | 28.0 | 1486 | A$6.32 | ||
| 34GPT-6 LunaOPENAIPARTIAL | 53.2 | 37.3 | A$0.28 | GLOBAL | ||
| 35Grok 4.6XAI | 53.1 | 44.3 | 1456 | A$4.21 | #9 | GLOBAL |
| 36GLM-5.2Z.AI | 51.3 | 33.7 | 1472 | A$3.02 | #21 | SELF-HOST |
| 37Claude Sonnet 5ANTHROPIC | 49.6 | 38.2 | 1461 | A$5.62 | #22 | IN AU |
| 38DeepSeek V4 ProDEEPSEEK | 48.5 | 36.0 | 1463 | A$2.78 | #25 | GLOBAL |
| 39Claude Sonnet 4.6ANTHROPIC | 47.5 | 30.1 | 1473 | A$8.42 | #33 | IN AU |
| 40Qwen3.7 MaxALIBABA | 47.2 | 29.5 | 1473 | A$5.26 | #31 | |
| 41DeepSeek V4 FlashDEEPSEEKPARTIAL | 46.3 | 34.3 | A$0.93 | #24 | GLOBAL | |
| 42GPT-5.6 LunaOPENAI | 43.1 | 37.3 | 1453 | A$0.63 | #27 | GLOBAL |
| 43MiMo V2.5 ProXIAOMI | 39.5 | 26.0 | 1467 | A$0.76 | #36 | SELF-HOST |
| 44Kimi K2.6MOONSHOT AI | 36.2 | 27.0 | 1460 | A$2.40 | #34 | SELF-HOST |
| 45Hy3TENCENT | 31.4 | 25.3 | 1456 | A$0.35 | #35 | SELF-HOST |
| 46Qwen3.8 27BALIBABA | 29.3 | 33.7 | 1437 | A$1.58 | #26 | SELF-HOST |
| 47Gemini 3.5 Flash-LiteGOOGLE | 27.9 | 22.2 | 1456 | A$1.19 | #40 | GLOBAL |
| 48MiniMax M3MINIMAX | 26.7 | 29.2 | 1441 | A$0.74 | #37 | SELF-HOST |
| 49GPT-5.4 miniOPENAI | 25.1 | 24.1 | 1448 | A$2.37 | #39 | |
| 50Gemma 4 31BGOOGLE | 21.1 | 19.0 | 1451 | A$0.29 | #41 | SELF-HOST |
| 51Mistral Medium 3.5MISTRAL AI | 0.0 | 14.2 | 1426 | A$4.21 | #42 | SELF-HOST |
As of 23 Sep 2026, Claude Opus 5.5 leads, ahead of Claude Fable 5 and Claude Fable 5.1, out of 51 ranked models. The highest-ranked open-weights model is GLM-5.3 (14th). The cheapest model in the top ten is Muse Spark 1.3, at A$2.81 per million tokens.
The leader for each kind of work.
One overall number hides a lot. These rankings weigh the tests that matter for each job.
Where you can run them in Australia.
9 of the 51 models we track can run with requests processed in Australia on at least one major cloud platform — mostly Claude, through AWS Bedrock in Sydney and Melbourne. The rest are reached through global routing, or self-hosted if the weights are open.
AUSTRALIAN AVAILABILITYWhat goes into the overall score.
Half the score is Artificial Analysis’s Intelligence Index (independent tests), half is LMArena’s text rating (which answers people prefer in blind votes). Scores from tests and votes don’t always agree, so using both gives a steadier ranking than either alone. New models often appear in one source a week or two before the other; they are ranked on what’s available and marked partial.
Artificial Analysis’s composite of its independent evaluations: reasoning, knowledge, maths, coding, long context and agent tasks.
People compare two anonymous answers to the same prompt and vote for the better one. Style-controlled, so longer or more heavily formatted answers don’t win on looks alone.
Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →
What people ask about AI model rankings.
What is the best AI model right now?
As of 23 Sep 2026, Claude Opus 5.5 ranks first, followed by Claude Fable 5 and Claude Fable 5.1. The ranking combines AA Intelligence Index and LMArena Text and is refreshed weekly.
What is the best open-weights model?
GLM-5.3 from Z.ai is the highest-ranked open-weights model (14th of 51). Open weights can be run on your own infrastructure, including in an Australian region.
Which top model is cheapest?
Of the top ten, Muse Spark 1.3 is cheapest at A$2.81 per million tokens (blended), against A$28.08 for Claude Fable 5. Prices are list prices converted at A$1 = US$0.7123, excluding GST.
Which of these models can keep data in Australia?
Of the top ten, Claude Opus 5.5, Claude Opus 5 and Claude Opus 4.7 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.
How is this ranking calculated?
Half the score is Artificial Analysis’s Intelligence Index (independent tests), half is LMArena’s text rating (which answers people prefer in blind votes). Scores from tests and votes don’t always agree, so using both gives a steadier ranking than either alone. New models often appear in one source a week or two before the other; they are ranked on what’s available and marked partial. Each input is rescaled 0–100 across the models we track, then weighted: AA Intelligence Index 50% and LMArena Text 50%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →