Best AI models for reasoning and analysis.
The models that handle hard, multi-step problems best, on independent tests rather than chat preferences. As of 23 Sep 2026, Claude Opus 5.5 leads for reasoning, ahead of Claude Fable 5.1 and GPT-6 Astra, out of 51 ranked models. The highest-ranked open-weights model is Kimi K3 (11th). The cheapest model in the top ten is MiMo V2.6 Pro, at A$0.76 per million tokens.
| MODEL | IN AUSTRALIA | |||||
|---|---|---|---|---|---|---|
| 01Claude Opus 5.5ANTHROPICPARTIAL | 100.0 | 57.6 | 61.4% | A$11.23 | IN AU | |
| 02Claude Fable 5.1ANTHROPIC | 91.1 | 53.4 | 59.1% | 93.7% | A$28.08 | GLOBAL |
| 03GPT-6 AstraOPENAI | 90.8 | 52.7 | 54.7% | 96.1% | A$28.08 | GLOBAL |
| 04Claude Opus 5ANTHROPIC | 85.3 | 50.8 | 54.9% | 93.2% | A$14.04 | IN AU |
| 05Claude Fable 5ANTHROPIC | 83.6 | 49.6 | 55.5% | 92.6% | A$28.08 | GLOBAL |
| 06Muse Spark 1.3META | 79.3 | 48.1 | 48.7% | 93.5% | A$2.81 | NOT IN AU |
| 07GPT-5.6 SolOPENAI | 79.2 | 47.0 | 49.5% | 94.1% | A$11.23 | GLOBAL |
| 08GPT-6 SolOPENAIPARTIAL | 75.0 | 47.5 | 47.9% | A$5.62 | GLOBAL | |
| 09MiMo V2.6 ProXIAOMIPARTIAL | 74.2 | 46.3 | 49.4% | A$0.76 | NOT IN AU | |
| 10Grok 4.6XAI | 73.6 | 44.3 | 42.9% | 94.9% | A$4.21 | GLOBAL |
| 11Kimi K3MOONSHOT AI | 73.2 | 43.6 | 46.9% | 93.5% | A$8.42 | GLOBAL |
| 12Gemini 3.8 FlashGOOGLE | 72.7 | 40.9 | 47.8% | 95.3% | A$2.11 | GLOBAL |
| 13Qwen3.8 MaxALIBABA | 72.5 | 45.4 | 43.1% | 92.8% | A$4.21 | NOT IN AU |
| 14Claude Opus 4.8ANTHROPIC | 70.3 | 41.8 | 48.7% | 92.0% | A$14.04 | IN AU |
| 15GLM-5.3Z.AI | 70.1 | 44.8 | 42.3% | 91.7% | A$3.02 | SELF-HOST |
| 16Grok 4.7XAIPARTIAL | 70.0 | 46.4 | 43.1% | A$4.21 | NOT IN AU | |
| 17Gemini 3.7 FlashGOOGLE | 69.7 | 39.1 | 47.9% | 94.5% | A$2.11 | GLOBAL |
| 18GPT-5.6 TerraOPENAI | 68.2 | 42.1 | 42.9% | 92.5% | A$6.32 | GLOBAL |
| 19Step 5STEPFUNPARTIAL | 68.2 | 43.7 | 46.5% | A$2.00 | ||
| 20GPT-5.5OPENAI | 66.6 | 38.4 | 45.8% | 93.5% | A$15.79 | GLOBAL |
| 21Claude Opus 4.7ANTHROPIC | 65.0 | 40.7 | 42.3% | 91.4% | A$14.04 | IN AU |
| 22Grok 4.5XAI | 65.0 | 38.8 | 42.7% | 93.1% | A$4.21 | NOT IN AU |
| 23GLM-5.3 FlashZ.AI | 64.8 | 41.8 | 39.9% | 91.2% | A$0.33 | SELF-HOST |
| 24GPT-5.4OPENAI | 64.5 | 39.0 | 43.7% | 92.0% | A$7.90 | IN AU |
| 25Muse Spark 1.2META | 64.2 | 39.6 | 45.5% | 90.4% | A$2.81 | NOT IN AU |
| 26Qwen3.8 Flash NextALIBABA | 62.7 | 39.8 | 38.0% | 92.3% | A$0.32 | SELF-HOST |
| 27Claude Sonnet 5ANTHROPIC | 61.2 | 38.2 | 41.3% | 91.1% | A$5.62 | IN AU |
| 28DeepSeek V4 ProDEEPSEEK | 60.5 | 36.0 | 41.0% | 92.8% | A$2.78 | GLOBAL |
| 29GPT-5.6 LunaOPENAI | 59.2 | 37.3 | 39.5% | 91.1% | A$0.63 | GLOBAL |
| 30Gemini 3.6 FlashGOOGLE | 58.1 | 34.0 | 40.8% | 92.8% | A$2.11 | GLOBAL |
| 31Gemini 3.1 ProGOOGLE | 57.9 | 29.7 | 47.0% | 94.1% | A$6.32 | GLOBAL |
| 32Muse Spark 1.1META | 57.1 | 33.7 | 46.2% | 89.8% | A$2.81 | |
| 33Gemini 3.5 FlashGOOGLE | 56.8 | 32.6 | 42.7% | 92.2% | A$4.74 | IN AU |
| 34DeepSeek V4.1 FlashDEEPSEEKPARTIAL | 56.7 | 39.5 | 39.2% | A$0.74 | SELF-HOST | |
| 35DeepSeek V4 FlashDEEPSEEK | 55.0 | 34.3 | 38.6% | 90.8% | A$0.93 | GLOBAL |
| 36GLM-5.2Z.AI | 54.1 | 33.7 | 41.1% | 89.5% | A$3.02 | SELF-HOST |
| 37GPT-6 LunaOPENAIPARTIAL | 52.8 | 37.3 | 38.5% | A$0.28 | GLOBAL | |
| 38Qwen3.7 MaxALIBABA | 52.2 | 29.5 | 40.5% | 92.3% | A$5.26 | |
| 39MiniMax M3MINIMAX | 51.8 | 29.2 | 39.0% | 92.9% | A$0.74 | SELF-HOST |
| 40Claude Opus 4.6ANTHROPIC | 51.5 | 31.9 | 39.9% | 89.6% | A$14.04 | IN AU |
| 41Qwen3.8 27BALIBABA | 51.4 | 33.7 | 33.9% | 90.5% | A$1.58 | SELF-HOST |
| 42Muse SparkMETA | 49.8 | 31.3 | 40.7% | 88.4% | ||
| 43Gemini 3 ProGOOGLE | 48.3 | 28.0 | 39.7% | 90.8% | A$6.32 | |
| 44Kimi K2.6MOONSHOT AI | 46.3 | 27.0 | 37.5% | 91.1% | A$2.40 | SELF-HOST |
| 45Claude Sonnet 4.6ANTHROPIC | 43.6 | 30.1 | 33.6% | 87.5% | A$8.42 | IN AU |
| 46Hy3TENCENT | 40.6 | 25.3 | 33.5% | 89.7% | A$0.35 | SELF-HOST |
| 47MiMo V2.5 ProXIAOMI | 38.9 | 26.0 | 35.7% | 86.6% | A$0.76 | SELF-HOST |
| 48GPT-5.4 miniOPENAI | 33.8 | 24.1 | 28.1% | 87.5% | A$2.37 | |
| 49Gemma 4 31BGOOGLE | 23.5 | 19.0 | 23.6% | 85.7% | A$0.29 | SELF-HOST |
| 50Gemini 3.5 Flash-LiteGOOGLE | 22.4 | 22.2 | 18.8% | 83.8% | A$1.19 | GLOBAL |
| 51Mistral Medium 3.5MISTRAL AI | 0.0 | 14.2 | 13.8% | 74.8% | A$4.21 | SELF-HOST |
What goes into the reasoning score.
The Intelligence Index already includes Humanity’s Last Exam and GPQA Diamond; this lens weights them up so the ranking leans towards expert-level questions where there is a single right answer.
Artificial Analysis’s composite of its independent evaluations: reasoning, knowledge, maths, coding, long context and agent tasks.
Very hard expert-written questions across dozens of fields; frontier models still get most of them wrong.
Graduate-level biology, physics and chemistry questions written to be hard to look up.
Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →
Choosing a model for reasoning.
What is the best AI model for reasoning right now?
As of 23 Sep 2026, Claude Opus 5.5 ranks first, followed by Claude Fable 5.1 and GPT-6 Astra. The ranking combines AA Intelligence Index, Humanity’s Last Exam and GPQA Diamond and is refreshed weekly.
What is the best open-weights model for reasoning?
Kimi K3 from Moonshot AI is the highest-ranked open-weights model (11th of 51). Open weights can be run on your own infrastructure, including in an Australian region.
Which top reasoning model is cheapest?
Of the top ten, MiMo V2.6 Pro is cheapest at A$0.76 per million tokens (blended), against A$28.08 for Claude Fable 5.1. Prices are list prices converted at A$1 = US$0.7123, excluding GST.
Which of these models can keep data in Australia?
Of the top ten, Claude Opus 5.5 and Claude Opus 5 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.
How is this ranking calculated?
The Intelligence Index already includes Humanity’s Last Exam and GPQA Diamond; this lens weights them up so the ranking leans towards expert-level questions where there is a single right answer. Each input is rescaled 0–100 across the models we track, then weighted: AA Intelligence Index 50%, Humanity’s Last Exam 25% and GPQA Diamond 25%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →