Best AI models for reasoning and analysis.

The models that handle hard, multi-step problems best, on independent tests rather than chat preferences. As of 23 Sep 2026, Claude Opus 5.5 leads for reasoning, ahead of Claude Fable 5.1 and GPT-6 Astra, out of 51 ranked models. The highest-ranked open-weights model is Kimi K3 (11th). The cheapest model in the top ten is MiMo V2.6 Pro, at A$0.76 per million tokens.

51 MODELS RANKEDA$1 = US$0.7123UPDATED 23 SEP 2026
REASONING RANKING51 MODELS · 23 SEP 2026
51 MODELS
Best AI models for reasoning and analysis, ranked, as of 23 Sep 2026
MODELIN AUSTRALIA
01Claude Opus 5.5ANTHROPICPARTIAL100.057.661.4%A$11.23IN AU
02Claude Fable 5.1ANTHROPIC91.153.459.1%93.7%A$28.08GLOBAL
03GPT-6 AstraOPENAI90.852.754.7%96.1%A$28.08GLOBAL
04Claude Opus 5ANTHROPIC85.350.854.9%93.2%A$14.04IN AU
05Claude Fable 5ANTHROPIC83.649.655.5%92.6%A$28.08GLOBAL
06Muse Spark 1.3META79.348.148.7%93.5%A$2.81NOT IN AU
07GPT-5.6 SolOPENAI79.247.049.5%94.1%A$11.23GLOBAL
08GPT-6 SolOPENAIPARTIAL75.047.547.9%A$5.62GLOBAL
09MiMo V2.6 ProXIAOMIPARTIAL74.246.349.4%A$0.76NOT IN AU
10Grok 4.6XAI73.644.342.9%94.9%A$4.21GLOBAL
11Kimi K3MOONSHOT AI73.243.646.9%93.5%A$8.42GLOBAL
12Gemini 3.8 FlashGOOGLE72.740.947.8%95.3%A$2.11GLOBAL
13Qwen3.8 MaxALIBABA72.545.443.1%92.8%A$4.21NOT IN AU
14Claude Opus 4.8ANTHROPIC70.341.848.7%92.0%A$14.04IN AU
15GLM-5.3Z.AI70.144.842.3%91.7%A$3.02SELF-HOST
16Grok 4.7XAIPARTIAL70.046.443.1%A$4.21NOT IN AU
17Gemini 3.7 FlashGOOGLE69.739.147.9%94.5%A$2.11GLOBAL
18GPT-5.6 TerraOPENAI68.242.142.9%92.5%A$6.32GLOBAL
19Step 5STEPFUNPARTIAL68.243.746.5%A$2.00
20GPT-5.5OPENAI66.638.445.8%93.5%A$15.79GLOBAL
21Claude Opus 4.7ANTHROPIC65.040.742.3%91.4%A$14.04IN AU
22Grok 4.5XAI65.038.842.7%93.1%A$4.21NOT IN AU
23GLM-5.3 FlashZ.AI64.841.839.9%91.2%A$0.33SELF-HOST
24GPT-5.4OPENAI64.539.043.7%92.0%A$7.90IN AU
25Muse Spark 1.2META64.239.645.5%90.4%A$2.81NOT IN AU
26Qwen3.8 Flash NextALIBABA62.739.838.0%92.3%A$0.32SELF-HOST
27Claude Sonnet 5ANTHROPIC61.238.241.3%91.1%A$5.62IN AU
28DeepSeek V4 ProDEEPSEEK60.536.041.0%92.8%A$2.78GLOBAL
29GPT-5.6 LunaOPENAI59.237.339.5%91.1%A$0.63GLOBAL
30Gemini 3.6 FlashGOOGLE58.134.040.8%92.8%A$2.11GLOBAL
31Gemini 3.1 ProGOOGLE57.929.747.0%94.1%A$6.32GLOBAL
32Muse Spark 1.1META57.133.746.2%89.8%A$2.81
33Gemini 3.5 FlashGOOGLE56.832.642.7%92.2%A$4.74IN AU
34DeepSeek V4.1 FlashDEEPSEEKPARTIAL56.739.539.2%A$0.74SELF-HOST
35DeepSeek V4 FlashDEEPSEEK55.034.338.6%90.8%A$0.93GLOBAL
36GLM-5.2Z.AI54.133.741.1%89.5%A$3.02SELF-HOST
37GPT-6 LunaOPENAIPARTIAL52.837.338.5%A$0.28GLOBAL
38Qwen3.7 MaxALIBABA52.229.540.5%92.3%A$5.26
39MiniMax M3MINIMAX51.829.239.0%92.9%A$0.74SELF-HOST
40Claude Opus 4.6ANTHROPIC51.531.939.9%89.6%A$14.04IN AU
41Qwen3.8 27BALIBABA51.433.733.9%90.5%A$1.58SELF-HOST
42Muse SparkMETA49.831.340.7%88.4%
43Gemini 3 ProGOOGLE48.328.039.7%90.8%A$6.32
44Kimi K2.6MOONSHOT AI46.327.037.5%91.1%A$2.40SELF-HOST
45Claude Sonnet 4.6ANTHROPIC43.630.133.6%87.5%A$8.42IN AU
46Hy3TENCENT40.625.333.5%89.7%A$0.35SELF-HOST
47MiMo V2.5 ProXIAOMI38.926.035.7%86.6%A$0.76SELF-HOST
48GPT-5.4 miniOPENAI33.824.128.1%87.5%A$2.37
49Gemma 4 31BGOOGLE23.519.023.6%85.7%A$0.29SELF-HOST
50Gemini 3.5 Flash-LiteGOOGLE22.422.218.8%83.8%A$1.19GLOBAL
51Mistral Medium 3.5MISTRAL AI0.014.213.8%74.8%A$4.21SELF-HOST
HOW THIS RANKING WORKS

What goes into the reasoning score.

The Intelligence Index already includes Humanity’s Last Exam and GPQA Diamond; this lens weights them up so the ranking leans towards expert-level questions where there is a single right answer.

50%AA Intelligence Index

Artificial Analysis’s composite of its independent evaluations: reasoning, knowledge, maths, coding, long context and agent tasks.

25%Humanity’s Last Exam

Very hard expert-written questions across dozens of fields; frontier models still get most of them wrong.

25%GPQA Diamond

Graduate-level biology, physics and chemistry questions written to be hard to look up.

Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →

QUESTIONS

Choosing a model for reasoning.

As of 23 Sep 2026, Claude Opus 5.5 ranks first, followed by Claude Fable 5.1 and GPT-6 Astra. The ranking combines AA Intelligence Index, Humanity’s Last Exam and GPQA Diamond and is refreshed weekly.

Kimi K3 from Moonshot AI is the highest-ranked open-weights model (11th of 51). Open weights can be run on your own infrastructure, including in an Australian region.

Of the top ten, MiMo V2.6 Pro is cheapest at A$0.76 per million tokens (blended), against A$28.08 for Claude Fable 5.1. Prices are list prices converted at A$1 = US$0.7123, excluding GST.

Of the top ten, Claude Opus 5.5 and Claude Opus 5 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.

The Intelligence Index already includes Humanity’s Last Exam and GPQA Diamond; this lens weights them up so the ranking leans towards expert-level questions where there is a single right answer. Each input is rescaled 0–100 across the models we track, then weighted: AA Intelligence Index 50%, Humanity’s Last Exam 25% and GPQA Diamond 25%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?