Best AI models for agents and automation.

Models ranked on multi-step work: following policy while using tools, getting jobs done in a terminal, and how people rate real agent sessions. As of 23 Sep 2026, Claude Fable 5.1 leads for agents, ahead of GPT-6 Astra and Claude Opus 5, out of 41 ranked models. The highest-ranked open-weights model is GLM-5.3 (8th). The cheapest model in the top ten is Gemini 3.8 Flash, at A$2.11 per million tokens.

41 MODELS RANKEDA$1 = US$0.7123UPDATED 23 SEP 2026
AGENTS RANKING41 MODELS · 23 SEP 2026
41 MODELS
Best AI models for agents and automation, ranked, as of 23 Sep 2026
MODELIN AUSTRALIA
01Claude Fable 5.1ANTHROPIC96.647.2%91.4%137A$28.08GLOBAL
02GPT-6 AstraOPENAI86.841.4%88.4%115A$28.08GLOBAL
03Claude Opus 5ANTHROPIC86.642.1%89.1%103A$14.04IN AU
04Grok 4.6XAI86.050.7%88.4%20A$4.21GLOBAL
05Muse Spark 1.3META85.550.5%84.3%42A$2.81NOT IN AU
06GPT-5.6 SolOPENAI84.844.3%88.0%71A$11.23GLOBAL
07Qwen3.8 MaxALIBABA84.847.8%88.8%33A$4.21NOT IN AU
08GLM-5.3Z.AI83.950.3%83.9%31A$3.02SELF-HOST
09Kimi K3MOONSHOT AI83.746.0%85.0%62A$8.42GLOBAL
10Gemini 3.8 FlashGOOGLE82.644.9%87.6%47A$2.11GLOBAL
11GLM-5.3 FlashZ.AI79.247.2%84.3%12A$0.33SELF-HOST
12Claude Fable 5ANTHROPIC78.338.1%84.6%88A$28.08GLOBAL
13Qwen3.8 Flash NextALIBABA77.345.4%86.1%-1A$0.32SELF-HOST
14Qwen3.8 27BALIBABA75.348.0%79.8%-6A$1.58SELF-HOST
15GPT-5.5OPENAI75.139.0%84.3%50A$15.79GLOBAL
16GPT-5.6 TerraOPENAI74.940.2%88.0%14A$6.32GLOBAL
17Grok 4.5XAI74.342.1%81.6%29A$4.21NOT IN AU
18Claude Opus 4.8ANTHROPIC73.934.2%84.6%82A$14.04IN AU
19Claude Sonnet 5ANTHROPIC72.137.3%80.5%60A$5.62IN AU
20DeepSeek V4 ProDEEPSEEK71.339.6%78.7%41A$2.78GLOBAL
21Claude Opus 4.7ANTHROPICPARTIAL70.034.6%83.1%A$14.04IN AU
22DeepSeek V4 FlashDEEPSEEK68.739.4%78.7%18A$0.93GLOBAL
23GPT-5.4OPENAI68.139.6%78.3%13A$7.90IN AU
24GLM-5.2Z.AI66.134.6%77.9%44A$3.02SELF-HOST
25Gemini 3.7 FlashGOOGLE64.232.8%85.8%-6A$2.11GLOBAL
26Gemini 3.5 FlashGOOGLEPARTIAL62.832.2%78.7%A$4.74IN AU
27Muse Spark 1.2META61.434.8%80.1%-18A$2.81NOT IN AU
28GPT-5.6 LunaOPENAI59.631.1%80.9%-4A$0.63GLOBAL
29Muse Spark 1.1META55.831.8%77.9%-30A$2.81
30Claude Sonnet 4.6ANTHROPIC55.734.4%71.2%-15A$8.42IN AU
31Gemini 3.6 FlashGOOGLE50.629.9%77.5%-59A$2.11GLOBAL
32Gemini 3.1 ProGOOGLE40.121.4%73.8%-58A$6.32GLOBAL
33Kimi K2.6MOONSHOT AIPARTIAL38.923.3%65.9%A$2.40SELF-HOST
34Hy3TENCENT36.322.9%64.4%-52A$0.35SELF-HOST
35GPT-5.4 miniOPENAIPARTIAL36.125.6%59.2%A$2.37
36Qwen3.7 MaxALIBABA33.411.8%74.5%-36A$5.26
37MiniMax M3MINIMAX28.815.3%65.2%-58A$0.74SELF-HOST
38MiMo V2.5 ProXIAOMI23.69.9%65.2%-57A$0.76SELF-HOST
39Mistral Medium 3.5MISTRAL AI14.215.1%50.6%-109A$4.21SELF-HOST
40Gemini 3.5 Flash-LiteGOOGLE13.817.5%53.6%-153A$1.19GLOBAL
41Gemma 4 31BGOOGLEPARTIAL6.914.8%43.4%A$0.29SELF-HOST
HOW THIS RANKING WORKS

What goes into the agents score.

τ²-Bench is the closest public test to a business agent — a customer conversation where the model has tools and rules and must resolve the request correctly. Terminal-Bench covers operational tasks; LMArena’s agent arena adds human judgement of real sessions. A model needs at least two of the three to be ranked.

40%τ²-Bench (banking)

Customer-service agent conversations in a banking setting, where the model must use tools and follow policy to resolve the request.

30%Terminal-Bench 2.1

Practical tasks completed by an agent working in a real command-line environment.

30%LMArena Agent

Multi-step agent sessions rated by the people running them. LMArena publishes a relative score, so the rank is the useful part.

Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →

QUESTIONS

Choosing a model for agents.

As of 23 Sep 2026, Claude Fable 5.1 ranks first, followed by GPT-6 Astra and Claude Opus 5. The ranking combines τ²-Bench (banking), Terminal-Bench 2.1 and LMArena Agent and is refreshed weekly.

GLM-5.3 from Z.ai is the highest-ranked open-weights model (8th of 41). Open weights can be run on your own infrastructure, including in an Australian region.

Of the top ten, Gemini 3.8 Flash is cheapest at A$2.11 per million tokens (blended), against A$28.08 for Claude Fable 5.1. Prices are list prices converted at A$1 = US$0.7123, excluding GST.

Of the top ten, Claude Opus 5 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.

τ²-Bench is the closest public test to a business agent — a customer conversation where the model has tools and rules and must resolve the request correctly. Terminal-Bench covers operational tasks; LMArena’s agent arena adds human judgement of real sessions. A model needs at least two of the three to be ranked. Each input is rescaled 0–100 across the models we track, then weighted: τ²-Bench (banking) 40%, Terminal-Bench 2.1 30% and LMArena Agent 30%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?