Best AI models for agents and automation.
Models ranked on multi-step work: following policy while using tools, getting jobs done in a terminal, and how people rate real agent sessions. As of 23 Sep 2026, Claude Fable 5.1 leads for agents, ahead of GPT-6 Astra and Claude Opus 5, out of 41 ranked models. The highest-ranked open-weights model is GLM-5.3 (8th). The cheapest model in the top ten is Gemini 3.8 Flash, at A$2.11 per million tokens.
| MODEL | IN AUSTRALIA | |||||
|---|---|---|---|---|---|---|
| 01Claude Fable 5.1ANTHROPIC | 96.6 | 47.2% | 91.4% | 137 | A$28.08 | GLOBAL |
| 02GPT-6 AstraOPENAI | 86.8 | 41.4% | 88.4% | 115 | A$28.08 | GLOBAL |
| 03Claude Opus 5ANTHROPIC | 86.6 | 42.1% | 89.1% | 103 | A$14.04 | IN AU |
| 04Grok 4.6XAI | 86.0 | 50.7% | 88.4% | 20 | A$4.21 | GLOBAL |
| 05Muse Spark 1.3META | 85.5 | 50.5% | 84.3% | 42 | A$2.81 | NOT IN AU |
| 06GPT-5.6 SolOPENAI | 84.8 | 44.3% | 88.0% | 71 | A$11.23 | GLOBAL |
| 07Qwen3.8 MaxALIBABA | 84.8 | 47.8% | 88.8% | 33 | A$4.21 | NOT IN AU |
| 08GLM-5.3Z.AI | 83.9 | 50.3% | 83.9% | 31 | A$3.02 | SELF-HOST |
| 09Kimi K3MOONSHOT AI | 83.7 | 46.0% | 85.0% | 62 | A$8.42 | GLOBAL |
| 10Gemini 3.8 FlashGOOGLE | 82.6 | 44.9% | 87.6% | 47 | A$2.11 | GLOBAL |
| 11GLM-5.3 FlashZ.AI | 79.2 | 47.2% | 84.3% | 12 | A$0.33 | SELF-HOST |
| 12Claude Fable 5ANTHROPIC | 78.3 | 38.1% | 84.6% | 88 | A$28.08 | GLOBAL |
| 13Qwen3.8 Flash NextALIBABA | 77.3 | 45.4% | 86.1% | -1 | A$0.32 | SELF-HOST |
| 14Qwen3.8 27BALIBABA | 75.3 | 48.0% | 79.8% | -6 | A$1.58 | SELF-HOST |
| 15GPT-5.5OPENAI | 75.1 | 39.0% | 84.3% | 50 | A$15.79 | GLOBAL |
| 16GPT-5.6 TerraOPENAI | 74.9 | 40.2% | 88.0% | 14 | A$6.32 | GLOBAL |
| 17Grok 4.5XAI | 74.3 | 42.1% | 81.6% | 29 | A$4.21 | NOT IN AU |
| 18Claude Opus 4.8ANTHROPIC | 73.9 | 34.2% | 84.6% | 82 | A$14.04 | IN AU |
| 19Claude Sonnet 5ANTHROPIC | 72.1 | 37.3% | 80.5% | 60 | A$5.62 | IN AU |
| 20DeepSeek V4 ProDEEPSEEK | 71.3 | 39.6% | 78.7% | 41 | A$2.78 | GLOBAL |
| 21Claude Opus 4.7ANTHROPICPARTIAL | 70.0 | 34.6% | 83.1% | A$14.04 | IN AU | |
| 22DeepSeek V4 FlashDEEPSEEK | 68.7 | 39.4% | 78.7% | 18 | A$0.93 | GLOBAL |
| 23GPT-5.4OPENAI | 68.1 | 39.6% | 78.3% | 13 | A$7.90 | IN AU |
| 24GLM-5.2Z.AI | 66.1 | 34.6% | 77.9% | 44 | A$3.02 | SELF-HOST |
| 25Gemini 3.7 FlashGOOGLE | 64.2 | 32.8% | 85.8% | -6 | A$2.11 | GLOBAL |
| 26Gemini 3.5 FlashGOOGLEPARTIAL | 62.8 | 32.2% | 78.7% | A$4.74 | IN AU | |
| 27Muse Spark 1.2META | 61.4 | 34.8% | 80.1% | -18 | A$2.81 | NOT IN AU |
| 28GPT-5.6 LunaOPENAI | 59.6 | 31.1% | 80.9% | -4 | A$0.63 | GLOBAL |
| 29Muse Spark 1.1META | 55.8 | 31.8% | 77.9% | -30 | A$2.81 | |
| 30Claude Sonnet 4.6ANTHROPIC | 55.7 | 34.4% | 71.2% | -15 | A$8.42 | IN AU |
| 31Gemini 3.6 FlashGOOGLE | 50.6 | 29.9% | 77.5% | -59 | A$2.11 | GLOBAL |
| 32Gemini 3.1 ProGOOGLE | 40.1 | 21.4% | 73.8% | -58 | A$6.32 | GLOBAL |
| 33Kimi K2.6MOONSHOT AIPARTIAL | 38.9 | 23.3% | 65.9% | A$2.40 | SELF-HOST | |
| 34Hy3TENCENT | 36.3 | 22.9% | 64.4% | -52 | A$0.35 | SELF-HOST |
| 35GPT-5.4 miniOPENAIPARTIAL | 36.1 | 25.6% | 59.2% | A$2.37 | ||
| 36Qwen3.7 MaxALIBABA | 33.4 | 11.8% | 74.5% | -36 | A$5.26 | |
| 37MiniMax M3MINIMAX | 28.8 | 15.3% | 65.2% | -58 | A$0.74 | SELF-HOST |
| 38MiMo V2.5 ProXIAOMI | 23.6 | 9.9% | 65.2% | -57 | A$0.76 | SELF-HOST |
| 39Mistral Medium 3.5MISTRAL AI | 14.2 | 15.1% | 50.6% | -109 | A$4.21 | SELF-HOST |
| 40Gemini 3.5 Flash-LiteGOOGLE | 13.8 | 17.5% | 53.6% | -153 | A$1.19 | GLOBAL |
| 41Gemma 4 31BGOOGLEPARTIAL | 6.9 | 14.8% | 43.4% | A$0.29 | SELF-HOST |
What goes into the agents score.
τ²-Bench is the closest public test to a business agent — a customer conversation where the model has tools and rules and must resolve the request correctly. Terminal-Bench covers operational tasks; LMArena’s agent arena adds human judgement of real sessions. A model needs at least two of the three to be ranked.
Customer-service agent conversations in a banking setting, where the model must use tools and follow policy to resolve the request.
Practical tasks completed by an agent working in a real command-line environment.
Multi-step agent sessions rated by the people running them. LMArena publishes a relative score, so the rank is the useful part.
Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →
Choosing a model for agents.
What is the best AI model for agents right now?
As of 23 Sep 2026, Claude Fable 5.1 ranks first, followed by GPT-6 Astra and Claude Opus 5. The ranking combines τ²-Bench (banking), Terminal-Bench 2.1 and LMArena Agent and is refreshed weekly.
What is the best open-weights model for agents?
GLM-5.3 from Z.ai is the highest-ranked open-weights model (8th of 41). Open weights can be run on your own infrastructure, including in an Australian region.
Which top agents model is cheapest?
Of the top ten, Gemini 3.8 Flash is cheapest at A$2.11 per million tokens (blended), against A$28.08 for Claude Fable 5.1. Prices are list prices converted at A$1 = US$0.7123, excluding GST.
Which of these models can keep data in Australia?
Of the top ten, Claude Opus 5 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.
How is this ranking calculated?
τ²-Bench is the closest public test to a business agent — a customer conversation where the model has tools and rules and must resolve the request correctly. Terminal-Bench covers operational tasks; LMArena’s agent arena adds human judgement of real sessions. A model needs at least two of the three to be ranked. Each input is rescaled 0–100 across the models we track, then weighted: τ²-Bench (banking) 40%, Terminal-Bench 2.1 30% and LMArena Agent 30%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →