How the benchmarks are built.

Where every number comes from, how the rankings are calculated, how prices become Australian dollars — and what none of it can tell you about your own work.

SNAPSHOT 23 SEP 2026UPDATED WEEKLY3 SOURCES
§ 01 — SOURCES

Three sources, all public.

We don’t run benchmarks ourselves. We combine two independent sources that measure models in different ways — controlled tests and blind human votes — and convert prices with the Reserve Bank’s published rate.

LMArena

Blind-vote ratings for text, web development, vision, search, documents and agents. Licence: CC BY 4.0.

SOURCE
Artificial Analysis

Independent evaluations (Intelligence and Coding indices, GPQA, HLE, AA-LCR, Terminal-Bench, τ²-Bench), list prices, output speed and latency. Licence: Used with attribution under the Artificial Analysis API terms.

SOURCE
Reserve Bank of Australia

The AUD/USD rate (table F11.1) used to convert US-dollar list prices. Licence: RBA statistical tables.

SOURCE
§ 02 — WHAT WE TRACK

A curated list, matched by exact name.

We track 51 models from 14 labs: the current and recent frontier models from the major labs, plus the strongest open-weights releases. 39 have their own page. Source rows are matched to a model only when their exact name is on our list — nothing is matched by guesswork — and new high-ranking models are flagged for review each week rather than added automatically. Where a source measures several settings of one model (reasoning effort levels), the headline figures come from the maker’s default or highest setting, and every setting is listed on the model’s page.

§ 03 — HOW SCORES WORK

Rescaled, weighted, compared.

Each measure is rescaled from 0 to 100 across the models we track: the best tracked model scores 100 and the lowest scores 0. Prices use a logarithmic scale with the cheapest at 100, so a tenfold price difference counts the same at any price level. A ranking’s score is the weighted average of the measures a model has, with the weights re-balanced when one is missing; its anchor measure is always required. Scores are relative — they say how a model compares with the others here, not how good it is in absolute terms.

RANKING DEFINITIONS6 RANKINGS
How each ranking is weighted
RANKINGINPUTS AND WEIGHTSREQUIRED
OverallAA Intelligence Index 50% · LMArena Text 50%AA Intelligence Index or LMArena Text
CodingAA Coding Index 60% · LMArena WebDev 40%AA Coding Index
Agentsτ²-Bench (banking) 40% · Terminal-Bench 2.1 30% · LMArena Agent 30%τ²-Bench (banking) or Terminal-Bench 2.1; at least 2 inputs
ReasoningAA Intelligence Index 50% · Humanity’s Last Exam 25% · GPQA Diamond 25%AA Intelligence Index; at least 2 inputs
Long contextAA-LCR 70% · LMArena Document 30%AA-LCR; documented context ≥ 128K
ValueAA Intelligence Index 60% · Blended price (AUD) 40%AA Intelligence Index; at least 2 inputs
§ 04 — PRICES IN AUD

List prices, in Australian dollars.

Providers publish US-dollar prices per million input and output tokens. We convert them at the RBA’s AUD/USD rate for the snapshot (A$1 = US$0.7123 on 22 Sep 2026), excluding GST. The blended price weights input and output 3 to 1, a typical mix for chat and document work. Reasoning models also bill their thinking as output tokens, which list prices and our job estimates don’t include — at high effort settings that can multiply the real cost.

Customer support reply

About 2,000 tokens of instructions, history and knowledge-base text in; a 300-token reply out. Priced per 1,000 replies.

Summarise a 30-page document

About 20,000 tokens in; an 800-token summary out. Priced per 100 documents.

Agentic coding task

About 150,000 tokens of code and tool output read across the task; 8,000 tokens written. Priced per 10 tasks.

§ 05 — AUSTRALIAN AVAILABILITY

Checked by hand, with the source.

For each model with a page we check the AWS Bedrock, Microsoft Azure and Google Vertex AI documentation for Sydney, Melbourne and Australia East, and record the page and the date we checked. Open-weights models are also marked self-hostable. This tracks what the platforms document, not a legal view of your obligations. See the Australia page →

IN REGIONIn region

Served from that Australian region. Requests are processed there.

AU ROUTINGAustralia routing

Routed between Australian regions only (on Bedrock, the “au.” inference profiles across Sydney and Melbourne). Data stays in Australia.

GLOBALGlobal only

You can call it from an Australian region, but requests may be processed anywhere the provider runs capacity.

NOT IN AUNot in Australia

The platform offers the model, but not from this Australian region.

§ 06 — UPDATES AND CHECKS

Weekly, and never silently wrong.

Every Monday the sources are fetched and checked before anything is published: each must return a plausible number of models, values in sensible ranges and dates that don’t go backwards. A source that fails keeps its previous figures, shown with their date. A model missing from its sources for a week is marked stale with figures carried forward; after eight weeks it leaves the rankings, but its page stays up, marked archived.

§ 07 — REUSE

Quoting these rankings.

You’re welcome to quote the rankings, scores and AUD figures with a link to the page you took them from. The underlying LMArena ratings are published under CC BY 4.0; Artificial Analysis evaluations, prices and speed figures are theirs and used with attribution; the exchange rate is published by the Reserve Bank of Australia.

§ 08 — LIMITS

What benchmarks can’t tell you.

Public benchmarks measure general ability on someone else’s tasks. They don’t know your documents, your customers or your processes, and models are tuned to do well on popular tests. Arena votes reflect the people who vote. Treat these rankings as a shortlist, then test the two or three front-runners on real examples of your own work before committing.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?