GPT-5.4 benchmarks
GPT-5.4 is a proprietary model from OpenAI, released 5 Mar 2026. It ranks 24th of 51 on our overall leaderboard and 18th of 51 for long context. At A$7.90 per million tokens (blended), it costs 2.2× the median model we track. It can run with requests kept in Australia on Microsoft Azure in Australia East.
- CONTEXT WINDOW
- 1,050,000 tokens
- MAX OUTPUT
- 128,000 tokens
- INPUT
- text, image
- WEIGHTS
- Proprietary
- RELEASED
- 5 Mar 2026
GPT-5.4 across our rankings.
Benchmark results.
Independent test scores are for the xhigh effort setting. “Among tracked” ranks GPT-5.4 against the 51 models on this site; LMArena ranks run across its full leaderboard.
| BENCHMARK | RESULT | AMONG TRACKED | SOURCE DETAIL |
|---|---|---|---|
| INDEPENDENT TESTS · ARTIFICIAL ANALYSIS | |||
| AA Intelligence Index | 39.0 | #25 of 51 | xhigh effort |
| AA Coding Index | 71.1 | #24 of 42 | xhigh effort |
| GPQA Diamond | 92.0% | #22 of 44 | xhigh effort |
| Humanity’s Last Exam | 43.7% | #19 of 51 | xhigh effort |
| AA-LCR | 82.0% | #16 of 51 | xhigh effort |
| Terminal-Bench 2.1 | 78.3% | #27 of 42 | xhigh effort |
| τ²-Bench (banking) | 39.6% | #16 of 41 | xhigh effort |
| BLIND HUMAN VOTES · LMARENA | |||
| LMArena Text60,537 votes · ±4 | 1476 | #23 of 42 | #26 of 402 on LMArena · 13 Sep 2026 |
| LMArena WebDev1,275 votes · ±19 | 1463 | #39 of 45 | #55 of 129 on LMArena · 22 Sep 2026 |
| LMArena Vision25,357 votes · ±6 | 1285 | #13 of 33 | #15 of 152 on LMArena · 13 Sep 2026 |
| LMArena Document33,331 votes · ±6 | 1470 | #14 of 26 | #18 of 44 on LMArena · 13 Sep 2026 |
| LMArena Search110,116 votes · ±6 | 1195 | #12 of 12 | #14 of 34 on LMArena · 24 Aug 2026 |
| LMArena Agent80,406 sessions | #24 | #21 of 37 | #24 of 46 on LMArena · 15 Sep 2026 |
What GPT-5.4 costs in Australian dollars.
List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.
| PER MILLION TOKENS | AUD | USD LIST |
|---|---|---|
| Input tokens | A$3.51 | US$2.50 |
| Output tokens | A$21.06 | US$15.00 |
| Blended (3 in : 1 out) | A$7.90 | US$5.63 |
| JOB | COST (AUD) | MEDIAN MODEL |
|---|---|---|
| Customer support replyPER 1,000 REPLIES | A$13.34 | ≈ A$6.58 |
| Summarise a 30-page documentPER 100 DOCUMENTS | A$8.70 | ≈ A$4.74 |
| Agentic coding taskPER 10 TASKS | A$6.95 | ≈ A$3.72 |
| SETTING | INTELLIGENCE | A$ / 1M | SPEED |
|---|---|---|---|
| xhigh effort | 39.0 | A$7.90 | |
| low effort | 27.6 | A$7.90 | |
| Non-reasoning | 18.2 | A$7.90 |
Running GPT-5.4 in Australia.
It can run with requests kept in Australia on Microsoft Azure in Australia East. How the platforms compare →
| PLATFORM | AVAILABILITY | DETAIL |
|---|---|---|
| AWS Bedrock · Sydney | NOT IN AU | |
| AWS Bedrock · Melbourne | NOT IN AU | |
| Azure · Australia East | IN REGION | Reserved capacity (Regional Provisioned) only; pay-as-you-go deployments in Australia East are global. Provider docs, checked 23 Sep 2026 |
| Google Vertex AI · Sydney | NOT OFFERED |
Compare it with.
GPT-5.4, answered.
How much does GPT-5.4 cost in Australian dollars?
GPT-5.4 lists at US$2.50 per million input tokens and US$15.00 per million output tokens — A$3.51 and A$21.06 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$13.34 per 1,000 replies, before any reasoning tokens.
Can I use GPT-5.4 in Australia with data kept onshore?
GPT-5.4 can run with requests kept in Australia on Microsoft Azure in Australia East. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.
Is GPT-5.4 good for coding?
It ranks 30th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 71.1 and an LMArena WebDev rating of 1463 (55th of 129). The current leader is Claude Fable 5.1.
How good is GPT-5.4 at agent and automation work?
It ranks 23rd of 41 on our agents ranking, completing 39.6% of τ²-Bench banking customer-service tasks and 78.3% of Terminal-Bench 2.1 tasks.
What is GPT-5.4’s context window?
OpenAI documents a context window of 1,050,000 tokens (1.05M), with up to 128,000 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 82.0%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →