GPT-5.5 benchmarks

GPT-5.5 is a proprietary model from OpenAI, released 23 Apr 2026. It ranks 22nd of 51 on our overall leaderboard and 8th of 51 for long context. At A$15.79 per million tokens (blended), it costs 4.4× the median model we track. It can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia.

#22 OF 51 OVERALLA$15.79 / 1M TOKENS1.05M CONTEXTUPDATED 23 SEP 2026
CONTEXT WINDOW
1,050,000 tokens
MAX OUTPUT
128,000 tokens
INPUT
text, image
WEIGHTS
Proprietary
RELEASED
23 Apr 2026
OPENAI DOCS
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the xhigh effort setting. “Among tracked” ranks GPT-5.5 against the 51 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS7 OF 8 CORE SIGNALS
GPT-5.5 benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index38.4#27 of 51xhigh effort
AA Coding Index74.9#13 of 42xhigh effort
GPQA Diamond93.5%#8 of 44xhigh effort
Humanity’s Last Exam45.8%#17 of 51xhigh effort
AA-LCR84.3%#6 of 51xhigh effort
Terminal-Bench 2.184.3%#14 of 42xhigh effort
τ²-Bench (banking)39.0%#19 of 41xhigh effort
SciCode55.8%#20 of 46xhigh effort
BLIND HUMAN VOTES · LMARENA
LMArena Text64,924 votes · ±41482#17 of 42#20 of 402 on LMArena · 13 Sep 2026
LMArena WebDev14,952 votes · ±61510#33 of 45#41 of 129 on LMArena · 22 Sep 2026
LMArena Vision23,423 votes · ±61287#11 of 33#13 of 152 on LMArena · 13 Sep 2026
LMArena Document21,279 votes · ±61483#8 of 26#10 of 44 on LMArena · 13 Sep 2026
LMArena Search89,873 votes · ±61224#3 of 12#3 of 34 on LMArena · 24 Aug 2026
LMArena Agent53,054 sessions#10#9 of 37#10 of 46 on LMArena · 15 Sep 2026
§ 03 — COST

What GPT-5.5 costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
GPT-5.5 price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$7.02US$5.00
Output tokensA$42.12US$30.00
Blended (3 in : 1 out)A$15.79US$11.25
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$26.67A$6.58
Summarise a 30-page documentPER 100 DOCUMENTSA$17.41A$4.74
Agentic coding taskPER 10 TASKSA$13.90A$3.72
SETTINGSMEASURED SEPARATELY
GPT-5.5 settings compared
SETTINGINTELLIGENCEA$ / 1MSPEED
xhigh effort38.4A$15.79
high effort37.0A$15.79
medium effort33.8A$15.79
low effort30.7A$15.79
Non-reasoning23.2A$15.79
§ 04 — AUSTRALIA

Running GPT-5.5 in Australia.

It can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
GPT-5.5 availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyNOT IN AU

Provider docs, checked 23 Sep 2026

AWS Bedrock · MelbourneNOT IN AU

Provider docs, checked 23 Sep 2026

Azure · Australia EastGLOBAL

Provider docs, checked 23 Sep 2026

Google Vertex AI · SydneyNOT OFFERED
§ 06 — QUESTIONS

GPT-5.5, answered.

GPT-5.5 lists at US$5.00 per million input tokens and US$30.00 per million output tokens — A$7.02 and A$42.12 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$26.67 per 1,000 replies, before any reasoning tokens.

GPT-5.5 can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

It ranks 18th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 74.9 and an LMArena WebDev rating of 1510 (41st of 129). The current leader is Claude Fable 5.1.

It ranks 15th of 41 on our agents ranking, completing 39.0% of τ²-Bench banking customer-service tasks and 84.3% of Terminal-Bench 2.1 tasks.

OpenAI documents a context window of 1,050,000 tokens (1.05M), with up to 128,000 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 84.3%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?