Claude Sonnet 5.5 benchmarks

Claude Sonnet 5.5 is a proprietary model from Anthropic, released 28 Sep 2026. It ranks 1st of 54 on our overall leaderboard and 2nd of 54 for reasoning, though some of its results are not yet published. At A$5.76 per million tokens (blended), it costs 1.3× the median model we track. It writes about 145 tokens a second, and answers start after 4.6 min on average, thinking included. It can be called from AWS Bedrock in Sydney and Melbourne and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia.

#1 OF 54 OVERALLA$5.76 / 1M TOKENS1M CONTEXTUPDATED 2 OCT 2026
CONTEXT WINDOW
1,000,000 tokens
MAX OUTPUT
128,000 tokens
INPUT
text, image
WEIGHTS
Proprietary
RELEASED
28 Sep 2026
OUTPUT SPEED
145 tok/s
ANSWER STARTS AFTER
4.6 min (thinking included)
ANTHROPIC DOCS ↗
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the adaptive reasoning, max effort setting. “Among tracked” ranks Claude Sonnet 5.5 against the 54 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS5 OF 8 CORE SIGNALS
Claude Sonnet 5.5 benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index56.0#2 of 54Adaptive reasoning, max effort
Humanity’s Last Exam55.0%#5 of 54Adaptive reasoning, max effort
AA-LCR82.7%#16 of 54Adaptive reasoning, max effort
SciCode61.0%#4 of 49Adaptive reasoning, max effort
BLIND HUMAN VOTES · LMARENA
LMArena WebDev1,810 votes · ±161709#5 of 52#5 of 137 on LMArena · 30 Sep 2026
§ 03 — COST

What Claude Sonnet 5.5 costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
Claude Sonnet 5.5 price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$2.88US$2.00
Output tokensA$14.39US$10.00
Blended (3 in : 1 out)A$5.76US$4.00
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$10.07≈ A$8.20
Summarise a 30-page documentPER 100 DOCUMENTSA$6.91≈ A$5.35
Agentic coding taskPER 10 TASKSA$5.47≈ A$4.27
SETTINGSMEASURED SEPARATELY
Claude Sonnet 5.5 settings compared
SETTINGINTELLIGENCEA$ / 1MSPEED
Adaptive reasoning, max effort56.0A$5.76145 tok/s
Adaptive reasoning, xhigh effort51.9A$5.76109 tok/s
Adaptive reasoning, high effort46.7A$5.76108 tok/s
Adaptive reasoning, medium effort40.7A$5.76109 tok/s
Adaptive reasoning, low effortA$5.76104 tok/s
§ 04 — AUSTRALIA

Running Claude Sonnet 5.5 in Australia.

It can be called from AWS Bedrock in Sydney and Melbourne and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
Claude Sonnet 5.5 availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyGLOBAL

Provider docs, checked 2 Oct 2026

AWS Bedrock · MelbourneGLOBAL

Provider docs, checked 2 Oct 2026

Azure · Australia EastNOT IN AU

Provider docs, checked 2 Oct 2026

Google Vertex AI · SydneyGLOBAL

Provider docs, checked 2 Oct 2026

§ 06 — QUESTIONS

Claude Sonnet 5.5, answered.

Claude Sonnet 5.5 lists at US$2.00 per million input tokens and US$10.00 per million output tokens — A$2.88 and A$14.39 at A$1 = US$0.6949 (RBA, 1 Oct 2026), excluding GST. A typical customer support reply works out at about A$10.07 per 1,000 replies, before any reasoning tokens.

Claude Sonnet 5.5 can be called from AWS Bedrock in Sydney and Melbourne and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

Artificial Analysis measures Claude Sonnet 5.5 at about 145 tok/s of output, with the answer starting after 4.6 min on average once thinking time is included (adaptive reasoning, max effort setting).

Anthropic documents a context window of 1,000,000 tokens (1M), with up to 128,000 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 82.7%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?