GPT-5.6 Luna benchmarks

GPT-5.6 Luna is a proprietary model from OpenAI, released 9 Jul 2026. It ranks 42nd of 51 on our overall leaderboard and 8th of 50 for value. At A$0.63 per million tokens (blended), it costs about 5.7× cheaper than the median model we track. It writes about 145 tokens a second, and answers start after 1.5 min on average, thinking included. It can be called from AWS Bedrock in Sydney and Melbourne and Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia.

#42 OF 51 OVERALLA$0.63 / 1M TOKENS1.05M CONTEXTUPDATED 23 SEP 2026
CONTEXT WINDOW
1,050,000 tokens
MAX OUTPUT
128,000 tokens
INPUT
text, image
WEIGHTS
Proprietary
RELEASED
9 Jul 2026
OUTPUT SPEED
145 tok/s
ANSWER STARTS AFTER
1.5 min (thinking included)
OPENAI DOCS
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the max effort setting. “Among tracked” ranks GPT-5.6 Luna against the 51 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS8 OF 8 CORE SIGNALS
GPT-5.6 Luna benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index37.3#29 of 51max effort
AA Coding Index71.4#22 of 42max effort
GPQA Diamond91.1%#27 of 44max effort
Humanity’s Last Exam39.5%#37 of 51max effort
AA-LCR83.7%#9 of 51max effort
Terminal-Bench 2.180.9%#20 of 42max effort
τ²-Bench (banking)31.1%#30 of 41max effort
SciCode53.6%#27 of 46max effort
BLIND HUMAN VOTES · LMARENA
LMArena Text28,547 votes · ±51453#37 of 42#67 of 402 on LMArena · 13 Sep 2026
LMArena WebDev10,128 votes · ±71520#30 of 45#37 of 129 on LMArena · 22 Sep 2026
LMArena Vision7,793 votes · ±81258#29 of 33#40 of 152 on LMArena · 13 Sep 2026
LMArena Document5,766 votes · ±81462#18 of 26#23 of 44 on LMArena · 13 Sep 2026
LMArena Agent29,186 sessions#27#24 of 37#27 of 46 on LMArena · 15 Sep 2026
§ 03 — COST

What GPT-5.6 Luna costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
GPT-5.6 Luna price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$0.28US$0.20
Output tokensA$1.68US$1.20
Blended (3 in : 1 out)A$0.63US$0.45
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$1.07A$6.58
Summarise a 30-page documentPER 100 DOCUMENTSA$0.70A$4.74
Agentic coding taskPER 10 TASKSA$0.56A$3.72
SETTINGSMEASURED SEPARATELY
GPT-5.6 Luna settings compared
SETTINGINTELLIGENCEA$ / 1MSPEED
max effort37.3A$0.63145 tok/s
xhigh effort34.6A$0.63141 tok/s
high effort32.1A$0.63137 tok/s
medium effort25.0A$0.63134 tok/s
low effort21.0A$0.63138 tok/s
Non-reasoning15.5A$0.63127 tok/s
§ 04 — AUSTRALIA

Running GPT-5.6 Luna in Australia.

It can be called from AWS Bedrock in Sydney and Melbourne and Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
GPT-5.6 Luna availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyGLOBAL

Provider docs, checked 23 Sep 2026

AWS Bedrock · MelbourneGLOBAL

Provider docs, checked 23 Sep 2026

Azure · Australia EastGLOBAL

Provider docs, checked 23 Sep 2026

Google Vertex AI · SydneyNOT OFFERED
§ 06 — QUESTIONS

GPT-5.6 Luna, answered.

GPT-5.6 Luna lists at US$0.20 per million input tokens and US$1.20 per million output tokens — A$0.28 and A$1.68 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$1.07 per 1,000 replies, before any reasoning tokens.

GPT-5.6 Luna can be called from AWS Bedrock in Sydney and Melbourne and Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

It ranks 27th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 71.4 and an LMArena WebDev rating of 1520 (37th of 129). The current leader is Claude Fable 5.1.

It ranks 28th of 41 on our agents ranking, completing 31.1% of τ²-Bench banking customer-service tasks and 80.9% of Terminal-Bench 2.1 tasks.

Artificial Analysis measures GPT-5.6 Luna at about 145 tok/s of output, with the answer starting after 1.5 min on average once thinking time is included (max effort setting).

OpenAI documents a context window of 1,050,000 tokens (1.05M), with up to 128,000 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 83.7%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?