Grok 4.6 benchmarks

Grok 4.6 is a proprietary model from xAI, released 12 Aug 2026. It ranks 35th of 51 on our overall leaderboard and 4th of 41 for agents. At A$4.21 per million tokens (blended), it costs 1.2× the median model we track. It writes about 70 tokens a second, and answers start after 35 s on average, thinking included. It can be called from AWS Bedrock in Sydney and Melbourne, Microsoft Azure in Australia East and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia.

#35 OF 51 OVERALLA$4.21 / 1M TOKENS500K CONTEXTUPDATED 23 SEP 2026
CONTEXT WINDOW
500,000 tokens
INPUT
text, image
WEIGHTS
Proprietary
RELEASED
12 Aug 2026
OUTPUT SPEED
70 tok/s
ANSWER STARTS AFTER
35 s (thinking included)
XAI DOCS
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the high effort setting. “Among tracked” ranks Grok 4.6 against the 51 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS8 OF 8 CORE SIGNALS
Grok 4.6 benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index44.3#13 of 51high effort
AA Coding Index76.8#5 of 42high effort
GPQA Diamond94.9%#3 of 44high effort
Humanity’s Last Exam42.9%#22 of 51high effort
AA-LCR80.3%#24 of 51high effort
Terminal-Bench 2.188.4%#4 of 42high effort
τ²-Bench (banking)50.7%#1 of 41high effort
SciCode56.5%#17 of 46high effort
BLIND HUMAN VOTES · LMARENA
LMArena Text15,521 votes · ±61456#36 of 42#63 of 402 on LMArena · 13 Sep 2026
LMArena WebDev6,447 votes · ±91616#12 of 45#16 of 129 on LMArena · 22 Sep 2026
LMArena Vision3,632 votes · ±111265#26 of 33#35 of 152 on LMArena · 13 Sep 2026
LMArena Document2,258 votes · ±121461#19 of 26#24 of 44 on LMArena · 13 Sep 2026
LMArena Agent22,548 sessions#21#18 of 37#21 of 46 on LMArena · 15 Sep 2026
§ 03 — COST

What Grok 4.6 costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
Grok 4.6 price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$2.81US$2.00
Output tokensA$8.42US$6.00
Blended (3 in : 1 out)A$4.21US$3.00
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$8.14A$6.58
Summarise a 30-page documentPER 100 DOCUMENTSA$6.29A$4.74
Agentic coding taskPER 10 TASKSA$4.89A$3.72
SETTINGSMEASURED SEPARATELY
Grok 4.6 settings compared
SETTINGINTELLIGENCEA$ / 1MSPEED
high effort44.3A$4.2170 tok/s
xhigh effort44.2A$4.2160 tok/s
medium effort42.8A$4.2165 tok/s
low effort35.1A$4.2152 tok/s
§ 04 — AUSTRALIA

Running Grok 4.6 in Australia.

It can be called from AWS Bedrock in Sydney and Melbourne, Microsoft Azure in Australia East and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
Grok 4.6 availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyGLOBAL

Provider docs, checked 23 Sep 2026

AWS Bedrock · MelbourneGLOBAL

Provider docs, checked 23 Sep 2026

Azure · Australia EastGLOBAL

Provider docs, checked 23 Sep 2026

Google Vertex AI · SydneyGLOBAL

Provider docs, checked 23 Sep 2026

§ 06 — QUESTIONS

Grok 4.6, answered.

Grok 4.6 lists at US$2.00 per million input tokens and US$6.00 per million output tokens — A$2.81 and A$8.42 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$8.14 per 1,000 replies, before any reasoning tokens.

Grok 4.6 can be called from AWS Bedrock in Sydney and Melbourne, Microsoft Azure in Australia East and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

It ranks 9th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 76.8 and an LMArena WebDev rating of 1616 (16th of 129). The current leader is Claude Fable 5.1.

It ranks 4th of 41 on our agents ranking, completing 50.7% of τ²-Bench banking customer-service tasks and 88.4% of Terminal-Bench 2.1 tasks.

Artificial Analysis measures Grok 4.6 at about 70 tok/s of output, with the answer starting after 35 s on average once thinking time is included (high effort setting).

xAI documents a context window of 500,000 tokens (500K). On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 80.3%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?