Grok 4.6 benchmarks
Grok 4.6 is a proprietary model from xAI, released 12 Aug 2026. It ranks 35th of 51 on our overall leaderboard and 4th of 41 for agents. At A$4.21 per million tokens (blended), it costs 1.2× the median model we track. It writes about 70 tokens a second, and answers start after 35 s on average, thinking included. It can be called from AWS Bedrock in Sydney and Melbourne, Microsoft Azure in Australia East and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia.
- CONTEXT WINDOW
- 500,000 tokens
- INPUT
- text, image
- WEIGHTS
- Proprietary
- RELEASED
- 12 Aug 2026
- OUTPUT SPEED
- 70 tok/s
- ANSWER STARTS AFTER
- 35 s (thinking included)
Grok 4.6 across our rankings.
Benchmark results.
Independent test scores are for the high effort setting. “Among tracked” ranks Grok 4.6 against the 51 models on this site; LMArena ranks run across its full leaderboard.
| BENCHMARK | RESULT | AMONG TRACKED | SOURCE DETAIL |
|---|---|---|---|
| INDEPENDENT TESTS · ARTIFICIAL ANALYSIS | |||
| AA Intelligence Index | 44.3 | #13 of 51 | high effort |
| AA Coding Index | 76.8 | #5 of 42 | high effort |
| GPQA Diamond | 94.9% | #3 of 44 | high effort |
| Humanity’s Last Exam | 42.9% | #22 of 51 | high effort |
| AA-LCR | 80.3% | #24 of 51 | high effort |
| Terminal-Bench 2.1 | 88.4% | #4 of 42 | high effort |
| τ²-Bench (banking) | 50.7% | #1 of 41 | high effort |
| SciCode | 56.5% | #17 of 46 | high effort |
| BLIND HUMAN VOTES · LMARENA | |||
| LMArena Text15,521 votes · ±6 | 1456 | #36 of 42 | #63 of 402 on LMArena · 13 Sep 2026 |
| LMArena WebDev6,447 votes · ±9 | 1616 | #12 of 45 | #16 of 129 on LMArena · 22 Sep 2026 |
| LMArena Vision3,632 votes · ±11 | 1265 | #26 of 33 | #35 of 152 on LMArena · 13 Sep 2026 |
| LMArena Document2,258 votes · ±12 | 1461 | #19 of 26 | #24 of 44 on LMArena · 13 Sep 2026 |
| LMArena Agent22,548 sessions | #21 | #18 of 37 | #21 of 46 on LMArena · 15 Sep 2026 |
What Grok 4.6 costs in Australian dollars.
List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.
| PER MILLION TOKENS | AUD | USD LIST |
|---|---|---|
| Input tokens | A$2.81 | US$2.00 |
| Output tokens | A$8.42 | US$6.00 |
| Blended (3 in : 1 out) | A$4.21 | US$3.00 |
| JOB | COST (AUD) | MEDIAN MODEL |
|---|---|---|
| Customer support replyPER 1,000 REPLIES | A$8.14 | ≈ A$6.58 |
| Summarise a 30-page documentPER 100 DOCUMENTS | A$6.29 | ≈ A$4.74 |
| Agentic coding taskPER 10 TASKS | A$4.89 | ≈ A$3.72 |
| SETTING | INTELLIGENCE | A$ / 1M | SPEED |
|---|---|---|---|
| high effort | 44.3 | A$4.21 | 70 tok/s |
| xhigh effort | 44.2 | A$4.21 | 60 tok/s |
| medium effort | 42.8 | A$4.21 | 65 tok/s |
| low effort | 35.1 | A$4.21 | 52 tok/s |
Running Grok 4.6 in Australia.
It can be called from AWS Bedrock in Sydney and Melbourne, Microsoft Azure in Australia East and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia. How the platforms compare →
| PLATFORM | AVAILABILITY | DETAIL |
|---|---|---|
| AWS Bedrock · Sydney | GLOBAL | |
| AWS Bedrock · Melbourne | GLOBAL | |
| Azure · Australia East | GLOBAL | |
| Google Vertex AI · Sydney | GLOBAL |
Compare it with.
Grok 4.6, answered.
How much does Grok 4.6 cost in Australian dollars?
Grok 4.6 lists at US$2.00 per million input tokens and US$6.00 per million output tokens — A$2.81 and A$8.42 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$8.14 per 1,000 replies, before any reasoning tokens.
Can I use Grok 4.6 in Australia with data kept onshore?
Grok 4.6 can be called from AWS Bedrock in Sydney and Melbourne, Microsoft Azure in Australia East and Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.
Is Grok 4.6 good for coding?
It ranks 9th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 76.8 and an LMArena WebDev rating of 1616 (16th of 129). The current leader is Claude Fable 5.1.
How good is Grok 4.6 at agent and automation work?
It ranks 4th of 41 on our agents ranking, completing 50.7% of τ²-Bench banking customer-service tasks and 88.4% of Terminal-Bench 2.1 tasks.
How fast is Grok 4.6?
Artificial Analysis measures Grok 4.6 at about 70 tok/s of output, with the answer starting after 35 s on average once thinking time is included (high effort setting).
What is Grok 4.6’s context window?
xAI documents a context window of 500,000 tokens (500K). On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 80.3%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →