Qwen3.8 Max benchmarks
Qwen3.8 Max is a proprietary model from Alibaba, released 3 Aug 2026. It ranks 16th of 51 on our overall leaderboard and 4th of 42 for coding. At A$4.21 per million tokens (blended), it costs 1.2× the median model we track. It writes about 42 tokens a second, and answers start after 50 s on average, thinking included. None of the major cloud platforms offer it in an Australian region yet.
- CONTEXT WINDOW
- 1,000,000 tokens
- MAX OUTPUT
- 131,072 tokens
- INPUT
- text, image, video
- WEIGHTS
- Proprietary
- RELEASED
- 3 Aug 2026
- OUTPUT SPEED
- 42 tok/s
- ANSWER STARTS AFTER
- 50 s (thinking included)
Qwen3.8 Max across our rankings.
Benchmark results.
Independent test scores are for the 0902 setting. “Among tracked” ranks Qwen3.8 Max against the 51 models on this site; LMArena ranks run across its full leaderboard.
| BENCHMARK | RESULT | AMONG TRACKED | SOURCE DETAIL |
|---|---|---|---|
| INDEPENDENT TESTS · ARTIFICIAL ANALYSIS | |||
| AA Intelligence Index | 45.4 | #11 of 51 | 0902 |
| AA Coding Index | 76.2 | #9 of 42 | 0902 |
| GPQA Diamond | 92.8% | #14 of 44 | 0902 |
| Humanity’s Last Exam | 43.1% | #20 of 51 | 0902 |
| AA-LCR | 80.3% | #24 of 51 | 0902 |
| Terminal-Bench 2.1 | 88.8% | #3 of 42 | 0902 |
| τ²-Bench (banking) | 47.8% | #5 of 41 | 0902 |
| SciCode | 52.1% | #29 of 46 | 0902 |
| BLIND HUMAN VOTES · LMARENA | |||
| LMArena Text16,670 votes · ±6 | 1481 | #19 of 42 | #22 of 402 on LMArena · 13 Sep 2026 |
| LMArena WebDev3,221 votes · ±13 | 1671 | #4 of 45 | #4 of 129 on LMArena · 22 Sep 2026 |
| LMArena Vision8,665 votes · ±8 | 1302 | #2 of 33 | #2 of 152 on LMArena · 13 Sep 2026 |
| LMArena Agent31,489 sessions | #17 | #15 of 37 | #17 of 46 on LMArena · 15 Sep 2026 |
What Qwen3.8 Max costs in Australian dollars.
List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.
| PER MILLION TOKENS | AUD | USD LIST |
|---|---|---|
| Input tokens | A$2.81 | US$2.00 |
| Output tokens | A$8.42 | US$6.00 |
| Blended (3 in : 1 out) | A$4.21 | US$3.00 |
| JOB | COST (AUD) | MEDIAN MODEL |
|---|---|---|
| Customer support replyPER 1,000 REPLIES | A$8.14 | ≈ A$6.58 |
| Summarise a 30-page documentPER 100 DOCUMENTS | A$6.29 | ≈ A$4.74 |
| Agentic coding taskPER 10 TASKS | A$4.89 | ≈ A$3.72 |
Running Qwen3.8 Max in Australia.
None of the major cloud platforms offer it in an Australian region yet. How the platforms compare →
| PLATFORM | AVAILABILITY | DETAIL |
|---|---|---|
| AWS Bedrock · Sydney | NOT OFFERED | |
| AWS Bedrock · Melbourne | NOT OFFERED | |
| Azure · Australia East | NOT OFFERED | |
| Google Vertex AI · Sydney | NOT OFFERED |
Compare it with.
Qwen3.8 Max, answered.
How much does Qwen3.8 Max cost in Australian dollars?
Qwen3.8 Max lists at US$2.00 per million input tokens and US$6.00 per million output tokens — A$2.81 and A$8.42 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$8.14 per 1,000 replies, before any reasoning tokens.
Can I use Qwen3.8 Max in Australia with data kept onshore?
None of the major cloud platforms offer it in an Australian region yet. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.
Is Qwen3.8 Max good for coding?
It ranks 4th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 76.2 and an LMArena WebDev rating of 1671 (4th of 129). The current leader is Claude Fable 5.1.
How good is Qwen3.8 Max at agent and automation work?
It ranks 7th of 41 on our agents ranking, completing 47.8% of τ²-Bench banking customer-service tasks and 88.8% of Terminal-Bench 2.1 tasks.
How fast is Qwen3.8 Max?
Artificial Analysis measures Qwen3.8 Max at about 42 tok/s of output, with the answer starting after 50 s on average once thinking time is included (0902 setting).
What is Qwen3.8 Max’s context window?
Alibaba documents a context window of 1,000,000 tokens (1M), with up to 131,072 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 80.3%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →