Qwen3.8 27B benchmarks

Qwen3.8 27B is an open-weights model from Alibaba, released 14 Aug 2026. It ranks 46th of 51 on our overall leaderboard and 13th of 51 for long context. At A$1.58 per million tokens (blended), it costs about 2.3× cheaper than the median model we track. It writes about 46 tokens a second, and answers start after 45 s on average, thinking included. Its weights are public, so it can be hosted in Australia on your own infrastructure or an Australian cloud region.

#46 OF 51 OVERALLA$1.58 / 1M TOKENS1M CONTEXTUPDATED 23 SEP 2026
CONTEXT WINDOW
1,000,000 tokens
MAX OUTPUT
131,072 tokens
INPUT
text, image, video
WEIGHTS
Open (Apache 2.0)
RELEASED
14 Aug 2026
OUTPUT SPEED
46 tok/s
ANSWER STARTS AFTER
45 s (thinking included)
ALIBABA DOCS
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the xhigh effort setting. “Among tracked” ranks Qwen3.8 27B against the 51 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS8 OF 8 CORE SIGNALS
Qwen3.8 27B benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index33.7#34 of 51xhigh effort
AA Coding Index68.1#31 of 42xhigh effort
GPQA Diamond90.5%#32 of 44xhigh effort
Humanity’s Last Exam33.9%#45 of 51xhigh effort
AA-LCR82.0%#16 of 51xhigh effort
Terminal-Bench 2.179.8%#23 of 42xhigh effort
τ²-Bench (banking)48.0%#4 of 41xhigh effort
SciCode46.6%#43 of 46xhigh effort
BLIND HUMAN VOTES · LMARENA
LMArena Text10,697 votes · ±71437#41 of 42#90 of 402 on LMArena · 13 Sep 2026
LMArena WebDev9,092 votes · ±81591#17 of 45#21 of 129 on LMArena · 22 Sep 2026
LMArena Vision4,577 votes · ±101244#31 of 33#49 of 152 on LMArena · 13 Sep 2026
LMArena Agent37,257 sessions#29#26 of 37#29 of 46 on LMArena · 15 Sep 2026
§ 03 — COST

What Qwen3.8 27B costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
Qwen3.8 27B price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$0.70US$0.50
Output tokensA$4.21US$3.00
Blended (3 in : 1 out)A$1.58US$1.13
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$2.67A$6.58
Summarise a 30-page documentPER 100 DOCUMENTSA$1.74A$4.74
Agentic coding taskPER 10 TASKSA$1.39A$3.72
SETTINGSMEASURED SEPARATELY
Qwen3.8 27B settings compared
SETTINGINTELLIGENCEA$ / 1MSPEED
xhigh effort33.7A$1.5846 tok/s
medium effort27.6A$1.5847 tok/s
low effort26.2A$1.5852 tok/s
Non-reasoning20.2A$1.5849 tok/s
§ 04 — AUSTRALIA

Running Qwen3.8 27B in Australia.

Its weights are public, so it can be hosted in Australia on your own infrastructure or an Australian cloud region. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
Qwen3.8 27B availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyNOT OFFERED
AWS Bedrock · MelbourneNOT OFFERED
Azure · Australia EastNOT OFFERED
Google Vertex AI · SydneyNOT OFFERED
Your own infrastructureSELF-HOST

Open weights: run it on your own servers or GPU instances in an Australian region.

§ 06 — QUESTIONS

Qwen3.8 27B, answered.

Qwen3.8 27B lists at US$0.50 per million input tokens and US$3.00 per million output tokens — A$0.70 and A$4.21 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$2.67 per 1,000 replies, before any reasoning tokens.

Qwen3.8 27B’s weights are public, so it can be hosted in Australia on your own infrastructure or an Australian cloud region. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

It ranks 26th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 68.1 and an LMArena WebDev rating of 1591 (21st of 129). The current leader is Claude Fable 5.1.

It ranks 14th of 41 on our agents ranking, completing 48.0% of τ²-Bench banking customer-service tasks and 79.8% of Terminal-Bench 2.1 tasks.

Artificial Analysis measures Qwen3.8 27B at about 46 tok/s of output, with the answer starting after 45 s on average once thinking time is included (xhigh effort setting).

Alibaba documents a context window of 1,000,000 tokens (1M), with up to 131,072 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 82.0%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?