DeepSeek V4 Flash benchmarks

DeepSeek V4 Flash is an open-weights model from DeepSeek, released 31 Jul 2026. It ranks 41st of 51 on our overall leaderboard and 16th of 50 for value, though some of its results are not yet published. At A$0.93 per million tokens (blended), it costs about 3.9× cheaper than the median model we track. It writes about 237 tokens a second, and answers start after 9.2 s on average, thinking included. It can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia.

#41 OF 51 OVERALLA$0.93 / 1M TOKENS1M CONTEXTUPDATED 23 SEP 2026
CONTEXT WINDOW
1,000,000 tokens
MAX OUTPUT
384,000 tokens
INPUT
text
WEIGHTS
Open (MIT)
RELEASED
31 Jul 2026
OUTPUT SPEED
237 tok/s
ANSWER STARTS AFTER
9.2 s (thinking included)
DEEPSEEK DOCS
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the reasoning, max effort setting. “Among tracked” ranks DeepSeek V4 Flash against the 51 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS7 OF 8 CORE SIGNALS
DeepSeek V4 Flash benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index34.3#32 of 51Reasoning, max effort
AA Coding Index69.1#27 of 42Reasoning, max effort
GPQA Diamond90.8%#30 of 44Reasoning, max effort
Humanity’s Last Exam38.6%#40 of 51Reasoning, max effort
AA-LCR79.7%#30 of 51Reasoning, max effort
Terminal-Bench 2.178.7%#24 of 42Reasoning, max effort
τ²-Bench (banking)39.4%#18 of 41Reasoning, max effort
SciCode50.3%#38 of 46Reasoning, max effort
BLIND HUMAN VOTES · LMARENA
LMArena WebDev4,739 votes · ±101580#20 of 45#24 of 129 on LMArena · 22 Sep 2026
LMArena Agent64,482 sessions#22#19 of 37#22 of 46 on LMArena · 15 Sep 2026
§ 03 — COST

What DeepSeek V4 Flash costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
DeepSeek V4 Flash price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$0.62US$0.44
Output tokensA$1.85US$1.32
Blended (3 in : 1 out)A$0.93US$0.66
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$1.79A$6.58
Summarise a 30-page documentPER 100 DOCUMENTSA$1.38A$4.74
Agentic coding taskPER 10 TASKSA$1.07A$3.72
SETTINGSMEASURED SEPARATELY
DeepSeek V4 Flash settings compared
SETTINGINTELLIGENCEA$ / 1MSPEED
Reasoning, max effort34.3A$0.93
Reasoning, max effort34.8A$0.93237 tok/s
§ 04 — AUSTRALIA

Running DeepSeek V4 Flash in Australia.

It can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
DeepSeek V4 Flash availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyNOT OFFERED
AWS Bedrock · MelbourneNOT OFFERED
Azure · Australia EastGLOBAL

Provider docs, checked 23 Sep 2026

Google Vertex AI · SydneyNOT OFFERED
Your own infrastructureSELF-HOST

Open weights: run it on your own servers or GPU instances in an Australian region.

§ 06 — QUESTIONS

DeepSeek V4 Flash, answered.

DeepSeek V4 Flash lists at US$0.44 per million input tokens and US$1.32 per million output tokens — A$0.62 and A$1.85 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$1.79 per 1,000 replies, before any reasoning tokens.

DeepSeek V4 Flash can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

It ranks 24th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 69.1 and an LMArena WebDev rating of 1580 (24th of 129). The current leader is Claude Fable 5.1.

It ranks 22nd of 41 on our agents ranking, completing 39.4% of τ²-Bench banking customer-service tasks and 78.7% of Terminal-Bench 2.1 tasks.

Artificial Analysis measures DeepSeek V4 Flash at about 237 tok/s of output, with the answer starting after 9.2 s on average once thinking time is included (reasoning, max effort setting).

DeepSeek documents a context window of 1,000,000 tokens (1M), with up to 384,000 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 79.7%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?