DeepSeek V4 Flash benchmarks
DeepSeek V4 Flash is an open-weights model from DeepSeek, released 31 Jul 2026. It ranks 41st of 51 on our overall leaderboard and 16th of 50 for value, though some of its results are not yet published. At A$0.93 per million tokens (blended), it costs about 3.9× cheaper than the median model we track. It writes about 237 tokens a second, and answers start after 9.2 s on average, thinking included. It can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia.
- CONTEXT WINDOW
- 1,000,000 tokens
- MAX OUTPUT
- 384,000 tokens
- INPUT
- text
- WEIGHTS
- Open (MIT)
- RELEASED
- 31 Jul 2026
- OUTPUT SPEED
- 237 tok/s
- ANSWER STARTS AFTER
- 9.2 s (thinking included)
DeepSeek V4 Flash across our rankings.
Benchmark results.
Independent test scores are for the reasoning, max effort setting. “Among tracked” ranks DeepSeek V4 Flash against the 51 models on this site; LMArena ranks run across its full leaderboard.
| BENCHMARK | RESULT | AMONG TRACKED | SOURCE DETAIL |
|---|---|---|---|
| INDEPENDENT TESTS · ARTIFICIAL ANALYSIS | |||
| AA Intelligence Index | 34.3 | #32 of 51 | Reasoning, max effort |
| AA Coding Index | 69.1 | #27 of 42 | Reasoning, max effort |
| GPQA Diamond | 90.8% | #30 of 44 | Reasoning, max effort |
| Humanity’s Last Exam | 38.6% | #40 of 51 | Reasoning, max effort |
| AA-LCR | 79.7% | #30 of 51 | Reasoning, max effort |
| Terminal-Bench 2.1 | 78.7% | #24 of 42 | Reasoning, max effort |
| τ²-Bench (banking) | 39.4% | #18 of 41 | Reasoning, max effort |
| SciCode | 50.3% | #38 of 46 | Reasoning, max effort |
| BLIND HUMAN VOTES · LMARENA | |||
| LMArena WebDev4,739 votes · ±10 | 1580 | #20 of 45 | #24 of 129 on LMArena · 22 Sep 2026 |
| LMArena Agent64,482 sessions | #22 | #19 of 37 | #22 of 46 on LMArena · 15 Sep 2026 |
What DeepSeek V4 Flash costs in Australian dollars.
List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.
| PER MILLION TOKENS | AUD | USD LIST |
|---|---|---|
| Input tokens | A$0.62 | US$0.44 |
| Output tokens | A$1.85 | US$1.32 |
| Blended (3 in : 1 out) | A$0.93 | US$0.66 |
| JOB | COST (AUD) | MEDIAN MODEL |
|---|---|---|
| Customer support replyPER 1,000 REPLIES | A$1.79 | ≈ A$6.58 |
| Summarise a 30-page documentPER 100 DOCUMENTS | A$1.38 | ≈ A$4.74 |
| Agentic coding taskPER 10 TASKS | A$1.07 | ≈ A$3.72 |
| SETTING | INTELLIGENCE | A$ / 1M | SPEED |
|---|---|---|---|
| Reasoning, max effort | 34.3 | A$0.93 | |
| Reasoning, max effort | 34.8 | A$0.93 | 237 tok/s |
Running DeepSeek V4 Flash in Australia.
It can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia. How the platforms compare →
| PLATFORM | AVAILABILITY | DETAIL |
|---|---|---|
| AWS Bedrock · Sydney | NOT OFFERED | |
| AWS Bedrock · Melbourne | NOT OFFERED | |
| Azure · Australia East | GLOBAL | |
| Google Vertex AI · Sydney | NOT OFFERED | |
| Your own infrastructure | SELF-HOST | Open weights: run it on your own servers or GPU instances in an Australian region. |
Compare it with.
DeepSeek V4 Flash, answered.
How much does DeepSeek V4 Flash cost in Australian dollars?
DeepSeek V4 Flash lists at US$0.44 per million input tokens and US$1.32 per million output tokens — A$0.62 and A$1.85 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$1.79 per 1,000 replies, before any reasoning tokens.
Can I use DeepSeek V4 Flash in Australia with data kept onshore?
DeepSeek V4 Flash can be called from Microsoft Azure in Australia East, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.
Is DeepSeek V4 Flash good for coding?
It ranks 24th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 69.1 and an LMArena WebDev rating of 1580 (24th of 129). The current leader is Claude Fable 5.1.
How good is DeepSeek V4 Flash at agent and automation work?
It ranks 22nd of 41 on our agents ranking, completing 39.4% of τ²-Bench banking customer-service tasks and 78.7% of Terminal-Bench 2.1 tasks.
How fast is DeepSeek V4 Flash?
Artificial Analysis measures DeepSeek V4 Flash at about 237 tok/s of output, with the answer starting after 9.2 s on average once thinking time is included (reasoning, max effort setting).
What is DeepSeek V4 Flash’s context window?
DeepSeek documents a context window of 1,000,000 tokens (1M), with up to 384,000 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 79.7%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →