Kimi K3 benchmarks
Kimi K3 is an open-weights model from Moonshot AI, released 16 Jul 2026. It ranks 15th of 51 on our overall leaderboard and 1st of 51 for long context. At A$8.42 per million tokens (blended), it costs 2.3× the median model we track. It writes about 38 tokens a second, and answers start after 56 s on average, thinking included. It can be called from AWS Bedrock in Sydney and Melbourne, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia.
- CONTEXT WINDOW
- 1,048,576 tokens
- INPUT
- text, image, video
- WEIGHTS
- Open (Kimi K3 license)
- RELEASED
- 16 Jul 2026
- OUTPUT SPEED
- 38 tok/s
- ANSWER STARTS AFTER
- 56 s (thinking included)
Kimi K3 across our rankings.
Benchmark results.
Independent test scores are for the max effort setting. “Among tracked” ranks Kimi K3 against the 51 models on this site; LMArena ranks run across its full leaderboard.
| BENCHMARK | RESULT | AMONG TRACKED | SOURCE DETAIL |
|---|---|---|---|
| INDEPENDENT TESTS · ARTIFICIAL ANALYSIS | |||
| AA Intelligence Index | 43.6 | #15 of 51 | max effort |
| AA Coding Index | 76.2 | #9 of 42 | max effort |
| GPQA Diamond | 93.5% | #8 of 44 | max effort |
| Humanity’s Last Exam | 46.9% | #14 of 51 | max effort |
| AA-LCR | 88.7% | #1 of 51 | max effort |
| Terminal-Bench 2.1 | 85.0% | #11 of 42 | max effort |
| τ²-Bench (banking) | 46.0% | #8 of 41 | max effort |
| SciCode | 59.5% | #5 of 46 | max effort |
| BLIND HUMAN VOTES · LMARENA | |||
| LMArena Text20,987 votes · ±5 | 1485 | #14 of 42 | #17 of 402 on LMArena · 13 Sep 2026 |
| LMArena WebDev13,140 votes · ±7 | 1658 | #5 of 45 | #7 of 129 on LMArena · 22 Sep 2026 |
| LMArena Agent108,615 sessions | #8 | #7 of 37 | #8 of 46 on LMArena · 15 Sep 2026 |
What Kimi K3 costs in Australian dollars.
List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.
| PER MILLION TOKENS | AUD | USD LIST |
|---|---|---|
| Input tokens | A$4.21 | US$3.00 |
| Output tokens | A$21.06 | US$15.00 |
| Blended (3 in : 1 out) | A$8.42 | US$6.00 |
| JOB | COST (AUD) | MEDIAN MODEL |
|---|---|---|
| Customer support replyPER 1,000 REPLIES | A$14.74 | ≈ A$6.58 |
| Summarise a 30-page documentPER 100 DOCUMENTS | A$10.11 | ≈ A$4.74 |
| Agentic coding taskPER 10 TASKS | A$8.00 | ≈ A$3.72 |
| SETTING | INTELLIGENCE | A$ / 1M | SPEED |
|---|---|---|---|
| max effort | 43.6 | A$8.42 | 38 tok/s |
| low effort | 34.5 | A$8.42 | 37 tok/s |
Running Kimi K3 in Australia.
It can be called from AWS Bedrock in Sydney and Melbourne, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia. How the platforms compare →
| PLATFORM | AVAILABILITY | DETAIL |
|---|---|---|
| AWS Bedrock · Sydney | GLOBAL | |
| AWS Bedrock · Melbourne | GLOBAL | |
| Azure · Australia East | NOT OFFERED | |
| Google Vertex AI · Sydney | NOT OFFERED | |
| Your own infrastructure | SELF-HOST | Open weights: run it on your own servers or GPU instances in an Australian region. |
Compare it with.
Kimi K3, answered.
How much does Kimi K3 cost in Australian dollars?
Kimi K3 lists at US$3.00 per million input tokens and US$15.00 per million output tokens — A$4.21 and A$21.06 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$14.74 per 1,000 replies, before any reasoning tokens.
Can I use Kimi K3 in Australia with data kept onshore?
Kimi K3 can be called from AWS Bedrock in Sydney and Melbourne, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.
Is Kimi K3 good for coding?
It ranks 5th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 76.2 and an LMArena WebDev rating of 1658 (7th of 129). The current leader is Claude Fable 5.1.
How good is Kimi K3 at agent and automation work?
It ranks 9th of 41 on our agents ranking, completing 46.0% of τ²-Bench banking customer-service tasks and 85.0% of Terminal-Bench 2.1 tasks.
How fast is Kimi K3?
Artificial Analysis measures Kimi K3 at about 38 tok/s of output, with the answer starting after 56 s on average once thinking time is included (max effort setting).
What is Kimi K3’s context window?
Moonshot AI documents a context window of 1,048,576 tokens (1M). On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 88.7%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →