Kimi K3 benchmarks

Kimi K3 is an open-weights model from Moonshot AI, released 16 Jul 2026. It ranks 15th of 51 on our overall leaderboard and 1st of 51 for long context. At A$8.42 per million tokens (blended), it costs 2.3× the median model we track. It writes about 38 tokens a second, and answers start after 56 s on average, thinking included. It can be called from AWS Bedrock in Sydney and Melbourne, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia.

#15 OF 51 OVERALLA$8.42 / 1M TOKENS1M CONTEXTUPDATED 23 SEP 2026
CONTEXT WINDOW
1,048,576 tokens
INPUT
text, image, video
WEIGHTS
Open (Kimi K3 license)
RELEASED
16 Jul 2026
OUTPUT SPEED
38 tok/s
ANSWER STARTS AFTER
56 s (thinking included)
MOONSHOT AI DOCS
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the max effort setting. “Among tracked” ranks Kimi K3 against the 51 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS8 OF 8 CORE SIGNALS
Kimi K3 benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index43.6#15 of 51max effort
AA Coding Index76.2#9 of 42max effort
GPQA Diamond93.5%#8 of 44max effort
Humanity’s Last Exam46.9%#14 of 51max effort
AA-LCR88.7%#1 of 51max effort
Terminal-Bench 2.185.0%#11 of 42max effort
τ²-Bench (banking)46.0%#8 of 41max effort
SciCode59.5%#5 of 46max effort
BLIND HUMAN VOTES · LMARENA
LMArena Text20,987 votes · ±51485#14 of 42#17 of 402 on LMArena · 13 Sep 2026
LMArena WebDev13,140 votes · ±71658#5 of 45#7 of 129 on LMArena · 22 Sep 2026
LMArena Agent108,615 sessions#8#7 of 37#8 of 46 on LMArena · 15 Sep 2026
§ 03 — COST

What Kimi K3 costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
Kimi K3 price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$4.21US$3.00
Output tokensA$21.06US$15.00
Blended (3 in : 1 out)A$8.42US$6.00
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$14.74A$6.58
Summarise a 30-page documentPER 100 DOCUMENTSA$10.11A$4.74
Agentic coding taskPER 10 TASKSA$8.00A$3.72
SETTINGSMEASURED SEPARATELY
Kimi K3 settings compared
SETTINGINTELLIGENCEA$ / 1MSPEED
max effort43.6A$8.4238 tok/s
low effort34.5A$8.4237 tok/s
§ 04 — AUSTRALIA

Running Kimi K3 in Australia.

It can be called from AWS Bedrock in Sydney and Melbourne, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
Kimi K3 availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyGLOBAL

Provider docs, checked 23 Sep 2026

AWS Bedrock · MelbourneGLOBAL

Provider docs, checked 23 Sep 2026

Azure · Australia EastNOT OFFERED
Google Vertex AI · SydneyNOT OFFERED
Your own infrastructureSELF-HOST

Open weights: run it on your own servers or GPU instances in an Australian region.

§ 06 — QUESTIONS

Kimi K3, answered.

Kimi K3 lists at US$3.00 per million input tokens and US$15.00 per million output tokens — A$4.21 and A$21.06 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$14.74 per 1,000 replies, before any reasoning tokens.

Kimi K3 can be called from AWS Bedrock in Sydney and Melbourne, but only through global routing, so requests may be processed outside Australia; its weights are public, so it can also be self-hosted in Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

It ranks 5th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 76.2 and an LMArena WebDev rating of 1658 (7th of 129). The current leader is Claude Fable 5.1.

It ranks 9th of 41 on our agents ranking, completing 46.0% of τ²-Bench banking customer-service tasks and 85.0% of Terminal-Bench 2.1 tasks.

Artificial Analysis measures Kimi K3 at about 38 tok/s of output, with the answer starting after 56 s on average once thinking time is included (max effort setting).

Moonshot AI documents a context window of 1,048,576 tokens (1M). On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 88.7%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?