Gemma 4 31B benchmarks

Gemma 4 31B is an open-weights model from Google, released 2 Apr 2026. It ranks 50th of 51 on our overall leaderboard and 35th of 50 for value. At A$0.29 per million tokens (blended), it costs about 13× cheaper than the median model we track. It writes about 35 tokens a second, and answers start after 50 s on average, thinking included. Its weights are public, so it can be hosted in Australia on your own infrastructure or an Australian cloud region.

#50 OF 51 OVERALLA$0.29 / 1M TOKENS250K CONTEXTUPDATED 23 SEP 2026
CONTEXT WINDOW
256,000 tokens
INPUT
text, image
WEIGHTS
Open (Apache 2.0)
RELEASED
2 Apr 2026
OUTPUT SPEED
35 tok/s
ANSWER STARTS AFTER
50 s (thinking included)
GOOGLE DOCS
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the reasoning setting. “Among tracked” ranks Gemma 4 31B against the 51 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS8 OF 8 CORE SIGNALS
Gemma 4 31B benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index19.0#50 of 51Reasoning
AA Coding Index43.4#42 of 42Reasoning
GPQA Diamond85.7%#42 of 44Reasoning
Humanity’s Last Exam23.6%#49 of 51Reasoning
AA-LCR69.7%#50 of 51Reasoning
Terminal-Bench 2.143.4%#42 of 42Reasoning
τ²-Bench (banking)14.8%#39 of 41Reasoning
SciCode45.5%#44 of 46Reasoning
BLIND HUMAN VOTES · LMARENA
LMArena Text5,894 votes · ±81451#38 of 42#68 of 402 on LMArena · 13 Sep 2026
LMArena WebDev11,584 votes · ±71364#44 of 45#91 of 129 on LMArena · 22 Sep 2026
LMArena Vision35,713 votes · ±61261#28 of 33#38 of 152 on LMArena · 13 Sep 2026
LMArena Document12,326 votes · ±81445#25 of 26#34 of 44 on LMArena · 13 Sep 2026
§ 03 — COST

What Gemma 4 31B costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
Gemma 4 31B price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$0.20US$0.14
Output tokensA$0.56US$0.40
Blended (3 in : 1 out)A$0.29US$0.20
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$0.56A$6.58
Summarise a 30-page documentPER 100 DOCUMENTSA$0.44A$4.74
Agentic coding taskPER 10 TASKSA$0.34A$3.72
SETTINGSMEASURED SEPARATELY
Gemma 4 31B settings compared
SETTINGINTELLIGENCEA$ / 1MSPEED
Reasoning19.035 tok/s
Non-reasoning13.9A$0.2938 tok/s
§ 04 — AUSTRALIA

Running Gemma 4 31B in Australia.

Its weights are public, so it can be hosted in Australia on your own infrastructure or an Australian cloud region. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
Gemma 4 31B availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyNOT IN AU

Provider docs, checked 23 Sep 2026

AWS Bedrock · MelbourneNOT IN AU

Provider docs, checked 23 Sep 2026

Azure · Australia EastNOT OFFERED
Google Vertex AI · SydneyNOT OFFERED
Your own infrastructureSELF-HOST

Open weights: run it on your own servers or GPU instances in an Australian region.

§ 06 — QUESTIONS

Gemma 4 31B, answered.

Gemma 4 31B lists at US$0.14 per million input tokens and US$0.40 per million output tokens — A$0.20 and A$0.56 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$0.56 per 1,000 replies, before any reasoning tokens.

Gemma 4 31B’s weights are public, so it can be hosted in Australia on your own infrastructure or an Australian cloud region. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

It ranks 41st of 42 on our coding ranking, with an Artificial Analysis Coding Index of 43.4 and an LMArena WebDev rating of 1364 (91st of 129). The current leader is Claude Fable 5.1.

It ranks 41st of 41 on our agents ranking, completing 14.8% of τ²-Bench banking customer-service tasks and 43.4% of Terminal-Bench 2.1 tasks.

Artificial Analysis measures Gemma 4 31B at about 35 tok/s of output, with the answer starting after 50 s on average once thinking time is included (reasoning setting).

Google documents a context window of 256,000 tokens (250K). On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 69.7%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?