Gemini 3.5 Flash-Lite benchmarks

Gemini 3.5 Flash-Lite is a proprietary model from Google, released 21 Jul 2026. It ranks 47th of 51 on our overall leaderboard and 42nd of 50 for value. At A$1.19 per million tokens (blended), it costs about 3.0× cheaper than the median model we track. It writes about 350 tokens a second, and answers start after 6.7 s on average, thinking included. It can be called from Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia.

#47 OF 51 OVERALLA$1.19 / 1M TOKENS1M CONTEXTUPDATED 23 SEP 2026
CONTEXT WINDOW
1,048,576 tokens
MAX OUTPUT
65,536 tokens
INPUT
text, image, video, audio
WEIGHTS
Proprietary
RELEASED
21 Jul 2026
OUTPUT SPEED
350 tok/s
ANSWER STARTS AFTER
6.7 s (thinking included)
GOOGLE DOCS
§ 02 — RESULTS

Benchmark results.

Independent test scores are for the default setting. “Among tracked” ranks Gemini 3.5 Flash-Lite against the 51 models on this site; LMArena ranks run across its full leaderboard.

ALL RESULTS8 OF 8 CORE SIGNALS
Gemini 3.5 Flash-Lite benchmark results
BENCHMARKRESULTAMONG TRACKEDSOURCE DETAIL
INDEPENDENT TESTS · ARTIFICIAL ANALYSIS
AA Intelligence Index22.2#49 of 51Default
AA Coding Index49.3#40 of 42Default
GPQA Diamond83.8%#43 of 44Default
Humanity’s Last Exam18.8%#50 of 51Default
AA-LCR76.0%#47 of 51Default
Terminal-Bench 2.153.6%#40 of 42Default
τ²-Bench (banking)17.5%#36 of 41Default
SciCode41.3%#45 of 46Default
BLIND HUMAN VOTES · LMARENA
LMArena Text26,165 votes · ±51456#34 of 42#59 of 402 on LMArena · 13 Sep 2026
LMArena WebDev228 votes · ±431447#40 of 45#58 of 129 on LMArena · 22 Sep 2026
LMArena Vision5,254 votes · ±101266#25 of 33#34 of 152 on LMArena · 13 Sep 2026
LMArena Agent23,248 sessions#45#37 of 37#45 of 46 on LMArena · 15 Sep 2026
§ 03 — COST

What Gemini 3.5 Flash-Lite costs in Australian dollars.

List prices converted at the snapshot’s RBA rate, excluding GST. Reasoning models are billed for their thinking as output tokens, so heavier settings cost more than the job estimates show.

LIST PRICE
Gemini 3.5 Flash-Lite price per million tokens
PER MILLION TOKENSAUDUSD LIST
Input tokensA$0.42US$0.30
Output tokensA$3.51US$2.50
Blended (3 in : 1 out)A$1.19US$0.85
EVERYDAY JOBSESTIMATES
Estimated cost of common jobs in Australian dollars
JOBCOST (AUD)MEDIAN MODEL
Customer support replyPER 1,000 REPLIESA$1.90A$6.58
Summarise a 30-page documentPER 100 DOCUMENTSA$1.12A$4.74
Agentic coding taskPER 10 TASKSA$0.91A$3.72
§ 04 — AUSTRALIA

Running Gemini 3.5 Flash-Lite in Australia.

It can be called from Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia. How the platforms compare →

AUSTRALIAN CLOUD REGIONS
Gemini 3.5 Flash-Lite availability in Australian cloud regions
PLATFORMAVAILABILITYDETAIL
AWS Bedrock · SydneyNOT OFFERED
AWS Bedrock · MelbourneNOT OFFERED
Azure · Australia EastNOT OFFERED
Google Vertex AI · SydneyGLOBAL

Provider docs, checked 23 Sep 2026

§ 06 — QUESTIONS

Gemini 3.5 Flash-Lite, answered.

Gemini 3.5 Flash-Lite lists at US$0.30 per million input tokens and US$2.50 per million output tokens — A$0.42 and A$3.51 at A$1 = US$0.7123 (RBA, 22 Sep 2026), excluding GST. A typical customer support reply works out at about A$1.90 per 1,000 replies, before any reasoning tokens.

Gemini 3.5 Flash-Lite can be called from Google Vertex AI in Sydney, but only through global routing, so requests may be processed outside Australia. Availability comes from the cloud providers’ own documentation; check your provider’s current terms before relying on it for data-residency obligations.

It ranks 40th of 42 on our coding ranking, with an Artificial Analysis Coding Index of 49.3 and an LMArena WebDev rating of 1447 (58th of 129). The current leader is Claude Fable 5.1.

It ranks 40th of 41 on our agents ranking, completing 17.5% of τ²-Bench banking customer-service tasks and 53.6% of Terminal-Bench 2.1 tasks.

Artificial Analysis measures Gemini 3.5 Flash-Lite at about 350 tok/s of output, with the answer starting after 6.7 s on average once thinking time is included (default setting).

Google documents a context window of 1,048,576 tokens (1M), with up to 65,536 output tokens. On AA-LCR, which tests reasoning across ~100,000-token document sets, it scores 76.0%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?