Best AI models for long documents.

Which models actually use a long context window: reasoning across large document sets, plus blind votes on questions about uploaded files. As of 23 Sep 2026, Kimi K3 leads for long context, ahead of Step 5 and MiMo V2.6 Pro, out of 51 ranked models. Kimi K3 has open weights, so it can also be self-hosted. The cheapest model in the top ten is DeepSeek V4.1 Flash, at A$0.74 per million tokens.

51 MODELS RANKEDA$1 = US$0.7123UPDATED 23 SEP 2026
LONG CONTEXT RANKING51 MODELS · 23 SEP 2026
51 MODELS
Best AI models for long documents, ranked, as of 23 Sep 2026
MODELIN AUSTRALIA
01Kimi K3MOONSHOT AIPARTIAL100.088.7%A$8.42GLOBAL
02Step 5STEPFUNPARTIAL97.988.3%A$2.00
03MiMo V2.6 ProXIAOMIPARTIAL87.686.3%A$0.76NOT IN AU
04Claude Fable 5.1ANTHROPIC85.085.3%A$28.08GLOBAL
05Claude Opus 5.5ANTHROPICPARTIAL79.484.7%A$11.23IN AU
06GPT-5.6 SolOPENAI78.484.0%A$11.23GLOBAL
07Claude Fable 5ANTHROPIC76.982.3%A$28.08GLOBAL
08GPT-5.5OPENAI76.884.3%A$15.79GLOBAL
09DeepSeek V4.1 FlashDEEPSEEKPARTIAL75.884.0%A$0.74SELF-HOST
10GPT-6 SolOPENAIPARTIAL74.283.7%A$5.62GLOBAL
11GPT-6 LunaOPENAIPARTIAL72.283.3%A$0.28GLOBAL
12GPT-5.6 TerraOPENAI69.383.0%A$6.32GLOBAL
13Qwen3.8 27BALIBABAPARTIAL65.582.0%A$1.58SELF-HOST
14Claude Sonnet 5ANTHROPIC65.382.0%A$5.62IN AU
15Muse Spark 1.3META65.383.0%A$2.81NOT IN AU
16GPT-5.6 LunaOPENAI65.283.7%A$0.63GLOBAL
17Gemini 3.7 FlashGOOGLEPARTIAL63.981.7%A$2.11GLOBAL
18GPT-5.4OPENAI62.782.0%A$7.90IN AU
19Claude Opus 5ANTHROPIC61.979.3%A$14.04IN AU
20Gemini 3.8 FlashGOOGLEPARTIAL61.981.3%A$2.11GLOBAL
21Claude Opus 4.7ANTHROPIC61.878.7%A$14.04IN AU
22Claude Opus 4.6ANTHROPIC59.878.0%A$14.04IN AU
23GPT-6 AstraOPENAI59.280.7%A$28.08GLOBAL
24Claude Sonnet 4.6ANTHROPIC59.080.0%A$8.42IN AU
25Gemini 3.1 ProGOOGLE57.682.0%A$6.32GLOBAL
26DeepSeek V4 ProDEEPSEEKPARTIAL56.780.3%A$2.78GLOBAL
27Qwen3.8 MaxALIBABAPARTIAL56.780.3%A$4.21NOT IN AU
28GLM-5.3 FlashZ.AIPARTIAL55.280.0%A$0.33SELF-HOST
29Claude Opus 4.8ANTHROPIC54.477.7%A$14.04IN AU
30DeepSeek V4 FlashDEEPSEEKPARTIAL53.679.7%A$0.93GLOBAL
31GLM-5.3Z.AIPARTIAL53.679.7%A$3.02SELF-HOST
32MiMo V2.5 ProXIAOMIPARTIAL53.679.7%A$0.76SELF-HOST
33Qwen3.8 Flash NextALIBABAPARTIAL53.679.7%A$0.32SELF-HOST
34Grok 4.6XAI52.580.3%A$4.21GLOBAL
35Muse Spark 1.1META51.577.7%A$2.81
36Grok 4.5XAI50.379.3%A$4.21NOT IN AU
37Hy3TENCENTPARTIAL50.079.0%A$0.35SELF-HOST
38Muse Spark 1.2METAPARTIAL50.079.0%A$2.81NOT IN AU
39Qwen3.7 MaxALIBABAPARTIAL50.079.0%A$5.26
40Kimi K2.6MOONSHOT AI49.981.0%A$2.40SELF-HOST
41MiniMax M3MINIMAX49.483.0%A$0.74SELF-HOST
42Gemini 3.6 FlashGOOGLE47.980.0%A$2.11GLOBAL
43Muse SparkMETA47.078.0%
44GLM-5.2Z.AIPARTIAL46.478.3%A$3.02SELF-HOST
45GPT-5.4 miniOPENAIPARTIAL39.777.0%A$2.37
46Grok 4.7XAIPARTIAL38.176.7%A$4.21NOT IN AU
47Gemini 3.5 Flash-LiteGOOGLEPARTIAL34.576.0%A$1.19GLOBAL
48Gemini 3 ProGOOGLE32.476.0%A$6.32
49Gemini 3.5 FlashGOOGLE27.073.3%A$4.74IN AU
50Gemma 4 31BGOOGLE6.769.7%A$0.29SELF-HOST
51Mistral Medium 3.5MISTRAL AIPARTIAL0.069.3%A$4.21SELF-HOST
HOW THIS RANKING WORKS

What goes into the long context score.

A big context window only helps if the model can find and connect what matters inside it. AA-LCR tests exactly that across roughly 100,000-token document sets; the document arena adds how people rate answers about their own files. Models with a documented window under 128K tokens are left out.

70%AA-LCR

Artificial Analysis’s long-context reasoning test: questions that need information pulled together from sets of documents of around 100,000 tokens.

30%LMArena Document

Blind votes on answers to questions about uploaded documents (style-controlled).

Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →

QUESTIONS

Choosing a model for long context.

As of 23 Sep 2026, Kimi K3 ranks first, followed by Step 5 and MiMo V2.6 Pro. The ranking combines AA-LCR and LMArena Document and is refreshed weekly.

Kimi K3 from Moonshot AI is the highest-ranked open-weights model (1st of 51). Open weights can be run on your own infrastructure, including in an Australian region.

Of the top ten, DeepSeek V4.1 Flash is cheapest at A$0.74 per million tokens (blended), against A$28.08 for Claude Fable 5.1. Prices are list prices converted at A$1 = US$0.7123, excluding GST.

Of the top ten, Claude Opus 5.5 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.

A big context window only helps if the model can find and connect what matters inside it. AA-LCR tests exactly that across roughly 100,000-token document sets; the document arena adds how people rate answers about their own files. Models with a documented window under 128K tokens are left out. Each input is rescaled 0–100 across the models we track, then weighted: AA-LCR 70% and LMArena Document 30%.

PUT THE COMPARISON TO WORK

Need help choosing and using AI for your business?