Best AI models for long documents.
Which models actually use a long context window: reasoning across large document sets, plus blind votes on questions about uploaded files. As of 23 Sep 2026, Kimi K3 leads for long context, ahead of Step 5 and MiMo V2.6 Pro, out of 51 ranked models. Kimi K3 has open weights, so it can also be self-hosted. The cheapest model in the top ten is DeepSeek V4.1 Flash, at A$0.74 per million tokens.
| MODEL | IN AUSTRALIA | |||
|---|---|---|---|---|
| 01Kimi K3MOONSHOT AIPARTIAL | 100.0 | 88.7% | A$8.42 | GLOBAL |
| 02Step 5STEPFUNPARTIAL | 97.9 | 88.3% | A$2.00 | |
| 03MiMo V2.6 ProXIAOMIPARTIAL | 87.6 | 86.3% | A$0.76 | NOT IN AU |
| 04Claude Fable 5.1ANTHROPIC | 85.0 | 85.3% | A$28.08 | GLOBAL |
| 05Claude Opus 5.5ANTHROPICPARTIAL | 79.4 | 84.7% | A$11.23 | IN AU |
| 06GPT-5.6 SolOPENAI | 78.4 | 84.0% | A$11.23 | GLOBAL |
| 07Claude Fable 5ANTHROPIC | 76.9 | 82.3% | A$28.08 | GLOBAL |
| 08GPT-5.5OPENAI | 76.8 | 84.3% | A$15.79 | GLOBAL |
| 09DeepSeek V4.1 FlashDEEPSEEKPARTIAL | 75.8 | 84.0% | A$0.74 | SELF-HOST |
| 10GPT-6 SolOPENAIPARTIAL | 74.2 | 83.7% | A$5.62 | GLOBAL |
| 11GPT-6 LunaOPENAIPARTIAL | 72.2 | 83.3% | A$0.28 | GLOBAL |
| 12GPT-5.6 TerraOPENAI | 69.3 | 83.0% | A$6.32 | GLOBAL |
| 13Qwen3.8 27BALIBABAPARTIAL | 65.5 | 82.0% | A$1.58 | SELF-HOST |
| 14Claude Sonnet 5ANTHROPIC | 65.3 | 82.0% | A$5.62 | IN AU |
| 15Muse Spark 1.3META | 65.3 | 83.0% | A$2.81 | NOT IN AU |
| 16GPT-5.6 LunaOPENAI | 65.2 | 83.7% | A$0.63 | GLOBAL |
| 17Gemini 3.7 FlashGOOGLEPARTIAL | 63.9 | 81.7% | A$2.11 | GLOBAL |
| 18GPT-5.4OPENAI | 62.7 | 82.0% | A$7.90 | IN AU |
| 19Claude Opus 5ANTHROPIC | 61.9 | 79.3% | A$14.04 | IN AU |
| 20Gemini 3.8 FlashGOOGLEPARTIAL | 61.9 | 81.3% | A$2.11 | GLOBAL |
| 21Claude Opus 4.7ANTHROPIC | 61.8 | 78.7% | A$14.04 | IN AU |
| 22Claude Opus 4.6ANTHROPIC | 59.8 | 78.0% | A$14.04 | IN AU |
| 23GPT-6 AstraOPENAI | 59.2 | 80.7% | A$28.08 | GLOBAL |
| 24Claude Sonnet 4.6ANTHROPIC | 59.0 | 80.0% | A$8.42 | IN AU |
| 25Gemini 3.1 ProGOOGLE | 57.6 | 82.0% | A$6.32 | GLOBAL |
| 26DeepSeek V4 ProDEEPSEEKPARTIAL | 56.7 | 80.3% | A$2.78 | GLOBAL |
| 27Qwen3.8 MaxALIBABAPARTIAL | 56.7 | 80.3% | A$4.21 | NOT IN AU |
| 28GLM-5.3 FlashZ.AIPARTIAL | 55.2 | 80.0% | A$0.33 | SELF-HOST |
| 29Claude Opus 4.8ANTHROPIC | 54.4 | 77.7% | A$14.04 | IN AU |
| 30DeepSeek V4 FlashDEEPSEEKPARTIAL | 53.6 | 79.7% | A$0.93 | GLOBAL |
| 31GLM-5.3Z.AIPARTIAL | 53.6 | 79.7% | A$3.02 | SELF-HOST |
| 32MiMo V2.5 ProXIAOMIPARTIAL | 53.6 | 79.7% | A$0.76 | SELF-HOST |
| 33Qwen3.8 Flash NextALIBABAPARTIAL | 53.6 | 79.7% | A$0.32 | SELF-HOST |
| 34Grok 4.6XAI | 52.5 | 80.3% | A$4.21 | GLOBAL |
| 35Muse Spark 1.1META | 51.5 | 77.7% | A$2.81 | |
| 36Grok 4.5XAI | 50.3 | 79.3% | A$4.21 | NOT IN AU |
| 37Hy3TENCENTPARTIAL | 50.0 | 79.0% | A$0.35 | SELF-HOST |
| 38Muse Spark 1.2METAPARTIAL | 50.0 | 79.0% | A$2.81 | NOT IN AU |
| 39Qwen3.7 MaxALIBABAPARTIAL | 50.0 | 79.0% | A$5.26 | |
| 40Kimi K2.6MOONSHOT AI | 49.9 | 81.0% | A$2.40 | SELF-HOST |
| 41MiniMax M3MINIMAX | 49.4 | 83.0% | A$0.74 | SELF-HOST |
| 42Gemini 3.6 FlashGOOGLE | 47.9 | 80.0% | A$2.11 | GLOBAL |
| 43Muse SparkMETA | 47.0 | 78.0% | ||
| 44GLM-5.2Z.AIPARTIAL | 46.4 | 78.3% | A$3.02 | SELF-HOST |
| 45GPT-5.4 miniOPENAIPARTIAL | 39.7 | 77.0% | A$2.37 | |
| 46Grok 4.7XAIPARTIAL | 38.1 | 76.7% | A$4.21 | NOT IN AU |
| 47Gemini 3.5 Flash-LiteGOOGLEPARTIAL | 34.5 | 76.0% | A$1.19 | GLOBAL |
| 48Gemini 3 ProGOOGLE | 32.4 | 76.0% | A$6.32 | |
| 49Gemini 3.5 FlashGOOGLE | 27.0 | 73.3% | A$4.74 | IN AU |
| 50Gemma 4 31BGOOGLE | 6.7 | 69.7% | A$0.29 | SELF-HOST |
| 51Mistral Medium 3.5MISTRAL AIPARTIAL | 0.0 | 69.3% | A$4.21 | SELF-HOST |
What goes into the long context score.
A big context window only helps if the model can find and connect what matters inside it. AA-LCR tests exactly that across roughly 100,000-token document sets; the document arena adds how people rate answers about their own files. Models with a documented window under 128K tokens are left out.
Artificial Analysis’s long-context reasoning test: questions that need information pulled together from sets of documents of around 100,000 tokens.
Blind votes on answers to questions about uploaded documents (style-controlled).
Scores are rescaled 0–100 across the models we track, so they compare models with each other rather than against a fixed bar. Full methodology →
Choosing a model for long context.
What is the best AI model for long context right now?
As of 23 Sep 2026, Kimi K3 ranks first, followed by Step 5 and MiMo V2.6 Pro. The ranking combines AA-LCR and LMArena Document and is refreshed weekly.
What is the best open-weights model for long context?
Kimi K3 from Moonshot AI is the highest-ranked open-weights model (1st of 51). Open weights can be run on your own infrastructure, including in an Australian region.
Which top long context model is cheapest?
Of the top ten, DeepSeek V4.1 Flash is cheapest at A$0.74 per million tokens (blended), against A$28.08 for Claude Fable 5.1. Prices are list prices converted at A$1 = US$0.7123, excluding GST.
Which of these models can keep data in Australia?
Of the top ten, Claude Opus 5.5 can run with requests kept in Australia on at least one major cloud platform. See the Australia page for each platform.
How is this ranking calculated?
A big context window only helps if the model can find and connect what matters inside it. AA-LCR tests exactly that across roughly 100,000-token document sets; the document arena adds how people rate answers about their own files. Models with a documented window under 128K tokens are left out. Each input is rescaled 0–100 across the models we track, then weighted: AA-LCR 70% and LMArena Document 30%.
Ratings: LMArena leaderboard dataset (CC BY 4.0), rescaled for the lens scores · leaderboards to 22 Sep 2026
Evaluations, prices and speed: Artificial Analysis (artificialanalysis.ai) · fetched 23 Sep 2026
Exchange rate: Reserve Bank of Australia, table F11.1 · A$1 = US$0.7123 on 22 Sep 2026 · prices exclude GST
Snapshot 23 Sep 2026 · updated weekly · How the rankings work →