Z.ai GLM release history.

Z.ai has shipped 16 major model releases and 3 other major launches since ChatGLM arrived in March 2023. The latest model recorded here is GLM-5.3-Flash, announced 26 Aug 2026. Every entry is dated from Z.ai’s own announcement and links to it.

19 MAJOR RELEASES2023–2026CHECKED 23 SEP 2026
NEWEST AIR & FLASH
GLM-5.3-Flash · 26 Aug 2026
NEWEST GLM FLAGSHIP
GLM-5.3 · 14 Aug 2026
NEWEST VISION (V)
GLM-5V-Turbo · 2 Apr 2026
§ 02 — MODELS

Every GLM model release.

16 model releases, newest first. “Current” is the newest release in its line; older ones are superseded, and retired ones have left Z.ai’s apps or API. Where a model is in our benchmark tables, the last column links to its scores, price in AUD and Australian availability.

MODEL RELEASES16 RELEASES
GLM model releases, newest first
RELEASEANNOUNCEDLINESTATUSBENCHMARKS
GLM-5.3-FlashAir & FlashCURRENTGLM-5.3 Flash
GLM-5.3GLM flagshipCURRENTGLM-5.3
GLM-5.2GLM flagshipSUPERSEDED
GLM-5.1GLM flagshipSUPERSEDED
GLM-5V-TurboVision (V)CURRENT
GLM-5GLM flagshipSUPERSEDED
GLM-4.7GLM flagshipSUPERSEDED
GLM-4.6VVision (V)SUPERSEDED
GLM-4.6GLM flagshipSUPERSEDED
GLM-4.5VVision (V)SUPERSEDED
GLM-4.5 and GLM-4.5-AirWATERSHEDGLM flagship · Air & FlashSUPERSEDED
GLM-4-0414, GLM-Z1 and Z.aiZ1 reasoning · GLM flagshipSUPERSEDED
GLM-4-9B open modelsGLM flagshipSUPERSEDED
ChatGLM3-6BChatGLMSUPERSEDED
ChatGLM2-6BChatGLMSUPERSEDED
ChatGLM and ChatGLM-6BWATERSHEDChatGLMSUPERSEDED
§ 03 — TIMELINE

Every major release, year by year.

19 OF 19

20267 RELEASES

  1. MODEL

    GLM-5.3-Flash

    • OPEN WEIGHTS
    • CURRENT
    • GLM-5 LONG-HORIZON

    GLM-5.3-Flash is a newly trained 320B-parameter model (18B active) and the first natively multimodal model in the GLM-5 series, using a hybrid of sparse and linear attention for cheaper 1M-token contexts. Z.ai says it beats GLM-5.2 at one-tenth the price and approaches Claude Opus 4.8 on coding and agent benchmarks; weights are MIT-licensed.

    Why it mattered. It offers close to flagship coding at a fraction of the cost, with three times GLM-5.2’s quota on the Coding Plan, and Z.ai serves it on a large cluster of Chinese AI chips.

    AVAILABLE TO
    Open weights · API · Coding Plan
  2. MODEL

    GLM-5.3

    • LAUNCHED
    • CURRENT
    • GLM-5 LONG-HORIZON

    GLM-5.3 keeps GLM-5.2’s base model and improves it through post-training alone: Z.ai reports a 50% gain on its Code Bench, Terminal-Bench 3.0 up from 4.6 to 28.3, and 84.5% on CyberGym. It launched on the Coding Plan with thinking always on at low, high or max effort; weights followed on Hugging Face under a new GLM-5.3 licence.

    Why it mattered. Z.ai also pitched it for security work: with partner teams it found 2,436 vulnerabilities in 269 real-world projects. Its licence requires a Z.ai security review for model-as-a-service providers earning over US$10 billion a year.

    AVAILABLE TO
    Coding Plan · API · Open weights
    BENCHMARKS
    GLM-5.3 →
  3. MODEL

    GLM-5.2

    • OPEN WEIGHTS
    • SUPERSEDED
    • GLM-5 LONG-HORIZON

    GLM-5.2 brings a 1M-token context built for long-horizon work, IndexShare sparse attention that cuts per-token compute 2.9× at 1M tokens, and multiple thinking-effort levels; it scores 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-bench Pro, and is MIT-licensed.

    Why it mattered. Z.ai reported it as the strongest open model on standard coding benchmarks, within a few points of Claude Opus 4.8 on Terminal-Bench 2.1, with context long enough for multi-hour agent work.

    AVAILABLE TO
    Open weights · API · Coding Plan
  4. PLAN

    GLM Coding Plan Team

    • LAUNCHED
    • GLM-5 LONG-HORIZON

    Z.ai launches a self-serve team edition of the GLM Coding Plan, its subscription for using GLM models in coding tools, with seat, permission, usage and budget management, central billing and invoicing, and code, prompts and conversations excluded from training by default.

    Why it mattered. It extended the individual Coding Plan to companies that need central administration and billing for AI coding.

    AVAILABLE TO
    Team
  5. MODEL

    GLM-5.1

    • OPEN WEIGHTS
    • SUPERSEDED
    • GLM-5 LONG-HORIZON

    GLM-5.1 targets long-horizon agentic engineering: Z.ai says it can work on its own for up to eight hours in a single run, scores a state-of-the-art 58.4 on SWE-Bench Pro, and leads GLM-5 by a wide margin on NL2Repo and Terminal-Bench 2.0; weights are MIT-licensed.

    Why it mattered. It shifted GLM from first-pass answers to sustained work, improving results over hundreds of rounds and thousands of tool calls.

    AVAILABLE TO
    Open weights · API · Coding Plan
  6. MODEL

    GLM-5V-Turbo

    • LAUNCHED
    • CURRENT
    • GLM-5 LONG-HORIZON

    GLM-5V-Turbo is Z.ai’s first multimodal coding foundation model: it takes images, video, text and files, has a 200K context, and is built for vision-based coding and GUI agents working with tools such as Claude Code and OpenClaw.

    Why it mattered. It carried the GLM-5 generation’s coding into tasks where the model has to see the result, such as front-end work and operating interfaces.

    AVAILABLE TO
    API
  7. MODEL

    GLM-5

    • OPEN WEIGHTS
    • SUPERSEDED
    • GLM-5 LONG-HORIZON

    GLM-5 scales to 744B total parameters (40B active) and 28.5T training tokens, adds DeepSeek Sparse Attention to cut deployment cost, and targets complex systems engineering and long-horizon agent tasks. Weights are MIT-licensed and can run on several Chinese chips, including Huawei Ascend.

    Why it mattered. Z.ai benchmarked it against Claude Opus 4.5 on its own coding evaluation; limited compute meant Coding Plan subscribers got access gradually.

    AVAILABLE TO
    Open weights · API · Coding Plan

20256 RELEASES

  1. MODEL

    GLM-4.7

    • OPEN WEIGHTS
    • SUPERSEDED
    • Z.AI AND AGENTIC GLM-4.X

    GLM-4.7 improves agentic coding over GLM-4.6, reaching 73.8% on SWE-bench and 41% on Terminal Bench 2.0, and adds preserved and turn-level thinking so reasoning carries across agent turns; weights are MIT-licensed.

    Why it mattered. Thinking between actions and staying consistent across turns made long sessions in coding agents such as Claude Code, Kilo Code, Cline and Roo Code more stable.

    AVAILABLE TO
    Open weights · API · Coding Plan
  2. MODEL

    GLM-4.6V

    • OPEN WEIGHTS
    • SUPERSEDED
    • Z.AI AND AGENTIC GLM-4.X

    GLM-4.6V (106B) and GLM-4.6V-Flash (9B) are open multimodal models with a 128K context and, for the first time in the series, native function calling, so images, screenshots and document pages can be passed straight into tools.

    Why it mattered. Vision-driven tool use let GLM’s multimodal models act on what they see rather than only describe it.

    AVAILABLE TO
    Open weights · API
  3. MODEL

    GLM-4.6

    • OPEN WEIGHTS
    • SUPERSEDED
    • Z.AI AND AGENTIC GLM-4.X

    GLM-4.6 extends the context window from 128K to 200K tokens, improves coding in agents such as Claude Code, Cline, Roo Code and Kilo Code, and uses about 15% fewer tokens than GLM-4.5 in Z.ai’s real-world coding tests; weights are MIT-licensed.

    Why it mattered. Z.ai reported it on par with Claude Sonnet 4 in those tests but still behind Claude Sonnet 4.5 on coding, and upgraded GLM Coding Plan subscribers to it automatically.

    AVAILABLE TO
    Open weights · API · Coding Plan
  4. MODEL

    GLM-4.5V

    • OPEN WEIGHTS
    • SUPERSEDED
    • Z.AI AND AGENTIC GLM-4.X

    GLM-4.5V is an open vision reasoning model built on GLM-4.5-Air (106B total, 12B active parameters) that handles images, video, documents and GUI agent tasks, with the same thinking-mode switch as GLM-4.5.

    Why it mattered. It brought the GLM-4.5 generation’s reasoning to screenshots, video and interface control in a model developers could run themselves.

    AVAILABLE TO
    Open weights · API
  5. MODEL

    GLM-4.5 and GLM-4.5-Air

    • WATERSHED
    • OPEN WEIGHTS
    • SUPERSEDED
    • Z.AI AND AGENTIC GLM-4.X

    GLM-4.5 (355B total, 32B active parameters) and GLM-4.5-Air (106B, 12B active) are hybrid reasoning models with thinking and non-thinking modes that unify reasoning, coding and agent work, released as MIT-licensed open weights with a 128K context.

    Why it mattered. Z.ai placed it third overall on 12 benchmarks against proprietary and open models and tested it inside Claude Code, making an open model a practical engine for coding agents.

    AVAILABLE TO
    Open weights · API
  6. MODEL

    GLM-4-0414, GLM-Z1 and Z.ai

    • OPEN WEIGHTS
    • SUPERSEDED
    • Z.AI AND AGENTIC GLM-4.X

    Zhipu releases the GLM-4-32B-0414 series under the MIT licence: a 32B chat model, the GLM-Z1 reasoning models (including a 9B version) and GLM-Z1-Rumination, a deep-research model that uses search while it thinks. The models are free to try on the Z.ai chat site.

    Why it mattered. It gave the GLM family open reasoning models under a permissive licence and a free English-language chat site at Z.ai.

    AVAILABLE TO
    Open weights · API · Web

20243 RELEASES

  1. AGENT

    AutoGLM and GLM-PC agents

    • BETA
    • GLM-4 AND FIRST AGENTS

    At its Agent OpenDay, Zhipu opens AutoGLM, first shown in October, to a large-scale beta: the phone agent carries out tasks of more than 50 steps across apps from a short instruction. AutoGLM also comes to the Qingyan browser plug-in, and the GLM-PC computer agent starts invitation testing.

    Why it mattered. AutoGLM moved Zhipu’s assistant from answering to acting inside ordinary apps and websites, a shift Zhipu describes as models going from chat to action.

    AVAILABLE TO
    App · Browser plug-in
  2. PRODUCT

    Video calls in Qingyan

    • BETA
    • GLM-4 AND FIRST AGENTS

    The Qingyan (ChatGLM) app adds video calling: users talk to the assistant while it sees through the camera and can be interrupted mid-answer. It opens to some users from 30 August, announced with the GLM-4-Plus base model and the GLM-4V-Plus video model.

    Why it mattered. Zhipu says it was the first video-call assistant opened to consumers in China, turning Qingyan into an assistant for text, voice, images and live video.

    AVAILABLE TO
    App
  3. MODEL

    GLM-4-9B open models

    • OPEN WEIGHTS
    • SUPERSEDED
    • GLM-4 AND FIRST AGENTS

    Zhipu AI open-sources GLM-4-9B, the open version of the GLM-4 generation it launched in January 2024. The chat model browses the web, runs code and calls tools over a 128K context, with a 1M-context variant and the GLM-4V-9B vision model alongside.

    Why it mattered. It brought GLM-4’s tool use and long context to a model developers could run and fine-tune themselves.

    AVAILABLE TO
    Open weights

20233 RELEASES

  1. MODEL

    ChatGLM3-6B

    • OPEN WEIGHTS
    • SUPERSEDED
    • CHATGLM OPEN MODELS

    Zhipu AI and Tsinghua’s KEG lab release ChatGLM3-6B with a new prompt format that natively supports function calling, a code interpreter and agent tasks, plus base and 32K long-context versions; weights are free for commercial use after registration.

    Why it mattered. Tool calling and code execution moved the open ChatGLM line from chat towards agent tasks.

    AVAILABLE TO
    Open weights
  2. MODEL

    ChatGLM2-6B

    • OPEN WEIGHTS
    • SUPERSEDED
    • CHATGLM OPEN MODELS

    The second-generation ChatGLM-6B extends context from 2K to 32K tokens, speeds up inference by 42% and lifts benchmark scores, including 23% on MMLU and 571% on GSM8K over the first version.

    Why it mattered. Longer conversations and faster inference on the same consumer hardware kept the open ChatGLM line practical for local deployment.

    AVAILABLE TO
    Open weights
  3. MODEL

    ChatGLM and ChatGLM-6B

    • WATERSHED
    • OPEN WEIGHTS
    • SUPERSEDED
    • CHATGLM OPEN MODELS

    Zhipu AI introduces ChatGLM, a bilingual Chinese–English dialogue model based on GLM-130B, at chatglm.cn, and open-sources ChatGLM-6B, a 6.2-billion-parameter version that runs locally in 6GB of GPU memory at INT4.

    Why it mattered. ChatGLM-6B gave developers a Chinese–English chat model they could fine-tune and deploy on consumer graphics cards, and started the GLM habit of pairing a hosted assistant with open weights.

    AVAILABLE TO
    Open weights · Web
§ 04 — ERAS

4 eras of GLM.

  1. ERA 01MARCH 2023OCTOBER 2023

    ChatGLM open models

    Zhipu AI (now Z.ai) launches ChatGLM and ships three generations of the open ChatGLM-6B, adding longer context and then native tool calling.

    3 releases, from ChatGLM and ChatGLM-6B
  2. ERA 02JANUARY 2024NOVEMBER 2024

    GLM-4 and first agents

    GLM-4 arrives with an open 9B version, the Qingyan app adds video calls, and AutoGLM starts operating phone apps and websites for users.

    3 releases, from GLM-4-9B open models
  3. ERA 03APRIL 2025DECEMBER 2025

    Z.ai and agentic GLM-4.x

    Open GLM-Z1 reasoning models launch with the Z.ai chat site, then GLM-4.5 to 4.7 and the 4.5V and 4.6V vision models turn GLM into an open, MIT-licensed choice for coding agents.

    6 releases, from GLM-4-0414, GLM-Z1 and Z.ai
  4. ERA 04FEBRUARY 2026NOW

    GLM-5 long-horizon

    GLM-5 scales to 744B parameters, followed by 5V-Turbo, 5.1’s eight-hour runs, a Coding Plan team tier, 5.2’s 1M context, the security-focused 5.3 and the multimodal 5.3-Flash.

    7 releases, from GLM-5
§ 05 — QUESTIONS

GLM releases, answered.

As of our 23 Sep 2026 review, GLM-5.3-Flash, announced by Z.ai on 26 Aug 2026. GLM-5.3-Flash is a newly trained 320B-parameter model (18B active) and the first natively multimodal model in the GLM-5 series, using a hybrid of sparse and linear attention for cheaper 1M-token contexts. Z.ai says it beats GLM-5.2 at one-tenth the price and approaches Claude Opus 4.8 on coding and agent benchmarks; weights are MIT-licensed.

14 Mar 2023: ChatGLM and ChatGLM-6B. Zhipu AI introduces ChatGLM, a bilingual Chinese–English dialogue model based on GLM-130B, at chatglm.cn, and open-sources ChatGLM-6B, a 6.2-billion-parameter version that runs locally in 6GB of GPU memory at INT4.

As of 23 Sep 2026, the newest release in each line is: GLM-5V-Turbo (2 Apr 2026), GLM-5.3 (14 Aug 2026) and GLM-5.3-Flash (26 Aug 2026). Older models in those lines have been superseded, and some have been retired from Z.ai’s apps.

19 on this timeline: 16 model releases and 3 product, platform, agent and plan launches. It leaves out incremental updates, settings and feature drops — only releases that changed what GLM could do or who could use it.

GLM-5.3: Its weights are public, so it can be hosted in Australia on your own infrastructure or an Australian cloud region. GLM-5.3 Flash: Its weights are public, so it can be hosted in Australia on your own infrastructure or an Australian cloud region. Our benchmark pages track Australian cloud availability for each model weekly.

Dates are the day Z.ai announced each release, taken from Z.ai’s own site and documentation.

Major releases only: new models, products, plans and capability launches. Retirements are noted on the model they retired.

Checked 23 Sep 2026 · How the current models compare →

MAKE SENSE OF THE NEXT RELEASE

Choose AI improvements around your business needs.