GLM-4.7-Flash

Updated
05.10.2026
Tools
Thinking
Reasoning
Code

GLM-4.7-Flash is Z.ai’s 30B-A3B MoE under MIT: 31.2B parameters, about 3B active per token, and 2-bit GGUF builds fit in 12 GB.

At a glance

  • License: MIT
  • Parameters: 31.2B total, about 3B active per token (30B-A3B MoE)
  • Max new tokens (vendor evals): 131,072
  • Modalities: Text in and out, English and Chinese
  • Minimum hardware: 12 GB of RAM or VRAM with the 2-bit GGUF builds

What is GLM-4.7-Flash?

GLM-4.7-Flash is a mixture of experts model from Z.ai. It carries 31.2B total parameters but activates only about 3B per token, which is what the 30B-A3B label means. Z.ai calls it the strongest model in the 30B class and presents it as a new option for lightweight deployment that balances performance and efficiency. The weights went up on Hugging Face on January 19, 2026 under the MIT license, and the card names vLLM and SGLang for local serving, plus a transformers example for running the weights directly.

SpecificationGLM-4.7-Flash
Total parameters31.2B (31,221,488,576 exactly)
Active parametersAbout 3B per token (30B-A3B)
ArchitectureMixture of experts (MoE)
TaskText generation, text in and out
LanguagesEnglish and Chinese
ReasoningThinking mode; Preserved Thinking recommended for multi-turn agent tasks
Speculative decodingMTP (vLLM) and EAGLE (SGLang) in the vendor serve configs
Vendor sampling defaultsTemperature 1.0, top-p 0.95
Max new tokens in vendor evals131,072
Serving stacksvLLM and SGLang (main branches only), plus Transformers
Parsersglm47 tool-call parser, glm45 reasoning parser
Release dateJanuary 19, 2026
LicenseMIT

Z.ai's own serve commands ship with speculative decoding switched on: MTP under vLLM with one speculative token, and EAGLE under SGLang with three speculative steps and four draft tokens. That is the mechanism covered in our speculative decoding guide. Both commands also name the same two parsers, glm47 for tool calls and glm45 for reasoning, and the vLLM one turns on automatic tool choice as well. For multi-turn agentic tasks, which Z.ai lists as τ²-Bench and Terminal Bench 2, the card says to turn on Preserved Thinking mode.

GLM-4.7-Flash benchmarks

Z.ai's launch numbers, from the model card, compare GLM-4.7-Flash with Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B:

BenchmarkGLM-4.7-FlashQwen3-30B-A3B-Thinking-2507GPT-OSS-20B
AIME 25
Competition math
91.685.091.7
GPQA
Expert science
75.273.471.5
LCB v6
Competitive coding
64.066.061.0
HLE
Expert questions
14.49.810.9
SWE-bench Verified
Software engineering
59.222.034.0
τ²-Bench
Agentic tools
79.549.047.7
BrowseComp
Web browsing
42.82.2928.3

GLM-4.7-Flash takes five of the seven rows, with the largest margins on SWE-bench Verified, τ²-Bench and BrowseComp. It gives up AIME 25 to GPT-OSS-20B by a tenth of a point and LCB v6 to the Qwen model by two.

Z.ai did not run every row at the same settings. Terminal Bench and SWE-bench Verified were scored at temperature 0.7, top-p 1.0 and a 16,384-token generation cap, and τ²-Bench at temperature 0 with the same cap, rather than at the defaults in the table above. For τ²-Bench it also added a prompt to the Retail and Telecom user interaction to avoid failures caused by users ending the interaction incorrectly, and applied the Airline domain fixes from the Claude Opus 4.5 release report. Z.ai also says to run that benchmark and Terminal Bench 2 with Preserved Thinking on, so the agentic scores describe a specific setup, not a plain chat.

GLM-4.7-Flash hardware requirements

The system requirement to check is memory. The GGUF builds below come from the community repo unsloth/GLM-4.7-Flash-GGUF, and the sizes are the real file sizes from that listing.

MemoryBuild to pickFile size
10 GBUD-TQ1_08.33 GB
12 GBUD-IQ2_M10.99 GB
16 GBUD-Q3_K_XL13.78 GB
20 GBIQ4_XS16.27 GB
24 GBQ4_K_M18.31 GB
32 GBQ6_K24.69 GB
48 GB and upQ8_031.84 GB

Neighbouring files differ by a gigabyte or two, so when two builds both fit, take the larger one. The listing goes further down than the table does, to a 9.25 GB UD-IQ1_S and a 9.81 GB UD-IQ1_M; reach for those 1-bit builds only when nothing else fits. The unquantized BF16 weights are also there, 59.91 GB across two files. If the format is new to you, start with what GGUF is.

Serving the full weights on GPUs is the other path: Z.ai's reference commands run at tensor parallel size 4 under both vLLM and SGLang, and both frameworks support the model only on their main branches. On Blackwell cards the SGLang command also needs the triton attention backends.

How to run GLM-4.7-Flash in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for GLM-4.7-Flash in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

Leave the sampler at Z.ai's defaults for most tasks, temperature 1.0 with top-p 0.95, and give thinking room: the vendor evals allow up to 131,072 new tokens per answer. For the rest of the family, see every GLM model you can run locally.

GLM-4.7-Flash license

GLM-4.7-Flash is released under the MIT license. That permits commercial use, modification and redistribution, provided the copyright and license notice travel with the software, which makes it one of the most permissive licenses an open-weight model can carry.

Get the weights from Hugging Face

huggingface-cli download zai-org/GLM-4.7-Flash
from transformers import AutoModel
model = AutoModel.from_pretrained("zai-org/GLM-4.7-Flash")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

GLM-4.7-Flash is a 31.2B-parameter open-weight language model from zai-org (Z.ai), the smallest member of the GLM-4.7 family. It uses a lightweight Mixture-of-Experts architecture and supports a 128K-token context, with capabilities for code, reasoning, tool calling, and a thinking mode. It is positioned as a fast option for coding and agentic work that can run on a single GPU.

A 4-bit quantized version runs on about 16GB of VRAM, which covers cards like the RTX 3090 and 4090 or an M-series Mac with enough unified memory. Running it at full BF16 precision needs roughly 64GB. Thanks to the MoE design you can also offload to system RAM and trade speed for lower VRAM use.

Yes. The weights are released under the MIT license, so you can download and run them at no cost. Running the model locally in Atomic Chat means there are no API fees or per-token charges, so you only pay for your own hardware and electricity.

Yes. Once you download the weights, GLM-4.7-Flash runs fully on-device with no internet connection required. Your prompts, code, and files never leave your machine, which makes it suitable for private or air-gapped work.

It is built for coding, reasoning, and agentic tasks. The code capability suits writing and debugging across files, the tools capability lets it call functions for local agent setups, and the thinking mode plus 128K context handle multi-step reasoning over long inputs. Reported throughput is around 60-100 tokens per second on consumer GPUs.