Active Parameters
358B
Context Length
200K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
MIT
Release Date
8 Jan 2026
Knowledge Cutoff
Sep 2024
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
42x RTX 4090
24GB VRAM
Datacenter
11x NVIDIA A100
80GB VRAM
Apple Silicon
9x Apple M3 Max
128GB VRAM
200,000 tokens
Consumer
47x RTX 4090
24GB VRAM
Datacenter
13x NVIDIA A100
80GB VRAM
Apple Silicon
10x Apple M3 Max
128GB VRAM
Rank
#35
| Benchmark | Score | Rank |
|---|---|---|
Graduate-Level QA | 0.857 | 8 |
Professional Knowledge | 0.83 | 28 |
Web Development | 1434 | 51 |
General Text | 1442 | 67 |
Overall Rank
#35
Coding Rank
#40
GLM-4.7 is a large-scale Mixture of Experts (MoE) model developed by Z.ai, specifically architected to support advanced agentic coding, complex reasoning, and multi-step tool orchestration. Building upon the GLM-4 series, the model integrates a sophisticated reasoning system that prioritizes logical consistency and task completion across extended interactions. It is designed to function as a primary engine for coding agents and terminal-based automation, featuring optimizations for multi-language programming and autonomous execution within complex software environments.
The model's technical foundation includes a triple-tier thinking architecture designed to maintain reasoning coherence. Interleaved Thinking allows the model to perform internal reasoning steps before every response and tool invocation, ensuring that generated instructions align with logical constraints. Preserved Thinking facilitates the retention of these reasoning blocks across multi-turn conversations, preventing the context decay typically seen in long-horizon tasks. Additionally, Turn-level Thinking provides a granular control mechanism, allowing developers to adjust reasoning depth based on the specific requirements of each interaction to manage computational overhead and latency effectively.
Beyond programming, GLM-4.7 features a refined approach to frontend and user interface development, often referred to as vibe coding. This capability focuses on generating aesthetically consistent and structurally sound UI code, including modern web pages and professional presentation layouts. The model's architecture also emphasizes robust tool integration, enabling it to navigate terminal environments, execute shell commands, and interact with external APIs while maintaining a high degree of stability and instruction adherence in diverse automation scenarios.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
96
Key-Value Heads
8
Attention Head Dimension
128
Position Embedding
Absolute Position Embedding
RoPE Theta
1,000,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
Swish
Dimensions
Hidden Dimension Size
5,120
Number of Layers
92
FFN Intermediate Size (Dense)
1,536
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
151,552
Mixture of Experts
Total Expert Parameters
32.0B
Number of Experts
160
Active Experts
8
Shared Experts
1
FFN Intermediate Size (per Expert)
1,536
Dense Layers Before MoE
3
GLM-4 is a series of bilingual (English and Chinese) language models developed by Zhipu AI. The models feature extended context windows, superior coding performance, advanced reasoning capabilities, and strong agent functionalities. GLM-4.6 offers improvements in tool use and search-based agents.
APX AI
Online