Active Parameters
744B
Context Length
205K
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
MIT
Release Date
12 Feb 2026
Knowledge Cutoff
Dec 2025
API Pricing (per 1M)
Input: $1.00 · Output: $3.20
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
98x RTX 4090
24GB VRAM
Datacenter
25x NVIDIA A100
80GB VRAM
Apple Silicon
21x Apple M3 Max
128GB VRAM
204,800 tokens
Consumer
119x RTX 4090
24GB VRAM
Datacenter
30x NVIDIA A100
80GB VRAM
Apple Silicon
26x Apple M3 Max
128GB VRAM
Rank
#60
| Benchmark | Score | Rank |
|---|---|---|
Software Engineering | 0.73 | 10 |
StackUnseen | 0.551 | 21 |
General Text | 1458 | 45 |
Web Development | 1436 | 53 |
Intelligence Index | 0.26 | 114 |
Overall Rank
#60
Coding Rank
#55
GLM-5 is Z.ai's 744B-parameter flagship Mixture-of-Experts multimodal foundation model activating 40B parameters for complex systems engineering. Featuring DeepSeek Sparse Attention and open weights, it supports up to 128K token generations over a 205K context window.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
64
Key-Value Heads
64
Attention Head Dimension
64
Position Embedding
Absolute Position Embedding
RoPE Theta
1,000,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
Swish
Dimensions
Hidden Dimension Size
6,144
Number of Layers
80
FFN Intermediate Size (Dense)
2,048
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
154,880
Mixture of Experts
Total Expert Parameters
40.0B
Number of Experts
256
Active Experts
8
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
3
GLM 5 is the fifth generation of General Language Models developed by Z.ai. It represents a significant leap in multimodal foundational capabilities, featuring advanced reasoning and long-horizon agentic capabilities across diverse systems engineering tasks.
APX AI
Online