Active Parameters
753.9B
Context Length
1.05M
Modality
Text
Architecture
Mixture of Experts (MoE)
License
other
Release Date
25 Aug 2026
Knowledge Cutoff
-
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
99x RTX 4090
24GB VRAM
Datacenter
25x NVIDIA A100
80GB VRAM
Apple Silicon
21x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
517x RTX 4090
24GB VRAM
Datacenter
111x NVIDIA A100
80GB VRAM
Apple Silicon
120x Apple M3 Max
128GB VRAM
Rank
#43
| Benchmark | Score | Rank |
|---|---|---|
Agentic Coding | 0.61 | 10 |
Web Development max | 1609 | ⭐ 12 |
General Text max | 1482 | ⭐ 18 |
Coding | 0.79 | 22 |
General | 0.76 | 23 |
Reasoning | 0.86 | 27 |
LiveBench Average | 0.76 | 27 |
Agent Arena max | 0.04 | 30 |
Data Analysis | 0.70 | 44 |
Mathematics | 0.88 | 45 |
Overall Rank
#43
Coding Rank
#36
GLM-5.3 is a large-scale mixture-of-experts language model developed by Zhipu AI / Z.ai utilizing dynamic sparse attention. It provides state-of-the-art multilingual comprehension, complex mathematical reasoning, and coding capabilities.
Attention
Attention Structure
DeepSeek Sparse Attention
Attention Heads
64
Key-Value Heads
64
Attention Head Dimension
192
Position Embedding
ROPE
RoPE Theta
8,000,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
No
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
6,144
Number of Layers
78
FFN Intermediate Size (Dense)
12,288
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
154,880
Mixture of Experts
Total Expert Parameters
51.6B
Number of Experts
256
Active Experts
8
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
3
General Language Models from Z.ai
APX AI
Online