Active Parameters
313.3B
Context Length
1.05M
Modality
Text
Architecture
Mixture of Experts (MoE)
License
MIT License
Release Date
25 Aug 2026
Knowledge Cutoff
-
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
36x RTX 4090
24GB VRAM
Datacenter
10x NVIDIA A100
80GB VRAM
Apple Silicon
8x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
154x RTX 4090
24GB VRAM
Datacenter
38x NVIDIA A100
80GB VRAM
Apple Silicon
33x Apple M3 Max
128GB VRAM
Rank
#100
| Benchmark | Score | Rank |
|---|---|---|
Web Development | 1604 | ⭐ 13 |
Agentic Coding | 0.57 | 16 |
Coding | 0.79 | 22 |
Data Analysis | 0.76 | 28 |
Agent Arena | 0.04 | 29 |
General Text | 1473 | 34 |
General | 0.72 | 42 |
Reasoning | 0.78 | 51 |
LiveBench Average | 0.72 | 52 |
Mathematics | 0.81 | 60 |
Overall Rank
#100
Coding Rank
#16
GLM-5.3 Flash is a high-efficiency open model designed for rapid inference and low-latency reasoning workloads. Built on the Glm5Next architecture, it offers optimized throughput for agentic applications and code generation.
Attention
Attention Structure
DeepSeek Sparse Attention
Attention Heads
64
Key-Value Heads
64
Attention Head Dimension
128
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
Yes
Linear Attention Ratio
75.6%
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
4,096
Number of Layers
45
FFN Intermediate Size (Dense)
12,288
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
154,880
Mixture of Experts
Total Expert Parameters
17.3B
Number of Experts
288
Active Experts
8
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
3
General Language Models from Z.ai
APX AI
Online