ApX logoApX logo

GLM-5.3

Active Parameters

753.9B

Context Length

1.05M

Modality

Text

Architecture

Mixture of Experts (MoE)

License

other

Release Date

25 Aug 2026

Knowledge Cutoff

-

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

1588.81 GB VRAM

Consumer

99x RTX 4090

24GB VRAM

Datacenter

25x NVIDIA A100

80GB VRAM

Apple Silicon

21x Apple M3 Max

128GB VRAM

1,048,576 tokens

5805.78 GB VRAM

Consumer

517x RTX 4090

24GB VRAM

Datacenter

111x NVIDIA A100

80GB VRAM

Apple Silicon

120x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 6.1k · Context: 1.05M · Vocab: 154.9kx 78 layersRMSNormPre-AttentionDeepSeek Sparse Attention64Q / 64KV headsHead dim: 192+RMSNormPre-FFNSparse MoE FFN (8/256 experts)SwiGLUIntermediate: 2k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#43

BenchmarkScoreRank

Agentic Coding

LiveBench Agentic

0.61

10

Web Development

max
WebDev Arena

1609

12

General Text

max
Text Arena

1482

18

0.79

22

0.76

23

0.86

27

LiveBench Average

LiveBench Average

0.76

27

Agent Arena

max
Agent Arena

0.04

30

0.70

44

0.88

45

Rankings

Overall Rank

#43

Coding Rank

#36

About GLM-5.3

GLM-5.3 is a large-scale mixture-of-experts language model developed by Zhipu AI / Z.ai utilizing dynamic sparse attention. It provides state-of-the-art multilingual comprehension, complex mathematical reasoning, and coding capabilities.

Technical Specifications

Attention

Attention Structure

DeepSeek Sparse Attention

Attention Heads

64

Key-Value Heads

64

Attention Head Dimension

192

Position Embedding

ROPE

RoPE Theta

8,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

No

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

6,144

Number of Layers

78

FFN Intermediate Size (Dense)

12,288

Multi-Token Prediction Heads

1

Tokenizer

Vocabulary Size

154,880

Mixture of Experts

Total Expert Parameters

51.6B

Number of Experts

256

Active Experts

8

Shared Experts

1

FFN Intermediate Size (per Expert)

2,048

Dense Layers Before MoE

3

About GLM Family

General Language Models from Z.ai


Other GLM Family Models
GLM-5.3: Specifications and GPU VRAM Requirements