ApX logoApX logo

GLM-5.3

Active Parameters

753.9B

Context Length

1.05M

Modality

Text

Architecture

Mixture of Experts (MoE)

License

other

Release Date

25 Aug 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Input: $1.40 · Output: $4.40

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

1588.81 GB VRAM

Consumer

99x RTX 4090

24GB VRAM

Datacenter

25x NVIDIA A100

80GB VRAM

Apple Silicon

21x Apple M3 Max

128GB VRAM

1,048,576 tokens

5805.78 GB VRAM

Consumer

517x RTX 4090

24GB VRAM

Datacenter

111x NVIDIA A100

80GB VRAM

Apple Silicon

120x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 6.1k · Context: 1.05M · Vocab: 154.9kx 78 layersRMSNormPre-AttentionDeepSeek Sparse Attention64Q / 64KV headsHead dim: 192+RMSNormPre-FFNSparse MoE FFN (8/256 experts)SwiGLUIntermediate: 2k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#10

BenchmarkScoreRank

Agentic Index

Artificial Analysis
max

0.53

6

Agentic Coding

LiveBench Agentic

0.61

11

Web Development

WebDev Arena
max

1614

16

Intelligence Index

Artificial Analysis
max

0.45

16

General Text

Text Arena
max

1483

17

0.79

18

Agent Arena

Agent Arena
max

0.03

22

LiveBench Average

LiveBench Average

0.76

22

0.76

23

max

0.75

24

0.86

26

0.88

32

0.70

37

Rankings

Overall Rank

#10

Coding Rank

#11

About GLM-5.3

GLM-5.3 is a large-scale mixture-of-experts language model developed by Zhipu AI / Z.ai utilizing dynamic sparse attention. It provides state-of-the-art multilingual comprehension, complex mathematical reasoning, and coding capabilities.

Technical Specifications

Attention

Attention Structure

DeepSeek Sparse Attention

Attention Heads

64

Key-Value Heads

64

Attention Head Dimension

192

Position Embedding

ROPE

RoPE Theta

8,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

No

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

6,144

Number of Layers

78

FFN Intermediate Size (Dense)

12,288

Multi-Token Prediction Heads

1

Tokenizer

Vocabulary Size

154,880

Mixture of Experts

Total Expert Parameters

51.6B

Number of Experts

256

Active Experts

8

Shared Experts

1

FFN Intermediate Size (per Expert)

2,048

Dense Layers Before MoE

3

About GLM Family

General Language Models from Z.ai


Other GLM Family Models