ApX logoApX logo

GLM-5.2

Active Parameters

744B

Context Length

1M

Modality

Text

Architecture

Mixture of Experts (MoE)

License

MIT

Release Date

13 Jun 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Input: $1.40 · Output: $4.40

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

1568.02 GB VRAM

Consumer

98x RTX 4090

24GB VRAM

Datacenter

25x NVIDIA A100

80GB VRAM

Apple Silicon

21x Apple M3 Max

128GB VRAM

1,000,000 tokens

5589.45 GB VRAM

Consumer

491x RTX 4090

24GB VRAM

Datacenter

106x NVIDIA A100

80GB VRAM

Apple Silicon

113x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 6.1k · Context: 1M · Vocab: 154.9kx 78 layersRMSNormPre-AttentionDeepSeek Sparse Attention64Q / 64KV headsHead dim: 192+RMSNormPre-FFNSparse MoE FFN (8/256 experts)SwishIntermediate: 2k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#20

BenchmarkScoreRank

0.867

9

0.80

17

Agent Arena

Agent Arena
max

0.04

18

Graduate-Level QA

GPQA

0.912

20

Web Development

WebDev Arena
max

1598

21

Agentic Coding

LiveBench Agentic

0.52

30

0.90

30

0.74

34

General Text

Text Arena
max

1472

37

0.73

39

LiveBench Average

LiveBench Average

0.73

39

0.79

43

Agentic Index

Artificial Analysis
max

0.38

45

max

0.69

Standard

0.47

50

98

Intelligence Index

Artificial Analysis
max

0.34

Standard

0.22

69

138

Rankings

Overall Rank

#20

Coding Rank

#14

About GLM-5.2

Z.ai's flagship open-weights foundation model released June 13, 2026. A 744-billion-parameter Mixture-of-Experts model with 40 billion active parameters per token, powered by IndexShare technology. Features a stable 1M-token context window optimised for long-horizon coding and agentic workflows, multiple thinking-effort levels (High and Max), and strong benchmark performance (81.0 on Terminal-Bench 2.1, 62.1 on SWE-bench Pro). Supports thinking mode, streaming, function calling, context caching, structured output, and MCP integration. Released under MIT license with no regional restrictions. Priced at $1.40 per million input tokens and $4.40 per million output tokens.

Technical Specifications

Attention

Attention Structure

DeepSeek Sparse Attention

Attention Heads

64

Key-Value Heads

64

Attention Head Dimension

192

Position Embedding

ROPE

RoPE Theta

8,000,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

6,144

Number of Layers

78

FFN Intermediate Size (Dense)

12,288

Multi-Token Prediction Heads

1

Tokenizer

Vocabulary Size

154,880

Mixture of Experts

Total Expert Parameters

40.0B

Number of Experts

256

Active Experts

8

Shared Experts

1

FFN Intermediate Size (per Expert)

2,048

Dense Layers Before MoE

3

About GLM-5.2

Z.ai's GLM-5.2 is a 744-billion-parameter Mixture-of-Experts flagship foundation model, released June 13, 2026, designed for long-horizon coding and agentic engineering tasks. It features a usable 1M-token context window, IndexShare architecture reducing compute to 1/20th of prior generations, and multiple thinking-effort levels. Open-weights under MIT license with no regional restrictions.


Other GLM-5.2 Models
  • No related models available
GLM-5.2: Specifications and GPU VRAM Requirements