ApX logoApX logo

DeepSeek-V3.2

Active Parameters

671B

Context Length

128K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

MIT

Release Date

10 Jan 2026

Knowledge Cutoff

May 2025

API Pricing (per 1M)

Input: $0.28 · Output: $0.42

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

1410.63 GB VRAM

Consumer

86x RTX 4090

24GB VRAM

Datacenter

22x NVIDIA A100

80GB VRAM

Apple Silicon

18x Apple M3 Max

128GB VRAM

128,000 tokens

1414.80 GB VRAM

Consumer

87x RTX 4090

24GB VRAM

Datacenter

22x NVIDIA A100

80GB VRAM

Apple Silicon

18x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 7.2k · Context: 128K · Vocab: 129.3kx 61 layersRMSNormPre-AttentionDeepSeek Sparse Attention128Q / 1KV headsHead dim: 56+RMSNormPre-FFNSparse MoE FFN (9/257 experts)SwiGLUIntermediate: 2k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#72

BenchmarkScoreRank

Professional Knowledge

MMLU Pro

0.85

12

Software Engineering

SWE-bench Verified
auto

0.60

Standard

0.70

27

15

Graduate-Level QA

GPQA

0.824

56

Web Development

WebDev Arena
auto

1361

Standard

1325

80

92

General Text

Text Arena
auto

1422

Standard

1425

88

85

auto

0.44

102

Intelligence Index

Artificial Analysis

0.16

155

Coding

Archived
Aider Coding

0.74

5

Rankings

Overall Rank

#72

Coding Rank

#79

About DeepSeek-V3.2

DeepSeek-V3.2 is an advanced 671B Mixture-of-Experts model from DeepSeek optimized for autonomous agents and complex reasoning. Featuring DeepSeek Sparse Attention and integrated thinking modes, it supports long-horizon execution across a 164K context window.

Technical Specifications

Attention

Attention Structure

DeepSeek Sparse Attention

Attention Heads

128

Key-Value Heads

1

Attention Head Dimension

-

Position Embedding

Absolute Position Embedding

RoPE Theta

10,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

7,168

Number of Layers

61

FFN Intermediate Size (Dense)

2,048

Multi-Token Prediction Heads

1

Tokenizer

Vocabulary Size

129,280

Mixture of Experts

Total Expert Parameters

37.0B

Number of Experts

257

Active Experts

9

Shared Experts

1

FFN Intermediate Size (per Expert)

2,048

Dense Layers Before MoE

3

About DeepSeek-V3

DeepSeek-V3 is a Mixture-of-Experts (MoE) language model comprising 671B parameters with 37B activated per token. Its architecture incorporates Multi-head Latent Attention and DeepSeekMoE for efficient inference and training. Innovations include an auxiliary-loss-free load balancing strategy and a multi-token prediction objective, trained on 14.8T tokens.


Other DeepSeek-V3 Models