Active Parameters
671B
Context Length
128K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
MIT
Release Date
10 Jan 2026
Knowledge Cutoff
May 2025
API Pricing (per 1M)
Input: $0.28 · Output: $0.42
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
86x RTX 4090
24GB VRAM
Datacenter
22x NVIDIA A100
80GB VRAM
Apple Silicon
18x Apple M3 Max
128GB VRAM
128,000 tokens
Consumer
87x RTX 4090
24GB VRAM
Datacenter
22x NVIDIA A100
80GB VRAM
Apple Silicon
18x Apple M3 Max
128GB VRAM
Rank
#72
| Benchmark | Score | Rank |
|---|---|---|
Professional Knowledge | 0.85 | 12 |
Software Engineering | auto 0.60 Standard 0.70 | 27 15 |
Graduate-Level QA | 0.824 | 56 |
Web Development | auto 1361 Standard 1325 | 80 92 |
General Text | auto 1422 Standard 1425 | 88 85 |
Coding Index | auto 0.44 | 102 |
Intelligence Index | 0.16 | 155 |
Coding Archived | 0.74 | 5 |
Overall Rank
#72
Coding Rank
#79
DeepSeek-V3.2 is an advanced 671B Mixture-of-Experts model from DeepSeek optimized for autonomous agents and complex reasoning. Featuring DeepSeek Sparse Attention and integrated thinking modes, it supports long-horizon execution across a 164K context window.
Attention
Attention Structure
DeepSeek Sparse Attention
Attention Heads
128
Key-Value Heads
1
Attention Head Dimension
-
Position Embedding
Absolute Position Embedding
RoPE Theta
10,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Auxiliary Parameters
-
Hidden Dimension Size
7,168
Number of Layers
61
FFN Intermediate Size (Dense)
2,048
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
129,280
Mixture of Experts
Total Expert Parameters
37.0B
Number of Experts
257
Active Experts
9
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
3
DeepSeek-V3 is a Mixture-of-Experts (MoE) language model comprising 671B parameters with 37B activated per token. Its architecture incorporates Multi-head Latent Attention and DeepSeekMoE for efficient inference and training. Innovations include an auxiliary-loss-free load balancing strategy and a multi-token prediction objective, trained on 14.8T tokens.
Assistant
Online