Active Parameters
671B
Context Length
128K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
MIT License
Release Date
21 Aug 2025
Knowledge Cutoff
-
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
87x RTX 4090
24GB VRAM
Datacenter
22x NVIDIA A100
80GB VRAM
Apple Silicon
18x Apple M3 Max
128GB VRAM
128,000 tokens
Consumer
127x RTX 4090
24GB VRAM
Datacenter
32x NVIDIA A100
80GB VRAM
Apple Silicon
27x Apple M3 Max
128GB VRAM
Rank
#130
| Benchmark | Score | Rank |
|---|---|---|
StackUnseen | 0.481 | 24 |
Professional Knowledge | 0.84 | 55 |
General Text | 1417 | 80 |
Overall Rank
#130
Coding Rank
#73
A hybrid model that supports both "thinking" and "non-thinking" modes for chat, reasoning, and coding. It's a Mixture-of-Experts (MoE) model with a massive context length and efficient architecture.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
128
Key-Value Heads
128
Attention Head Dimension
-
Position Embedding
ROPE
RoPE Theta
10,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
7,168
Number of Layers
61
FFN Intermediate Size (Dense)
2,048
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
129,280
Mixture of Experts
Total Expert Parameters
37.0B
Number of Experts
257
Active Experts
8
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
3
DeepSeek-V3 is a Mixture-of-Experts (MoE) language model comprising 671B parameters with 37B activated per token. Its architecture incorporates Multi-head Latent Attention and DeepSeekMoE for efficient inference and training. Innovations include an auxiliary-loss-free load balancing strategy and a multi-token prediction objective, trained on 14.8T tokens.
APX AI
Online