Active Parameters
671B
Context Length
128K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
MIT License
Release Date
21 Aug 2025
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $1.64 · Output: $2.75
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
80x RTX 4090
24GB VRAM
Datacenter
21x NVIDIA A100
80GB VRAM
Apple Silicon
17x Apple M3 Max
128GB VRAM
128,000 tokens
Consumer
117x RTX 4090
24GB VRAM
Datacenter
29x NVIDIA A100
80GB VRAM
Apple Silicon
25x Apple M3 Max
128GB VRAM
Rank
#104
| Benchmark | Score | Rank |
|---|---|---|
Professional Knowledge | 0.837 | 18 |
StackUnseen | 0.481 | 24 |
Graduate-Level QA | 0.749 | 78 |
General Text | auto 1416 Standard 1417 | 98 97 |
Coding Index | 0.43 | 105 |
Agentic Index | auto 0.07 | 107 |
Intelligence Index | 0.14 | 197 |
Overall Rank
#104
Coding Rank
#88
A hybrid model that supports both "thinking" and "non-thinking" modes for chat, reasoning, and coding. It's a Mixture-of-Experts (MoE) model with a massive context length and efficient architecture.
Attention
Attention Structure
Multi-Head Attention
Attention Heads
128
Key-Value Heads
128
Attention Head Dimension
-
Position Embedding
ROPE
RoPE Theta
10,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Auxiliary Parameters
-
Hidden Dimension Size
7,168
Number of Layers
61
FFN Intermediate Size (Dense)
2,048
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
129,280
Mixture of Experts
Total Expert Parameters
37.0B
Number of Experts
257
Active Experts
8
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
3
DeepSeek-V3 is a Mixture-of-Experts (MoE) language model comprising 671B parameters with 37B activated per token. Its architecture incorporates Multi-head Latent Attention and DeepSeekMoE for efficient inference and training. Innovations include an auxiliary-loss-free load balancing strategy and a multi-token prediction objective, trained on 14.8T tokens.
Assistant
Online