Active Parameters
671B
Context Length
131K
Modality
Text
Architecture
Mixture of Experts (MoE)
License
DeepSeek Model License
Release Date
27 Dec 2024
Knowledge Cutoff
Jul 2024
API Pricing (per 1M)
Input: $0.32 · Output: $0.89
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
87x RTX 4090
24GB VRAM
Datacenter
22x NVIDIA A100
80GB VRAM
Apple Silicon
18x Apple M3 Max
128GB VRAM
131,072 tokens
Consumer
128x RTX 4090
24GB VRAM
Datacenter
32x NVIDIA A100
80GB VRAM
Apple Silicon
28x Apple M3 Max
128GB VRAM
Rank
#114
| Benchmark | Score | Rank |
|---|---|---|
Professional Knowledge | 0.812 | 18 |
StackUnseen | 0.439 | 27 |
Software Engineering | 0.42 | 32 |
Graduate-Level QA | 0.684 | 73 |
General Text | 1396 | 88 |
Agentic Index | 0.8 | 111 |
Coding Index | 0.23 | 112 |
Intelligence Index | 0.10 | 194 |
StackEval Archived | 0.976 | 🥈 2 |
General Knowledge Reference | 0.885 | 5 |
QA Assistant Archived | 0.953 | 9 |
Coding Archived | 0.55 | 11 |
Summarization Archived | 0.806 | 11 |
Overall Rank
#114
Coding Rank
#94
DeepSeek-V3 is a 671B-parameter Mixture-of-Experts foundation model activating 37B parameters per token for efficient general language processing. Utilizing Multi-Head Latent Attention and Multi-Token Prediction, it excels across coding, math, and 128K long contexts.
Attention
Attention Structure
Multi-Layer Attention
Attention Heads
128
Key-Value Heads
128
Attention Head Dimension
-
Position Embedding
ROPE
RoPE Theta
10,000
Sliding Window Attention
No
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
Swish
Dimensions
Hidden Dimension Size
7,168
Number of Layers
61
FFN Intermediate Size (Dense)
2,048
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
129,280
Mixture of Experts
Total Expert Parameters
37.0B
Number of Experts
257
Active Experts
9
Shared Experts
1
FFN Intermediate Size (per Expert)
2,048
Dense Layers Before MoE
3
DeepSeek-V3 is a Mixture-of-Experts (MoE) language model comprising 671B parameters with 37B activated per token. Its architecture incorporates Multi-head Latent Attention and DeepSeekMoE for efficient inference and training. Innovations include an auxiliary-loss-free load balancing strategy and a multi-token prediction objective, trained on 14.8T tokens.
APX AI
Online