Active Parameters
1.61T
Context Length
1.05M
Modality
Reasoning
Architecture
Mixture of Experts (MoE)
License
MIT License
Release Date
13 Aug 2026
Knowledge Cutoff
-
API Pricing (per 1M)
Input: $1.32 · Output: $3.96
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
254x RTX 4090
24GB VRAM
Datacenter
59x NVIDIA A100
80GB VRAM
Apple Silicon
56x Apple M3 Max
128GB VRAM
1,048,576 tokens
Consumer
267x RTX 4090
24GB VRAM
Datacenter
62x NVIDIA A100
80GB VRAM
Apple Silicon
59x Apple M3 Max
128GB VRAM
Rank
#29
| Benchmark | Score | Rank |
|---|---|---|
Mathematics | 0.95 | 8 |
Data Analysis | 0.79 | 11 |
General | 0.77 | 15 |
LiveBench Average | 0.77 | 15 |
Agent Arena | high 0.04 | 16 |
Agentic Coding | 0.55 | 20 |
Web Development | high 1580 | 21 |
Reasoning | 0.86 | 24 |
Coding | 0.77 | 27 |
Agentic Index | max 0.42 | 32 |
General Text | high 1460 | 43 |
Intelligence Index | max 0.36 | 47 |
Coding Index | max 0.69 | 50 |
Overall Rank
#29
Coding Rank
#24
DeepSeek V4 Pro 0813 represents the general availability release of DeepSeek's flagship mixture-of-experts model. It provides frontier reasoning, complex software engineering capabilities, and extended context analysis.
Attention
Attention Structure
DeepSeek Sparse Attention
Attention Heads
128
Key-Value Heads
1
Attention Head Dimension
512
Position Embedding
ROPE
RoPE Theta
10,000
Sliding Window Attention
Yes
Sliding Window Size
128
Sliding Window Ratio
-
Linear Attention
No
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
7,168
Number of Layers
61
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
1
Tokenizer
Vocabulary Size
129,280
Mixture of Experts
Total Expert Parameters
87.8B
Number of Experts
384
Active Experts
6
Shared Experts
1
FFN Intermediate Size (per Expert)
3,072
Dense Layers Before MoE
0
DeepSeek-V4 is DeepSeek's latest generation of highly efficient Mixture-of-Experts language models, featuring a novel hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) that dramatically improves long-context efficiency. Pre-trained on 32T+ tokens with a comprehensive post-training pipeline including domain-specific expert cultivation and unified model consolidation. Both V4-Pro and V4-Flash support 1M context length as standard, with three reasoning effort modes (Non-think, Think High, Think Max). Released open-source under MIT license on April 24, 2026.
APX AI
Online