ApX logoApX logo

DeepSeek-V3 671B

Active Parameters

671B

Context Length

131K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

DeepSeek Model License

Release Date

27 Dec 2024

Knowledge Cutoff

Jul 2024

API Pricing (per 1M)

Input: $0.32 · Output: $0.89

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

1317.83 GB VRAM

Consumer

80x RTX 4090

24GB VRAM

Datacenter

21x NVIDIA A100

80GB VRAM

Apple Silicon

17x Apple M3 Max

128GB VRAM

131,072 tokens

1826.23 GB VRAM

Consumer

118x RTX 4090

24GB VRAM

Datacenter

29x NVIDIA A100

80GB VRAM

Apple Silicon

25x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 7.2k · Context: 131K · Vocab: 129.3kx 61 layersRMSNormPre-AttentionMulti-Layer Attention128Q / 128KV headsHead dim: 56+RMSNormPre-FFNSparse MoE FFN (9/257 experts)SwishIntermediate: 2k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#138

BenchmarkScoreRank

0.439

28

Software Engineering

SWE-bench Verified

0.42

33

Professional Knowledge

MMLU Pro

0.759

34

Graduate-Level QA

GPQA

0.684

87

General Text

Text Arena

1396

115

0.23

131

Agentic Index

Artificial Analysis

0.8

131

Intelligence Index

Artificial Analysis

0.10

228

StackEval

Archived
ProLLM Stack Eval

0.976

🥈

2

General Knowledge

Reference
MMLU

0.885

6

QA Assistant

Archived
ProLLM QA Assistant

0.953

9

Coding

Archived
Aider Coding

0.55

11

Summarization

Archived
ProLLM Summarization

0.806

11

Rankings

Overall Rank

#138

Coding Rank

#112

About DeepSeek-V3 671B

DeepSeek-V3 is a 671B-parameter Mixture-of-Experts foundation model activating 37B parameters per token for efficient general language processing. Utilizing Multi-Head Latent Attention and Multi-Token Prediction, it excels across coding, math, and 128K long contexts.

Technical Specifications

Attention

Attention Structure

Multi-Layer Attention

Attention Heads

128

Key-Value Heads

128

Attention Head Dimension

-

Position Embedding

ROPE

RoPE Theta

10,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

7,168

Number of Layers

61

FFN Intermediate Size (Dense)

2,048

Multi-Token Prediction Heads

1

Tokenizer

Vocabulary Size

129,280

Mixture of Experts

Total Expert Parameters

37.0B

Number of Experts

257

Active Experts

9

Shared Experts

1

FFN Intermediate Size (per Expert)

2,048

Dense Layers Before MoE

3

About DeepSeek-V3

DeepSeek-V3 is a Mixture-of-Experts (MoE) language model comprising 671B parameters with 37B activated per token. Its architecture incorporates Multi-head Latent Attention and DeepSeekMoE for efficient inference and training. Innovations include an auxiliary-loss-free load balancing strategy and a multi-token prediction objective, trained on 14.8T tokens.


Other DeepSeek-V3 Models