ApX 标志ApX 标志

趋近智

DeepSeek V4 Flash (0731)

活跃参数

284B

上下文长度

1.31M

模态

Reasoning

架构

Mixture of Experts (MoE)

许可证

MIT License

发布日期

31 Jul 2026

训练数据截止日期

-

API 价格 (每 1M)

输入: $0.44 · 输出: $1.32

系统要求

不同量化方法和上下文大小的显存要求

1024 个令牌

597.99 GB VRAM

消费级

32x RTX 4090

24GB VRAM

数据中心

9x NVIDIA A100

80GB VRAM

Apple Silicon

7x Apple M3 Max

128GB VRAM

1310720 个令牌

719.10 GB VRAM

消费级

40x RTX 4090

24GB VRAM

数据中心

11x NVIDIA A100

80GB VRAM

Apple Silicon

8x Apple M3 Max

128GB VRAM

架构图

Input TokensToken EmbeddingPosition: RoPEHidden: 4.1k · Context: 1.31M · Vocab: 129.3kx 43 layersRMSNormPre-AttentionDeepSeek Sparse Attention64Q / 1KV heads · SW: 128Head dim: 512+RMSNormPre-FFNSparse MoE FFN (6/256 experts)SwiGLUIntermediate: 2k+Final RMSNormOutput Logits

评估基准

排名

#31

基准分数排名

0.79

8

0.87

23

Agent Arena

Agent Arena
high

0.02

24

0.74

30

LiveBench 平均分

LiveBench Average

0.74

30

智能体指数

Artificial Analysis
max

0.42

35

0.75

35

智能编程

LiveBench Agentic

0.47

35

0.87

35

max

0.69

49

max

0.34

51

排名

排名

#31

编程排名

#34

关于 DeepSeek V4 Flash (0731)

DeepSeek V4 Flash 0731 是 DeepSeek 推出的一款 284B 稀疏混合专家模型,每个 token 包含 13B 激活参数。该模型是一个经过重新后训练的修订版本,专为快速编程、推理和多轮智能体工作流而设计。

技术规格

注意力

注意力结构

DeepSeek Sparse Attention

注意力头

64

键值头

1

注意力头维度

512

位置嵌入

ROPE

RoPE Theta

10,000

滑动窗口注意力

Yes

滑动窗口大小

128

滑动窗口比例

-

线性注意力

No

线性注意力比例

-

归一化

RMS Normalization

激活函数

SwigLU

维度

隐藏维度大小

4,096

层数

43

FFN 中间层大小(稠密层)

-

多 Token 预测头数

1

分词器

词汇量大小

129,280

混合专家

专家参数总数

13.0B

专家数量

256

活跃专家

6

共享专家数

1

FFN 中间层大小(每专家)

2,048

MoE 前的稠密层数

-

关于 DeepSeek V4

DeepSeek-V4 is DeepSeek's latest generation of highly efficient Mixture-of-Experts language models, featuring a novel hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) that dramatically improves long-context efficiency. Pre-trained on 32T+ tokens with a comprehensive post-training pipeline including domain-specific expert cultivation and unified model consolidation. Both V4-Pro and V4-Flash support 1M context length as standard, with three reasoning effort modes (Non-think, Think High, Think Max). Released open-source under MIT license on April 24, 2026.


其他 DeepSeek V4 模型