ApX logoApX logo

DeepSeek-R1 70B

Parameters

70B

Context Length

33K

Modality

Text

Architecture

Dense

License

MIT License

Release Date

27 Dec 2024

Knowledge Cutoff

-

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

153.43 GB VRAM

Consumer

8x RTX 4090

24GB VRAM

Datacenter

3x NVIDIA A100

80GB VRAM

Apple Silicon

2x Apple M3 Max

128GB VRAM

32,768 tokens

306.34 GB VRAM

Consumer

16x RTX 4090

24GB VRAM

Datacenter

5x NVIDIA A100

80GB VRAM

Apple Silicon

3x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 8.2k · Context: 33K · Vocab: 128.3kx 80 layersRMSNormPre-AttentionMulti-Layer Attention112Q / 112KV headsHead dim: 128+RMSNormPre-FFNFeed-Forward NetworkSwishIntermediate: 28.7k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for DeepSeek-R1 70B available.

Rankings

Overall Rank

-

Coding Rank

-

About DeepSeek-R1 70B

DeepSeek-R1 is a family of advanced large language models developed by DeepSeek, designed with a primary focus on enhancing reasoning capabilities. The DeepSeek-R1-Distill-Llama-70B variant is a product of knowledge distillation, leveraging the reasoning strengths of the larger DeepSeek-R1 model and transferring them to a Llama-3.3-70B-Instruct base architecture. This distillation process aims to create a highly capable model that maintains the efficiency and operational characteristics of its base while inheriting sophisticated reasoning patterns.

Architecturally, DeepSeek-R1-Distill-Llama-70B is a dense transformer model, distinguishing it from the Mixture of Experts (MoE) architecture of the original DeepSeek-R1. It employs a Multi-Head Attention (MLA) mechanism with 112 attention heads, facilitating comprehensive processing of input sequences. The model integrates Rotary Position Embeddings (RoPE) for effective handling of positional information within sequences and utilizes Flash Attention for optimized computational efficiency. This configuration enables the model to process substantial context lengths, supporting complex problem-solving.

This model is engineered for general text generation, code generation, and sophisticated problem-solving across domains requiring logical inference and multi-step reasoning. Its design prioritizes efficient deployment, making it suitable for applications where computational resources are a consideration, including those on consumer-grade hardware. The DeepSeek-R1-Distill-Llama-70B is particularly adept at tasks demanding structured thought processes, such as mathematical problem-solving and generating coherent code, extending its utility across various technical and research applications.

Technical Specifications

Attention

Attention Structure

Multi-Layer Attention

Attention Heads

112

Key-Value Heads

112

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

500,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Hidden Dimension Size

8,192

Number of Layers

80

FFN Intermediate Size (Dense)

28,672

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

128,256

About DeepSeek-R1

DeepSeek-R1 is a model family developed for logical reasoning tasks. It incorporates a Mixture-of-Experts architecture for computational efficiency and scalability. The family utilizes Multi-Head Latent Attention and employs reinforcement learning in its training, with some variants integrating cold-start data.


Other DeepSeek-R1 Models