ApX logoApX logo

Llama 3.1 405B

Parameters

405B

Context Length

128K

Modality

Text

Architecture

Dense

License

Llama 3.1 Community License Agreement

Release Date

23 Jul 2024

Knowledge Cutoff

Dec 2023

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

852.55 GB VRAM

Consumer

48x RTX 4090

24GB VRAM

Datacenter

13x NVIDIA A100

80GB VRAM

Apple Silicon

10x Apple M3 Max

128GB VRAM

128,000 tokens

921.36 GB VRAM

Consumer

53x RTX 4090

24GB VRAM

Datacenter

14x NVIDIA A100

80GB VRAM

Apple Silicon

11x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 16.4k · Context: 128Kx 126 layersRMSNormPre-AttentionGrouped-Query Attention128Q / 8KV headsHead dim: 128+RMSNormPre-FFNFeed-Forward NetworkSwiGLU+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#113

BenchmarkScoreRank

General Knowledge

MMLU

0.873

10

0.884

16

Professional Knowledge

MMLU Pro

0.73

61

Web Development

WebDev Arena

1335

65

General Text

Text Arena

1334

73

Rankings

Overall Rank

#113

Coding Rank

#74

About Llama 3.1 405B

Meta Llama 3.1 405B is the largest generative AI model within the Llama 3.1 collection, which also includes 8B and 70B parameter variants. This model is engineered to serve a broad spectrum of commercial and research applications, focusing on multilingual dialogue and advanced text generation. It is designed to expand the accessibility of sophisticated AI capabilities, supporting a comprehensive set of eight languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.

Architecturally, Llama 3.1 405B employs an optimized decoder-only Transformer. A notable innovation in its structure is the integration of Grouped-Query Attention (GQA), which is implemented to enhance inference scalability. The model was trained on an extensive dataset exceeding 15 trillion tokens, leveraging a substantial computational infrastructure of over 16,000 H100 GPUs. This training scale is distinct for a Llama model. Post-training refinement involves multiple iterative rounds of Supervised Fine-Tuning (SFT), Rejection Sampling (RS), and Direct Preference Optimization (DPO) to align the model's responses. The internal mechanisms feature Rotary Positional Embedding (RoPE) for positional encoding and Root Mean Square Normalization (RMSNorm) for internal state normalization. Its activation function is SwiGLU. The architecture prioritizes training stability and scalability by intentionally not incorporating a Mixture-of-Experts (MoE) design.

Functionally, Llama 3.1 405B offers a significantly expanded context length of 128,000 tokens, enabling processing of extended textual inputs. It demonstrates advanced capabilities in various domains, including general knowledge comprehension, steerability, mathematical problem-solving, and the use of external tools. Practical applications include long-form text summarization, development of multilingual conversational agents, and assistance with coding tasks. Additionally, the model is designed to facilitate advanced workflows such as generating synthetic data to enhance the training of smaller models and supporting model distillation processes. Its substantial parameter count contributes to its capacity for generating detailed and contextually relevant text.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

128

Key-Value Heads

8

Attention Head Dimension

-

Position Embedding

ROPE

RoPE Theta

-

Sliding Window Attention

-

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

16,384

Number of Layers

126

FFN Intermediate Size (Dense)

-

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

-

Model Integrity

Total Score

B+

72 / 100

About Llama 3.1

Llama 3.1 is Meta's advanced large language model family, building upon Llama 3. It features an optimized decoder-only transformer architecture, available in 8B, 70B, and 405B parameter versions. Significant enhancements include an expanded 128K token context window and improved multilingual capabilities across eight languages, refined through data and post-training procedures.


Other Llama 3.1 Models
Llama 3.1 405B: Specifications and GPU VRAM Requirements