ApX logoApX logo

Llama 3.1 405B

Parameters

405B

Context Length

128K

Modality

Text

Architecture

Dense

License

Llama 3.1 Community License Agreement

Release Date

23 Jul 2024

Knowledge Cutoff

Dec 2023

API Pricing (per 1M)

Input: $4.00 · Output: $4.00

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

852.55 GB VRAM

Consumer

48x RTX 4090

24GB VRAM

Datacenter

13x NVIDIA A100

80GB VRAM

Apple Silicon

10x Apple M3 Max

128GB VRAM

128,000 tokens

921.36 GB VRAM

Consumer

53x RTX 4090

24GB VRAM

Datacenter

14x NVIDIA A100

80GB VRAM

Apple Silicon

11x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 16.4k · Context: 128Kx 126 layersRMSNormPre-AttentionGrouped-Query Attention128Q / 8KV headsHead dim: 128+RMSNormPre-FFNFeed-Forward NetworkSwiGLU+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#122

BenchmarkScoreRank

Professional Knowledge

MMLU Pro

0.733

30

Graduate-Level QA

GPQA

0.507

87

General Text

Text Arena

1335

112

Intelligence Index

Artificial Analysis

0.07

219

General Knowledge

Reference
MMLU

0.873

8

StackEval

Archived
ProLLM Stack Eval

0.8

14

QA Assistant

Archived
ProLLM QA Assistant

0.884

16

Rankings

Overall Rank

#122

Coding Rank

-

About Llama 3.1 405B

Llama 3.1 405B is Meta's flagship dense open foundation model pretrained on over 15T tokens across 16,000 GPUs. It delivers state-of-the-art steerability, synthetic data generation, and complex multi-step reasoning over a 128K context window.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

128

Key-Value Heads

8

Attention Head Dimension

-

Position Embedding

ROPE

RoPE Theta

-

Sliding Window Attention

-

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

16,384

Number of Layers

126

FFN Intermediate Size (Dense)

-

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

-

About Llama 3.1

Llama 3.1 is Meta's advanced large language model family, building upon Llama 3. It features an optimized decoder-only transformer architecture, available in 8B, 70B, and 405B parameter versions. Significant enhancements include an expanded 128K token context window and improved multilingual capabilities across eight languages, refined through data and post-training procedures.


Other Llama 3.1 Models