Parameters
405B
Context Length
128K
Modality
Text
Architecture
Dense
License
Llama 3.1 Community License Agreement
Release Date
23 Jul 2024
Knowledge Cutoff
Dec 2023
API Pricing (per 1M)
Input: $4.00 · Output: $4.00
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
48x RTX 4090
24GB VRAM
Datacenter
13x NVIDIA A100
80GB VRAM
Apple Silicon
10x Apple M3 Max
128GB VRAM
128,000 tokens
Consumer
53x RTX 4090
24GB VRAM
Datacenter
14x NVIDIA A100
80GB VRAM
Apple Silicon
11x Apple M3 Max
128GB VRAM
Rank
#122
| Benchmark | Score | Rank |
|---|---|---|
Professional Knowledge | 0.733 | 30 |
Graduate-Level QA | 0.507 | 87 |
General Text | 1335 | 112 |
Intelligence Index | 0.07 | 219 |
General Knowledge Reference | 0.873 | 8 |
StackEval Archived | 0.8 | 14 |
QA Assistant Archived | 0.884 | 16 |
Overall Rank
#122
Coding Rank
-
Llama 3.1 405B is Meta's flagship dense open foundation model pretrained on over 15T tokens across 16,000 GPUs. It delivers state-of-the-art steerability, synthetic data generation, and complex multi-step reasoning over a 128K context window.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
128
Key-Value Heads
8
Attention Head Dimension
-
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
16,384
Number of Layers
126
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
-
Llama 3.1 is Meta's advanced large language model family, building upon Llama 3. It features an optimized decoder-only transformer architecture, available in 8B, 70B, and 405B parameter versions. Significant enhancements include an expanded 128K token context window and improved multilingual capabilities across eight languages, refined through data and post-training procedures.
APX AI
Online