Parameters
70B
Context Length
128K
Modality
Text
Architecture
Dense
License
Llama 3.1 Community License Agreement
Release Date
23 Jul 2024
Knowledge Cutoff
Dec 2023
API Pricing (per 1M)
Input: $0.56 · Output: $0.56
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
7x RTX 4090
24GB VRAM
Datacenter
2x NVIDIA A100
80GB VRAM
Apple Silicon
2x Apple M3 Max
128GB VRAM
128,000 tokens
Consumer
10x RTX 4090
24GB VRAM
Datacenter
3x NVIDIA A100
80GB VRAM
Apple Silicon
2x Apple M3 Max
128GB VRAM
Rank
#158
| Benchmark | Score | Rank |
|---|---|---|
Professional Knowledge | 0.664 | 48 |
Graduate-Level QA | 0.417 | 115 |
General Text | 1293 | 145 |
Intelligence Index | 0.07 | 224 |
General Knowledge Reference | 0.836 | 16 |
Summarization Archived | 0.598 | 23 |
Overall Rank
#158
Coding Rank
-
Llama 3.1 70B is a high-performance open language model from Meta built for advanced coding assistance, multilingual conversation, and enterprise RAG. It supports long-sequence comprehension and tool integration across a 128K context window.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
64
Key-Value Heads
8
Attention Head Dimension
-
Position Embedding
ROPE
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
-
Activation Function
-
Dimensions
Auxiliary Parameters
-
Hidden Dimension Size
8,192
Number of Layers
80
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
-
Llama 3.1 is Meta's advanced large language model family, building upon Llama 3. It features an optimized decoder-only transformer architecture, available in 8B, 70B, and 405B parameter versions. Significant enhancements include an expanded 128K token context window and improved multilingual capabilities across eight languages, refined through data and post-training procedures.
Assistant
Online