Active Parameters
2T
Context Length
10M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Llama 4 Community License Agreement
Release Date
-
Knowledge Cutoff
Aug 2024
API Pricing (per 1M)
Self-hosted only
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
306x RTX 4090
24GB VRAM
Datacenter
70x NVIDIA A100
80GB VRAM
Apple Silicon
69x Apple M3 Max
128GB VRAM
10,000,000 tokens
Consumer
1161x RTX 4090
24GB VRAM
Datacenter
226x NVIDIA A100
80GB VRAM
Apple Silicon
283x Apple M3 Max
128GB VRAM
No evaluation benchmarks for Llama 4 Behemoth available.
Overall Rank
-
Coding Rank
-
Llama 4 Behemoth is Meta's 2T-parameter multimodal Mixture-of-Experts research foundation model activating 288B parameters per token. Engineered with native early fusion, it serves as the high-capacity knowledge source for downstream distillation.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
128
Key-Value Heads
8
Attention Head Dimension
-
Position Embedding
Absolute Position Embedding
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Auxiliary Parameters
-
Hidden Dimension Size
16,384
Number of Layers
160
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
-
Mixture of Experts
Total Expert Parameters
288.0B
Number of Experts
16
Active Experts
2
Shared Experts
-
FFN Intermediate Size (per Expert)
-
Dense Layers Before MoE
-
Meta's Llama 4 model family implements a Mixture-of-Experts (MoE) architecture for efficient scaling. It features native multimodality through early fusion of text, images, and video. This iteration also supports significantly extended context lengths, with models capable of processing up to 10 million tokens.
Assistant
Online