Active Parameters
2T
Context Length
10M
Modality
Multimodal
Architecture
Mixture of Experts (MoE)
License
Llama 4 Community License Agreement
Release Date
-
Knowledge Cutoff
Aug 2024
API Pricing (per 1M)
Self-hosted only
VRAM requirements for different quantization methods and context sizes
1,024 tokens
Consumer
336x RTX 4090
24GB VRAM
Datacenter
76x NVIDIA A100
80GB VRAM
Apple Silicon
76x Apple M3 Max
128GB VRAM
10,000,000 tokens
Consumer
1288x RTX 4090
24GB VRAM
Datacenter
248x NVIDIA A100
80GB VRAM
Apple Silicon
316x Apple M3 Max
128GB VRAM
No evaluation benchmarks for Llama 4 Behemoth available.
Overall Rank
-
Coding Rank
-
Llama 4 Behemoth is Meta's 2T-parameter multimodal Mixture-of-Experts research foundation model activating 288B parameters per token. Engineered with native early fusion, it serves as the high-capacity knowledge source for downstream distillation.
Attention
Attention Structure
Grouped-Query Attention
Attention Heads
128
Key-Value Heads
8
Attention Head Dimension
-
Position Embedding
Absolute Position Embedding
RoPE Theta
-
Sliding Window Attention
-
Sliding Window Size
-
Sliding Window Ratio
-
Linear Attention
-
Linear Attention Ratio
-
Normalization
RMS Normalization
Activation Function
SwigLU
Dimensions
Hidden Dimension Size
16,384
Number of Layers
160
FFN Intermediate Size (Dense)
-
Multi-Token Prediction Heads
-
Tokenizer
Vocabulary Size
-
Mixture of Experts
Total Expert Parameters
288.0B
Number of Experts
16
Active Experts
2
Shared Experts
-
FFN Intermediate Size (per Expert)
-
Dense Layers Before MoE
-
Meta's Llama 4 model family implements a Mixture-of-Experts (MoE) architecture for efficient scaling. It features native multimodality through early fusion of text, images, and video. This iteration also supports significantly extended context lengths, with models capable of processing up to 10 million tokens.
APX AI
Online