ApX logoApX logo

Llama 4 Behemoth

Active Parameters

2T

Context Length

10M

Modality

Multimodal

Architecture

Mixture of Experts (MoE)

License

Llama 4 Community License Agreement

Release Date

-

Knowledge Cutoff

Aug 2024

API Pricing (per 1M)

Self-hosted only

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

4202.20 GB VRAM

Consumer

336x RTX 4090

24GB VRAM

Datacenter

76x NVIDIA A100

80GB VRAM

Apple Silicon

76x Apple M3 Max

128GB VRAM

10,000,000 tokens

11082.78 GB VRAM

Consumer

1288x RTX 4090

24GB VRAM

Datacenter

248x NVIDIA A100

80GB VRAM

Apple Silicon

316x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 16.4k · Context: 10Mx 160 layersRMSNormPre-AttentionGrouped-Query Attention128Q / 8KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (2/16 experts)SwiGLU+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for Llama 4 Behemoth available.

Rankings

Overall Rank

-

Coding Rank

-

About Llama 4 Behemoth

Llama 4 Behemoth is Meta's 2T-parameter multimodal Mixture-of-Experts research foundation model activating 288B parameters per token. Engineered with native early fusion, it serves as the high-capacity knowledge source for downstream distillation.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

128

Key-Value Heads

8

Attention Head Dimension

-

Position Embedding

Absolute Position Embedding

RoPE Theta

-

Sliding Window Attention

-

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

SwigLU

Dimensions

Hidden Dimension Size

16,384

Number of Layers

160

FFN Intermediate Size (Dense)

-

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

-

Mixture of Experts

Total Expert Parameters

288.0B

Number of Experts

16

Active Experts

2

Shared Experts

-

FFN Intermediate Size (per Expert)

-

Dense Layers Before MoE

-

About Llama 4

Meta's Llama 4 model family implements a Mixture-of-Experts (MoE) architecture for efficient scaling. It features native multimodality through early fusion of text, images, and video. This iteration also supports significantly extended context lengths, with models capable of processing up to 10 million tokens.


Other Llama 4 Models
Llama 4 Behemoth: Specifications and GPU VRAM Requirements