ApX logoApX logo

Llama 4 Maverick

Active Parameters

400B

Context Length

1M

Modality

Multimodal

Architecture

Mixture of Experts (MoE)

License

Llama 4 Community License Agreement

Release Date

5 Apr 2025

Knowledge Cutoff

Aug 2024

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

842.03 GB VRAM

Consumer

47x RTX 4090

24GB VRAM

Datacenter

13x NVIDIA A100

80GB VRAM

Apple Silicon

10x Apple M3 Max

128GB VRAM

1,000,000 tokens

1357.60 GB VRAM

Consumer

83x RTX 4090

24GB VRAM

Datacenter

21x NVIDIA A100

80GB VRAM

Apple Silicon

17x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingHidden: 12.3k · Context: 1M · Vocab: 202kx 120 layersRMSNormPre-AttentionGrouped-Query Attention96Q / 8KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (2/128 experts)Swish+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#120

BenchmarkScoreRank

0.949

10

General Knowledge

MMLU

0.855

12

0.72

21

0.319

30

0.16

31

Professional Knowledge

MMLU Pro

0.79

39

General Text

Text Arena

1327

128

Rankings

Overall Rank

#120

Coding Rank

#99

About Llama 4 Maverick

The Llama 4 Maverick model is a natively multimodal large language model developed by Meta, released as part of the Llama 4 model family. Its primary purpose is to deliver advanced capabilities in text and image understanding, supporting a wide range of applications including assistant-like conversational AI, creative content generation, complex reasoning, and code generation. Designed for both commercial and research deployment, Llama 4 Maverick aims to provide high-quality performance with improved cost efficiency.

From an architectural perspective, Llama 4 Maverick leverages a Mixture-of-Experts (MoE) design, a significant departure from previous dense transformer models. It comprises 400 billion total parameters, with only 17 billion parameters actively engaged per token during inference. This efficiency is achieved through the use of 128 experts, where processing involves alternating dense and MoE layers. The model integrates different modalities, such as text and images, through an early fusion mechanism, allowing for comprehensive multimodal processing from the initial stages. The internal architecture also incorporates iRoPE for managing and scaling context, further enhancing its capabilities.

Llama 4 Maverick demonstrates robust performance across diverse benchmarks, including coding, reasoning, and multilingual tasks, as well as long-context processing and image understanding. It is engineered for high model throughput and is suitable for production environments that demand low latency and precision. The model's design facilitates its deployment in scenarios requiring sophisticated multimodal interaction and efficient resource utilization, addressing modern AI application requirements.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

96

Key-Value Heads

8

Attention Head Dimension

128

Position Embedding

Irope

RoPE Theta

500,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Hidden Dimension Size

12,288

Number of Layers

120

FFN Intermediate Size (Dense)

8,192

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

202,048

Mixture of Experts

Total Expert Parameters

17.0B

Number of Experts

128

Active Experts

2

Shared Experts

-

FFN Intermediate Size (per Expert)

-

Dense Layers Before MoE

-

About Llama 4

Meta's Llama 4 model family implements a Mixture-of-Experts (MoE) architecture for efficient scaling. It features native multimodality through early fusion of text, images, and video. This iteration also supports significantly extended context lengths, with models capable of processing up to 10 million tokens.


Other Llama 4 Models