ApX logoApX logo

Llama 4 Maverick

Active Parameters

400B

Context Length

1M

Modality

Multimodal

Architecture

Mixture of Experts (MoE)

License

Llama 4 Community License Agreement

Release Date

5 Apr 2025

Knowledge Cutoff

Aug 2024

API Pricing (per 1M)

Input: $0.26 · Output: $0.91

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

842.03 GB VRAM

Consumer

47x RTX 4090

24GB VRAM

Datacenter

13x NVIDIA A100

80GB VRAM

Apple Silicon

10x Apple M3 Max

128GB VRAM

1,000,000 tokens

1357.60 GB VRAM

Consumer

83x RTX 4090

24GB VRAM

Datacenter

21x NVIDIA A100

80GB VRAM

Apple Silicon

17x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 12.3k · Context: 1M · Vocab: 202kx 120 layersRMSNormPre-AttentionGrouped-Query Attention96Q / 8KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (2/128 experts)Swish+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#144

BenchmarkScoreRank

Professional Knowledge

MMLU Pro

0.805

29

0.319

32

Software Engineering

SWE-bench Verified

0.21

37

Graduate-Level QA

GPQA

0.698

85

Agentic Index

Artificial Analysis

0.6

126

General Text

Text Arena

1327

132

0.16

136

Intelligence Index

Artificial Analysis

0.09

198

StackEval

Archived
ProLLM Stack Eval

0.923

7

QA Assistant

Archived
ProLLM QA Assistant

0.949

10

General Knowledge

Reference
MMLU

0.855

11

Coding

Archived
Aider Coding

0.16

20

Summarization

Archived
ProLLM Summarization

0.72

20

Rankings

Overall Rank

#144

Coding Rank

#119

About Llama 4 Maverick

Llama 4 Maverick is Meta's natively multimodal Mixture-of-Experts model with 400B total parameters activating 17B per token for production inference. It provides low-latency image and text understanding with early fusion across long contexts.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

96

Key-Value Heads

8

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

500,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Hidden Dimension Size

12,288

Number of Layers

120

FFN Intermediate Size (Dense)

8,192

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

202,048

Mixture of Experts

Total Expert Parameters

17.0B

Number of Experts

128

Active Experts

2

Shared Experts

-

FFN Intermediate Size (per Expert)

-

Dense Layers Before MoE

-

About Llama 4

Meta's Llama 4 model family implements a Mixture-of-Experts (MoE) architecture for efficient scaling. It features native multimodality through early fusion of text, images, and video. This iteration also supports significantly extended context lengths, with models capable of processing up to 10 million tokens.


Other Llama 4 Models