ApX logoApX logo

Llama 4 Maverick

Active Parameters

400B

Context Length

1M

Modality

Multimodal

Auxiliary Parameters

850M

Architecture

Mixture of Experts (MoE)

License

Llama 4 Community License Agreement

Release Date

5 Apr 2025

Knowledge Cutoff

Aug 2024

API Pricing (per 1M)

Input: $0.26 · Output: $0.91

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

784.30 GB VRAM

Consumer

44x RTX 4090

24GB VRAM

Datacenter

12x NVIDIA A100

80GB VRAM

Apple Silicon

9x Apple M3 Max

128GB VRAM

1,000,000 tokens

1264.46 GB VRAM

Consumer

76x RTX 4090

24GB VRAM

Datacenter

20x NVIDIA A100

80GB VRAM

Apple Silicon

16x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 12.3k · Context: 1M · Vocab: 202kx 120 layersRMSNormPre-AttentionGrouped-Query Attention96Q / 8KV headsHead dim: 128+RMSNormPre-FFNSparse MoE FFN (2/128 experts)Swish+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#152

BenchmarkScoreRank

Professional Knowledge

MMLU Pro

0.805

28

0.319

32

Software Engineering

SWE-bench Verified

0.21

38

Graduate-Level QA

GPQA

0.698

85

Agentic Index

Artificial Analysis

0.6

135

0.16

145

General Text

Text Arena

1327

151

Intelligence Index

Artificial Analysis

0.10

221

StackEval

Archived
ProLLM Stack Eval

0.923

7

QA Assistant

Archived
ProLLM QA Assistant

0.949

10

General Knowledge

Reference
MMLU

0.855

11

Coding

Archived
Aider Coding

0.16

20

Summarization

Archived
ProLLM Summarization

0.72

20

Rankings

Overall Rank

#152

Coding Rank

#133

About Llama 4 Maverick

Llama 4 Maverick is Meta's natively multimodal Mixture-of-Experts model with 400B total parameters activating 17B per token for production inference. It provides low-latency image and text understanding with early fusion across long contexts.

Technical Specifications

Attention

Attention Structure

Grouped-Query Attention

Attention Heads

96

Key-Value Heads

8

Attention Head Dimension

128

Position Embedding

ROPE

RoPE Theta

500,000

Sliding Window Attention

No

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Swish

Dimensions

Auxiliary Parameters

850M

Hidden Dimension Size

12,288

Number of Layers

120

FFN Intermediate Size (Dense)

8,192

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

202,048

Mixture of Experts

Total Expert Parameters

17.0B

Number of Experts

128

Active Experts

2

Shared Experts

-

FFN Intermediate Size (per Expert)

-

Dense Layers Before MoE

-

About Llama 4

Meta's Llama 4 model family implements a Mixture-of-Experts (MoE) architecture for efficient scaling. It features native multimodality through early fusion of text, images, and video. This iteration also supports significantly extended context lengths, with models capable of processing up to 10 million tokens.


Other Llama 4 Models