ApX logoApX logo

Gemma 4 12B

Parameters

11.95B

Context Length

262K

Modality

Multimodal

Auxiliary Parameters

550M

Architecture

Dense

License

Apache-2.0

Release Date

3 Jun 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Input: $0.10 · Output: $0.30

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

25.27 GB VRAM

Consumer

2x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

262,144 tokens

125.67 GB VRAM

Consumer

6x RTX 4090

24GB VRAM

Datacenter

2x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 3.8k · Context: 262K · Vocab: 262.1kx 48 layersRMSNormPre-AttentionMulti-Head Attention16Q / 8KV heads · SW: 1kHead dim: 256+RMSNormPre-FFNFeed-Forward NetworkGELUIntermediate: 15.4k+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#109

BenchmarkScoreRank

Professional Knowledge

MMLU Pro

0.772

32

Graduate-Level QA

GPQA

0.788

68

auto

0.31

124

Intelligence Index

Artificial Analysis
auto

0.14

Standard

0.09

213

256

Rankings

Overall Rank

#109

Coding Rank

#116

About Gemma 4 12B

Google DeepMind's 12B dense open-weights model released June 3, 2026, bridging the gap between the edge-friendly E4B and the more advanced 26B MoE. Uniquely features an encoder-free unified architecture that projects raw image patches and audio waveforms directly into the LLM embedding space through lightweight linear layers, eliminating the latency and memory overhead of separate encoders. Supports 256K token context, native text/image/audio inputs, configurable thinking mode, and runs on consumer laptops with 16GB of RAM.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

16

Key-Value Heads

8

Attention Head Dimension

256

Position Embedding

Absolute Position Embedding

RoPE Theta

10,000

Sliding Window Attention

Yes

Sliding Window Size

1,024

Sliding Window Ratio

83.3%

Linear Attention

No

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

GELU

Dimensions

Auxiliary Parameters

550M

Hidden Dimension Size

3,840

Number of Layers

48

FFN Intermediate Size (Dense)

15,360

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

262,144

About Gemma 4

Gemma 4 is Google DeepMind's most advanced open model family, built from Gemini 3 research and technology. Featuring both Dense and Mixture-of-Experts (MoE) architectures, these multimodal models handle text, images, and audio (on smaller variants), with context windows up to 256K tokens. Designed for frontier-level performance across reasoning, coding, and agentic workflows, Gemma 4 delivers unprecedented intelligence-per-parameter from mobile devices to enterprise servers. Released under Apache 2.0 license.


Other Gemma 4 Models
Gemma 4 12B - VRAM, Specs & Benchmarks