ApX logoApX logo

Gemma 3n E2B IT

Active Parameters

6B

Context Length

33K

Modality

Text

Architecture

Mixture of Experts (MoE)

License

Google Gemma License

Release Date

20 May 2025

Knowledge Cutoff

Jun 2024

API Pricing (per 1M)

Input: $0.00 · Output: $0.00

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

14.23 GB VRAM

Consumer

1x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

32,768 tokens

18.33 GB VRAM

Consumer

1x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: AbsoluteHidden: 2.6k · Context: 33Kx 30 layersRMSNormPre-AttentionMulti-Head Attention+RMSNormPre-FFNSparse MoE FFNActivation+Final RMSNormOutput Logits

Evaluation Benchmarks

Rank

#209

BenchmarkScoreRank

Professional Knowledge

MMLU Pro

0.405

60

Graduate-Level QA

GPQA

0.248

124

Intelligence Index

Artificial Analysis

0.05

301

General Knowledge

Reference
MMLU

0.601

32

Rankings

Overall Rank

#209

Coding Rank

-

About Gemma 3n E2B IT

Gemma 3n E2B IT is an edge-optimized multimodal model from Google built on the Matryoshka Transformer architecture for on-device deployment. It processes text, image, and audio inputs with native function calling over a 32K context window.

Technical Specifications

Attention

Attention Structure

Multi-Head Attention

Attention Heads

-

Key-Value Heads

-

Attention Head Dimension

-

Position Embedding

Absolute Position Embedding

RoPE Theta

-

Sliding Window Attention

-

Sliding Window Size

-

Sliding Window Ratio

-

Linear Attention

-

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

-

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

2,560

Number of Layers

30

FFN Intermediate Size (Dense)

-

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

-

Mixture of Experts

Total Expert Parameters

2.0B

Number of Experts

-

Active Experts

-

Shared Experts

-

FFN Intermediate Size (per Expert)

-

Dense Layers Before MoE

-

About Gemma 3

Gemma 3 is a family of open, lightweight models from Google. It introduces multimodal image and text processing, supports over 140 languages, and features extended context windows up to 128K tokens. Models are available in multiple parameter sizes for diverse applications.


Other Gemma 3 Models