ApX logoApX logo

EmbeddingGemma 2

Parameters

2B

Context Length

8K

Modality

Text

Architecture

Dense

License

Apache 2.0

Release Date

14 Sept 2026

Knowledge Cutoff

-

API Pricing (per 1M)

Self-hosted only

System Requirements

VRAM requirements for different quantization methods and context sizes

1,024 tokens

5.46 GB VRAM

Consumer

1x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

8,192 tokens

5.81 GB VRAM

Consumer

1x RTX 4090

24GB VRAM

Datacenter

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

Architecture Diagram

Input TokensToken EmbeddingPosition: RoPEHidden: 512 · Context: 8K · Vocab: 262.1kx 24 layersRMSNormPre-AttentionSliding-Window Attention4Q / 2KV heads · SW: 512Head dim: 256+RMSNormPre-FFNFeed-Forward NetworkGated GELUIntermediate: 2k+Final RMSNormOutput Logits

Evaluation Benchmarks

No evaluation benchmarks for EmbeddingGemma 2 available.

Rankings

Overall Rank

-

Coding Rank

-

About EmbeddingGemma 2

EmbeddingGemma 2 is a specialized text embedding model based on Google's Gemma 2 architecture. It is designed for high-performance semantic search, retrieval-augmented generation, and text similarity tasks.

Technical Specifications

Attention

Attention Structure

Single-Head Attention

Attention Heads

4

Key-Value Heads

2

Attention Head Dimension

256

Position Embedding

ROPE

RoPE Theta

10,000

Sliding Window Attention

Yes

Sliding Window Size

512

Sliding Window Ratio

83.3%

Linear Attention

No

Linear Attention Ratio

-

Normalization

RMS Normalization

Activation Function

Gated GELU

Dimensions

Auxiliary Parameters

-

Hidden Dimension Size

512

Number of Layers

24

FFN Intermediate Size (Dense)

2,048

Multi-Token Prediction Heads

-

Tokenizer

Vocabulary Size

262,144

About Gemma 2

Gemma 2 is Google's family of open large language models, offering 2B, 9B, and 27B parameter sizes. Built upon the Gemma architecture, it incorporates innovations such as interleaved local and global attention, logit soft-capping for training stability, and Grouped Query Attention for inference efficiency. The smaller models leverage knowledge distillation.


Other Gemma 2 Models
EmbeddingGemma 2 - VRAM, Specs & Benchmarks