ApX 标志ApX 标志

趋近智

EmbeddingGemma 2

参数

2B

上下文长度

8K

模态

Text

架构

Dense

许可证

Apache 2.0

发布日期

14 Sept 2026

训练数据截止日期

-

API 价格 (每 1M)

仅支持自建托管

系统要求

不同量化方法和上下文大小的显存要求

1024 个令牌

5.46 GB VRAM

消费级

1x RTX 4090

24GB VRAM

数据中心

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

8192 个令牌

5.81 GB VRAM

消费级

1x RTX 4090

24GB VRAM

数据中心

1x NVIDIA A100

80GB VRAM

Apple Silicon

1x Apple M3 Max

128GB VRAM

架构图

Input TokensToken EmbeddingPosition: RoPEHidden: 512 · Context: 8K · Vocab: 262.1kx 24 layersRMSNormPre-AttentionSliding-Window Attention4Q / 2KV heads · SW: 512Head dim: 256+RMSNormPre-FFNFeed-Forward NetworkGated GELUIntermediate: 2k+Final RMSNormOutput Logits

评估基准

没有可用的 EmbeddingGemma 2 评估基准。

排名

排名

-

编程排名

-

关于 EmbeddingGemma 2

EmbeddingGemma 2 is a specialized text embedding model based on Google's Gemma 2 architecture. It is designed for high-performance semantic search, retrieval-augmented generation, and text similarity tasks.

技术规格

注意力

注意力结构

Single-Head Attention

注意力头

4

键值头

2

注意力头维度

256

位置嵌入

ROPE

RoPE Theta

10,000

滑动窗口注意力

Yes

滑动窗口大小

512

滑动窗口比例

83.3%

线性注意力

No

线性注意力比例

-

归一化

RMS Normalization

激活函数

Gated GELU

维度

辅助参数

-

隐藏维度大小

512

层数

24

FFN 中间层大小(稠密层)

2,048

多 Token 预测头数

-

分词器

词汇量大小

262,144

关于 Gemma 2

Gemma 2 是 Google 推出的开放大语言模型系列,提供 2B、9B 和 27B 三种参数规模。该系列基于 Gemma 架构构建,并引入了多项创新技术,包括交替式局部与全局注意力机制、旨在提升训练稳定性的 Logit 软截断 (logit soft-capping),以及用于优化推理效率的分组查询注意力 (Grouped Query Attention)。此外,较小规模的模型还采用了知识蒸馏技术。


其他 Gemma 2 模型