趋近智
参数
2B
上下文长度
8K
模态
Text
架构
Dense
许可证
Apache 2.0
发布日期
14 Sept 2026
训练数据截止日期
-
API 价格 (每 1M)
仅支持自建托管
不同量化方法和上下文大小的显存要求
1024 个令牌
消费级
1x RTX 4090
24GB VRAM
数据中心
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
8192 个令牌
消费级
1x RTX 4090
24GB VRAM
数据中心
1x NVIDIA A100
80GB VRAM
Apple Silicon
1x Apple M3 Max
128GB VRAM
没有可用的 EmbeddingGemma 2 评估基准。
排名
-
编程排名
-
EmbeddingGemma 2 is a specialized text embedding model based on Google's Gemma 2 architecture. It is designed for high-performance semantic search, retrieval-augmented generation, and text similarity tasks.
注意力
注意力结构
Single-Head Attention
注意力头
4
键值头
2
注意力头维度
256
位置嵌入
ROPE
RoPE Theta
10,000
滑动窗口注意力
Yes
滑动窗口大小
512
滑动窗口比例
83.3%
线性注意力
No
线性注意力比例
-
归一化
RMS Normalization
激活函数
Gated GELU
维度
辅助参数
-
隐藏维度大小
512
层数
24
FFN 中间层大小(稠密层)
2,048
多 Token 预测头数
-
分词器
词汇量大小
262,144
Gemma 2 是 Google 推出的开放大语言模型系列,提供 2B、9B 和 27B 三种参数规模。该系列基于 Gemma 架构构建,并引入了多项创新技术,包括交替式局部与全局注意力机制、旨在提升训练稳定性的 Logit 软截断 (logit soft-capping),以及用于优化推理效率的分组查询注意力 (Grouped Query Attention)。此外,较小规模的模型还采用了知识蒸馏技术。
Assistant
在线