EmbeddingGemma 2:Google 把文字、图片、声音、视频放进同一个向量空间

EmbeddingGemma 2: Google Puts Text, Image, Audio, and Video Into a Single Embedding Space

Tech-Experiment #嵌入模型#多模态#Google#端侧AI#RAG#开源模型
🇨🇳 中文

大多数语义搜索系统都有同一个问题:文字用一个 embedding 模型,图片用另一个,音频又是另一个,跨模态检索要手动对齐。你有一段录音,想用一句话去搜相关片段——通常需要先把音频转文字,再嵌入文字,才能和文字库比较。能直接用的媒介就是不一样的。

EmbeddingGemma 2 把这个问题在架构上消掉了:一个模型,四种模态,一个向量空间。

HuggingFace: https://huggingface.co/google/embeddinggemma-2 | ⭐ 759 | Apache-2.0


架构:四个模态,一个坐标系

模型由三个部分组成,可以按需加载:

模块参数量职责
文本基座(backbone + embedder)270M(130M + 140M)处理文本和代码
视觉编码器170M处理图片和视频帧
音频编码器300M处理音频

所有模态的输出都投影到同一个 768 维向量空间。一句话和一张图片、一段录音的语义如果相近,它们的余弦相似度就会高——不需要跨模态转译,不需要把音频先转文字。


规格

项目值
总参数740M
输出维度768d(原生),支持 128/256/512d 截断
上下文长度8,192 tokens(上一代 2K,4 倍提升)
架构基座Gemma 4
支持模态文本(100+ 语言 + 代码)、图片、视频、音频
许可证Apache-2.0,可商用

基准测试

模态基准EmbeddingGemma 2EmbeddingGemma 1
文本(多语言)MTEB multilingual v261.3661.15
文本(代码)MTEB code v178.6868.76
图片MIEB lite64.64—
图片+视觉文档MMEB v2 Image57.28—
视频MMEB v2 Video50.67—
音频检索MSEB Retrieval69.54—

代码检索的提升最显著:78.68 对比 68.76,涨了约 +14.4 个百分点。多语言文本整体略有提升。

图片、视频、音频是新增模态,没有前代可比,但 MIEB 64.64 和 MSEB 69.54 是目前公开 768M 级别里的上游水位。


MRL:向量可以截断

模型原生支持 Matryoshka Representation Learning(MRL)——输出的 768 维向量可以截断到更低维度并重新归一化,质量损失在可接受范围内:

输出维度压缩比MTEB 多语言MTEB 代码
768d(完整)1:161.3678.68
512d1:1.561.1777.24
256d1:360.4176.18
128d1:657.8971.41

256d 以上截断质量损失很小,128d 更适合纯文本场景(图片和视频在 128d 时性能下降明显)。对于大规模本地知识库,256d 能把存储需求减到原来的 1/3。


端侧设计

模块化加载是关键设计决策:

  • 纯文本任务只加载 270M 基座,量化后约 191MB
  • 浏览器推理:20–70ms 单次查询(JavaScript)
  • 手机/笔记本可直接运行,不需要 GPU

这在实践中意味着:本地知识库索引、设备上 RAG、浏览器插件语义搜索,都不需要向云端发数据。


多模态统一的实际用法

文本检索的 API 是标准 sentence-transformers:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

query = "下周会议里讨论预算的部分"
# 可以是文字、图片路径、音频文件路径——同一个 encode() 调用
query_emb = model.encode(query, prompt_name="SearchQuery")

跨模态检索示例:用一句中文搜几小时录音里的片段——把文字和音频都 encode,直接算余弦相似度,不需要 ASR 中间步骤。


一句话说清楚

EmbeddingGemma 2 是一个把文字、图片、声音、视频的”意思”投影到同一张地图上的引擎。你的问题和相关内容在这张地图上越近,检索就越准——不管它们原来是什么格式。


和上一代的区别

特性EmbeddingGemma 1EmbeddingGemma 2
模态文本文本 + 图片 + 视频 + 音频
上下文2K8K
MTEB 代码68.7678.68
MRL 截断无✅ 128/256/512/768d
架构基座Gemma 2Gemma 4

Apache-2.0 开源。Google DeepMind 发布,740M 参数,HuggingFace 已上线,sentence-transformers 直接使用。开源仅供学习参考。


🇬🇧 English

EmbeddingGemma 2: Google Puts Text, Image, Audio, and Video Into a Single Embedding Space

Most semantic search systems share a common limitation: text uses one embedding model, images use another, audio requires yet another — and cross-modal retrieval needs manual alignment. A single audio segment is hard to search with a text query without first transcribing it.

EmbeddingGemma 2 eliminates this at the architecture level: one model, four modalities, one vector space.

HuggingFace: https://huggingface.co/google/embeddinggemma-2 | ⭐ 759 | Apache-2.0


Architecture: Four Modalities, One Coordinate System

The model has three loadable components:

ModuleParametersRole
Text backbone (backbone + embedder)270M (130M + 140M)Text and code
Vision encoder170MImages and video frames
Audio encoder300MAudio

All modalities project into a shared 768-dimensional vector space. A sentence and a semantically similar image or audio clip have high cosine similarity — no cross-modal translation, no forced ASR intermediate step.


Specifications

ItemValue
Total parameters740M
Output dimension768d native; 128/256/512d MRL truncation
Context length8,192 tokens (4x previous gen’s 2K)
Architecture baseGemma 4
Supported modalitiesText (100+ languages + code), images, video, audio
LicenseApache-2.0

Benchmarks

ModalityBenchmarkEmbeddingGemma 2EmbeddingGemma 1
Text (multilingual)MTEB multilingual v261.3661.15
Text (code)MTEB code v178.6868.76
ImageMIEB lite64.64—
Image + VisDocMMEB v257.28 / 67.84—
VideoMMEB v2 Video50.67—
Audio retrievalMSEB Retrieval69.54—

Code retrieval saw the biggest jump: +14.4 percentage points over previous gen. Multilingual text improved marginally.


MRL: Vectors Can Be Truncated

Native Matryoshka Representation Learning (MRL) support lets the 768d output be truncated to 512d, 256d, or 128d with minimal quality loss:

  • 256d: 3x storage reduction, minor quality loss (recommended for mixed-modality workloads)
  • 128d: 6x storage reduction, better suited for text-only workloads

For large local knowledge bases, 256d cuts storage requirements to one-third.


Edge Deployment

  • Text-only tasks: load only the 270M backbone, ~191MB quantized
  • Browser inference: 20–70ms per query (JavaScript)
  • Runs on mobile/laptop without GPU

Local knowledge base indexing, on-device RAG, and browser semantic search become possible without sending data to the cloud.


One-Line Summary

EmbeddingGemma 2 is an engine that maps the meaning of text, images, sounds, and video onto the same coordinate system — so your query and the relevant content are geometrically close, regardless of their original format.


Apache-2.0. Released by Google DeepMind, 740M parameters, available on HuggingFace via sentence-transformers. For technical reference only.

💬 评论与讨论

使用 GitHub 账号登录后发表评论

关于本站 · 免责声明

🍄 Mushroom Research Blog 是非营利、免费公开的个人科技观察博客与公众号 XStack18,不接受商业合作、不代表任何企业或机构立场,也不谋求商业利益。我们以个人视角客观中立地记录和分析 AI、Web3 等领域的最新模型发布与技术动态——不止转述新闻标题或二手信息,而是给出有独立思考的深入分析,希望帮更多人获得有价值的一手科技认知。

⚠️ 文中介绍的开源代码与模型,仅供学习交流与技术借鉴。它们大多仍处于早期阶段,有待进一步研究和验证,请勿直接用于工作或生产环境;如需采用,请先自行充分测试,并核实其许可证与安全性。
Open-source code and models featured here are shared for learning and reference only. Most are early-stage and still need further study and verification — please don't use them directly in your work or in production. Test them thoroughly and check their licenses and security first.

  1. 本站文章均为作者基于公开信息的个人研究与观点整理,不代表文中提及的任何公司、产品、模型的官方立场,未与其构成商业关联或合作关系。
  2. 科技行业信息更新极快,我们尽力保证内容准确、及时,但不对完整性、实时性做绝对保证,具体请以相关企业/项目官方公告为准。
  3. 文中引用的第三方商标、产品名称、图片、数据等版权归原权利人所有,我们会尽量注明来源;如你认为存在版权疑问或侵权,请通过下方邮箱联系我们,收到通知后会尽快核实处理(更正、加注来源或删除)。
  4. 文章内容仅为技术科普与个人观点,不构成投资、法律或其他专业建议,据此进行任何决策的后果需自行判断和承担。

📮 侵权 / 勘误 / 合作咨询:[email protected]