
News
EmbeddingGemma 2: Google's Multimodal Embedding Model, Specs and Benchmarks
On this page9 sections
Google DeepMind released EmbeddingGemma 2 on October 6, 2026: an open embedding model that maps text, code, images, video, and audio into a single 768-dimensional vector space. It is built on the Gemma 4 architecture, has 740 million parameters in its full form and 270 million for text alone, and ships under an Apache 2.0 license. Google says the first EmbeddingGemma passed 20 million downloads.
What is an embedding model?
An embedding model does not write answers. It turns content into vectors so that similar things sit close together, which is what powers semantic search, clustering, and the retrieval step in retrieval-augmented generation. What is new here is that one small model does this for five content types at once, on a phone or laptop, so a spoken query can find a video clip and a text query can find a diagram. Below: what changed, how the modular design and context budget work, what Google's benchmark table shows and does not, the license, where to run it, and what is not yet independently measured. No hands-on test is included: the article works from Google's published model card and announcement.
EmbeddingGemma 2 at a glance
- Release date: October 6, 2026
- What it does: embeds text, code, images, visual documents, video, and audio into one shared 768-dimensional space; cross-modal search works directly
- Size: 740M parameters full; 270M text and code only; 440M text plus vision; 570M text plus audio
- Context: 8,192 tokens shared across modalities
- Output dimensions: 768, truncatable to 512, 256, or 128
- Languages: more than 100
- License: Apache 2.0
- Weights: Hugging Face and Kaggle now; Google Cloud Model Garden coming soon
- Training data cutoff: January 2025
What is new in EmbeddingGemma 2?
1. One space for text, code, images, video, and audio
EmbeddingGemma 1 embedded text only. EmbeddingGemma 2 keeps a 270M text backbone, made of a 130M transformer and a 140M embedder, and adds two optional encoders: vision at 170M for images, PDFs, slides, charts, and video frames, and audio at 300M for speech and other sounds. All inputs pass through the shared backbone and land in the same space, so you can load only the encoders your data needs and still compare results: a query embedded with the text-only setup matches against documents embedded with the full model, and an index built with EmbeddingGemma 2 on text can later gain image or audio entries without re-embedding what is already there. That compatibility is between configurations of EmbeddingGemma 2 only. Vectors produced by EmbeddingGemma 1 live in a different space, and an index built with the old model has to be re-embedded to move to the new one.
Inputs can be interleaved. A product listing with text, two photos, and a video becomes one embedding, with placeholder tokens marking where each media item sits. Text inputs take a short task prefix (search query, document, code retrieval, classification, and so on) that steers the embedding; media inputs take none.
2. An 8K context with a token budget per modality
The context window is 8,192 tokens, four times EmbeddingGemma 1's 2,000, and every modality draws on the same budget at a fixed rate.
| Modality | Token cost at default settings | Maximum per input, single modality |
|---|---|---|
| Text | 1 token per subword | 8,192 tokens |
| Image | 280 tokens per image | about 29 images |
| Video | 140 tokens per frame, sampled at 1 frame per second | about 58 frames |
| Audio | 25 tokens per second, 16 kHz mono | about 327 seconds |
Source: Hugging Face model card. Mixing modalities in one input reduces how much of each fits. The vision token budget per image is adjustable from 70 to 1,120 soft tokens; at the low end roughly 114 images fit, at a cost in fine-grained detail.
3. Matryoshka dimensions, with Google's own quality table
Matryoshka Representation Learning lets you keep only the leading dimensions of a vector. In bfloat16, a million 768-dimensional vectors take about 1.5 GB of storage; at 128 dimensions, about 250 MB. Google describes this as up to a 6x reduction in vector storage costs. The trade-off is in the model card.
| Output dimension | Storage vs 768 | MTEB multilingual | MTEB code | MIEB lite (image) | MMEB v2 overall | MSEB retrieval (audio) |
|---|---|---|---|---|---|---|
| 768 | 1x | 61.36 | 78.68 | 64.64 | 59.01 | 69.54 |
| 512 | 0.67x | 61.17 | 77.24 | 64.32 | 58.38 | 69.18 |
| 256 | 0.33x | 60.41 | 76.18 | 63.13 | 56.24 | 66.76 |
| 128 | 0.17x | 57.89 | 71.41 | 59.06 | 45.65 | 56.71 |
Source: Hugging Face model card, full-precision checkpoint. Down to 256 dimensions the loss is small across the board. At 128 the multimodal scores fall hard: MMEB overall drops 13 points and audio retrieval 13 points, which is why Google recommends 128 only for text-heavy indexes and first-stage shortlisting. Two operational rules from the card: re-normalize after truncating, or cosine similarity degrades silently, and never score a query at one dimension against a corpus at another.
4. Apache 2.0 instead of Gemma terms
EmbeddingGemma 1's model card points to the Gemma Terms of Use. EmbeddingGemma 2 is released under Apache 2.0, a standard permissive license that permits commercial use and modification. Google separately requires deployments to follow the Gemma Prohibited Use Policy, so review that document for your use case. The model has had no safety tuning or output moderation, because it produces vectors rather than text; Google places responsibility for retrieval filtering and fairness testing on the deployer.
5. Built for a phone
Google quotes roughly 191 MB of active RAM for the text-only weights and about 567 MB for the full multimodal model on a Pixel 11 Pro with quantization. Because EmbeddingGemma 2 shares its text tokenizer and audio encoder architecture with Gemma 4, running both in one on-device RAG pipeline lowers the combined footprint. One precision warning from the model card matters in practice: run in bfloat16 or float32, never float16, which overflows and returns NaN or quietly degraded embeddings with no error.
EmbeddingGemma 2 benchmark results
All figures are Google's, reported on public benchmarks at full precision and 768 dimensions. The MTEB and MIEB families are public leaderboards where others can reproduce the runs; as of October 7, 2026, no independent evaluation of EmbeddingGemma 2 had been published.
| Benchmark | What it measures | Metric | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|---|---|
| MTEB multilingual v2 | Text tasks across languages | Mean over tasks, mixed metrics | 61.36 | 61.15 |
| MTEB English v2 | English text tasks | Mean over tasks, mixed metrics | 68.46 | 69.67 |
| MTEB code v1 | Code retrieval | NDCG@10 | 78.68 | 68.76 |
| MIEB lite | Image embedding tasks | Mean over task types, mixed metrics | 64.64 | not applicable |
| MMEB v2, image | Multimodal image tasks | Hit@1 | 57.28 | not applicable |
| MMEB v2, visual documents | PDFs, slides, charts | NDCG@5 | 67.84 | not applicable |
| MMEB v2, video | Video tasks | Hit@1 | 50.67 | not applicable |
| MSEB retrieval | Audio retrieval | MRR@10 | 69.54 | not applicable |
| MAEB | Audio embedding tasks | Mean over tasks, mixed metrics | 49.39 | not applicable |
Sources: Hugging Face model card for EmbeddingGemma 2; Google AI for Developers model card for EmbeddingGemma 1. The rows use different metrics, so they cannot be read as one scale: a video Hit@1 of 50.67 and an audio MRR@10 of 69.54 do not mean video retrieval is worse than audio retrieval. Compare each row only with its own predecessor or with other models on the same benchmark. Three readings. The code gain is the one Google leads with, 9.92 points. Multilingual text is flat, which Google describes as matching its predecessor. English text is slightly lower than EmbeddingGemma 1, a 1.2-point drop that Google's announcement does not mention. The multimodal rows have no predecessor to compare against; Google's claim that they are best in class among sub-1 billion-parameter multimodal embedders, and that the model beats some specialist models more than twice its size, is the company's reading of the leaderboards, not an independent result.
Where it runs on day one
Google lists weights on Hugging Face and Kaggle, with Model Garden on Google Cloud to follow, and names these integrations: Hugging Face Transformers and sentence-transformers (version 6.1.0 or later for the multimodal inputs), vLLM, SGLang, MLX, llama.cpp, Ollama, and LM Studio for serving; LiteRT and MediaPipe for on-device apps; transformers.js and WebGPU for the browser; Qdrant for vector storage; Unsloth for fine-tuning. Being listed does not mean every runtime accepts every input type. Ollama's model page, for example, offers the 270M, 440M, 570M, and 740M variants but lists text and image inputs only, with no audio or video. Check modality support and model format for your chosen runtime before planning a multimodal index around it.
Who should use EmbeddingGemma 2
- Developers indexing a codebase locally or building coding-agent retrieval: the code score is the main gain, and the 270M text-only setup is the lightest way to get it.
- Teams building private, offline search over mixed media: photos, screenshots, slide decks, call recordings, and documents in one index on a laptop or phone is the use case the model was built for.
- Anyone already on EmbeddingGemma 1 for English text: the English score is marginally lower, so test on your own corpus before migrating; the gains are in code and in modalities the old model did not have.
- Teams that need the strongest text embeddings regardless of size: this is a sub-1B on-device model. Larger cloud embedding models score higher on text; the trade is privacy, latency, and cost.
What is not yet known
| Question | Status on October 7, 2026 |
|---|---|
| Independent benchmark runs | None published; all scores are Google's |
| Scores for the quantized on-device variants | Memory figures given; quality at quantization not in the model card |
| Comparison against other small embedders such as Qwen3 Embedding 0.6B or BGE | Not published by Google; check the MTEB leaderboard |
| Google Cloud Model Garden availability | "Coming soon" |
| Video and audio limits in practice | Defaults documented (1 fps, 16 kHz mono); quality at other settings not reported |
What this means for small business owners
Embedding models sit underneath the tools people use rather than in front of them. They are what lets a search find the right document, photo, or recording by meaning rather than by exact words, and a small model that runs on a phone keeps that search private and fast. Owners of small service businesses usually meet this technology inside finished products, not as a model to set up. Empacer is a voice-first AI business assistant for small service business owners. You say what needs doing, review the finished work, and decide whether it goes out.
FAQ
What is EmbeddingGemma 2?
EmbeddingGemma 2 is an open embedding model from Google DeepMind, released on October 6, 2026. It turns text, code, images, video, and audio into vectors in one shared 768-dimensional space, so content of different types can be searched together. It has 740 million parameters in its full form, 270 million for text alone, and runs on phones and laptops.
Is EmbeddingGemma 2 free for commercial use?
Apache 2.0 permits commercial use and modification. Google also requires deployments to follow the separately stated Gemma Prohibited Use Policy, so review it for your case. EmbeddingGemma 1 shipped under the Gemma Terms of Use instead.
Can I keep an index built with EmbeddingGemma 1?
No. The compatibility Google describes is between the 270M, 440M, 570M, and 740M configurations of EmbeddingGemma 2, which share one vector space. EmbeddingGemma 1 vectors are not in that space, so moving an existing index to EmbeddingGemma 2 means re-embedding the corpus.
Do I need the full 740M model?
Only if you embed audio and video as well as images. Text and code need 270M, text plus images 440M, text plus audio 570M. All four setups load from the same checkpoint and share one vector space, so indexes built with one setup remain compatible with another.
How much quality do I lose by truncating to 128 dimensions?
A measurable amount for text and code and a large amount for media. MTEB code falls from 78.68 to 71.41, about 7 points; Google's MMEB overall score drops from 59.01 to 45.65 and audio retrieval from 69.54 to 56.71. Google positions 128 dimensions for text-heavy indexes and first-stage shortlisting before re-ranking, recommends validating it on your own data, and treats 256 as the floor for multimodal indexes.