Illustration of the Empacer lightning mascot holding a pencil next to written pages and stars

News

EmbeddingGemma 2: Google's Multimodal Embedding Model, Specs and Benchmarks

Empacer team11 min read
On this page9 sections

Google DeepMind released EmbeddingGemma 2 on October 6, 2026: an open embedding model that maps text, code, images, video, and audio into a single 768-dimensional vector space. It is built on the Gemma 4 architecture, has 740 million parameters in its full form and 270 million for text alone, and ships under an Apache 2.0 license. Google says the first EmbeddingGemma passed 20 million downloads.

What is an embedding model?

An embedding model does not write answers. It turns content into vectors so that similar things sit close together, which is what powers semantic search, clustering, and the retrieval step in retrieval-augmented generation. What is new here is that one small model does this for five content types at once, on a phone or laptop, so a spoken query can find a video clip and a text query can find a diagram. Below: what changed, how the modular design and context budget work, what Google's benchmark table shows and does not, the license, where to run it, and what is not yet independently measured. No hands-on test is included: the article works from Google's published model card and announcement.

EmbeddingGemma 2 at a glance

  • Release date: October 6, 2026
  • What it does: embeds text, code, images, visual documents, video, and audio into one shared 768-dimensional space; cross-modal search works directly
  • Size: 740M parameters full; 270M text and code only; 440M text plus vision; 570M text plus audio
  • Context: 8,192 tokens shared across modalities
  • Output dimensions: 768, truncatable to 512, 256, or 128
  • Languages: more than 100
  • License: Apache 2.0
  • Weights: Hugging Face and Kaggle now; Google Cloud Model Garden coming soon
  • Training data cutoff: January 2025

What is new in EmbeddingGemma 2?

1. One space for text, code, images, video, and audio

EmbeddingGemma 1 embedded text only. EmbeddingGemma 2 keeps a 270M text backbone, made of a 130M transformer and a 140M embedder, and adds two optional encoders: vision at 170M for images, PDFs, slides, charts, and video frames, and audio at 300M for speech and other sounds. All inputs pass through the shared backbone and land in the same space, so you can load only the encoders your data needs and still compare results: a query embedded with the text-only setup matches against documents embedded with the full model, and an index built with EmbeddingGemma 2 on text can later gain image or audio entries without re-embedding what is already there. That compatibility is between configurations of EmbeddingGemma 2 only. Vectors produced by EmbeddingGemma 1 live in a different space, and an index built with the old model has to be re-embedded to move to the new one.

Inputs can be interleaved. A product listing with text, two photos, and a video becomes one embedding, with placeholder tokens marking where each media item sits. Text inputs take a short task prefix (search query, document, code retrieval, classification, and so on) that steers the embedding; media inputs take none.

2. An 8K context with a token budget per modality

The context window is 8,192 tokens, four times EmbeddingGemma 1's 2,000, and every modality draws on the same budget at a fixed rate.

ModalityToken cost at default settingsMaximum per input, single modality
Text1 token per subword8,192 tokens
Image280 tokens per imageabout 29 images
Video140 tokens per frame, sampled at 1 frame per secondabout 58 frames
Audio25 tokens per second, 16 kHz monoabout 327 seconds

Source: Hugging Face model card. Mixing modalities in one input reduces how much of each fits. The vision token budget per image is adjustable from 70 to 1,120 soft tokens; at the low end roughly 114 images fit, at a cost in fine-grained detail.

3. Matryoshka dimensions, with Google's own quality table

Matryoshka Representation Learning lets you keep only the leading dimensions of a vector. In bfloat16, a million 768-dimensional vectors take about 1.5 GB of storage; at 128 dimensions, about 250 MB. Google describes this as up to a 6x reduction in vector storage costs. The trade-off is in the model card.

Output dimensionStorage vs 768MTEB multilingualMTEB codeMIEB lite (image)MMEB v2 overallMSEB retrieval (audio)
7681x61.3678.6864.6459.0169.54
5120.67x61.1777.2464.3258.3869.18
2560.33x60.4176.1863.1356.2466.76
1280.17x57.8971.4159.0645.6556.71

Source: Hugging Face model card, full-precision checkpoint. Down to 256 dimensions the loss is small across the board. At 128 the multimodal scores fall hard: MMEB overall drops 13 points and audio retrieval 13 points, which is why Google recommends 128 only for text-heavy indexes and first-stage shortlisting. Two operational rules from the card: re-normalize after truncating, or cosine similarity degrades silently, and never score a query at one dimension against a corpus at another.

4. Apache 2.0 instead of Gemma terms

EmbeddingGemma 1's model card points to the Gemma Terms of Use. EmbeddingGemma 2 is released under Apache 2.0, a standard permissive license that permits commercial use and modification. Google separately requires deployments to follow the Gemma Prohibited Use Policy, so review that document for your use case. The model has had no safety tuning or output moderation, because it produces vectors rather than text; Google places responsibility for retrieval filtering and fairness testing on the deployer.

5. Built for a phone

Google quotes roughly 191 MB of active RAM for the text-only weights and about 567 MB for the full multimodal model on a Pixel 11 Pro with quantization. Because EmbeddingGemma 2 shares its text tokenizer and audio encoder architecture with Gemma 4, running both in one on-device RAG pipeline lowers the combined footprint. One precision warning from the model card matters in practice: run in bfloat16 or float32, never float16, which overflows and returns NaN or quietly degraded embeddings with no error.

EmbeddingGemma 2 benchmark results

All figures are Google's, reported on public benchmarks at full precision and 768 dimensions. The MTEB and MIEB families are public leaderboards where others can reproduce the runs; as of October 7, 2026, no independent evaluation of EmbeddingGemma 2 had been published.

BenchmarkWhat it measuresMetricEmbeddingGemma 2EmbeddingGemma 1
MTEB multilingual v2Text tasks across languagesMean over tasks, mixed metrics61.3661.15
MTEB English v2English text tasksMean over tasks, mixed metrics68.4669.67
MTEB code v1Code retrievalNDCG@1078.6868.76
MIEB liteImage embedding tasksMean over task types, mixed metrics64.64not applicable
MMEB v2, imageMultimodal image tasksHit@157.28not applicable
MMEB v2, visual documentsPDFs, slides, chartsNDCG@567.84not applicable
MMEB v2, videoVideo tasksHit@150.67not applicable
MSEB retrievalAudio retrievalMRR@1069.54not applicable
MAEBAudio embedding tasksMean over tasks, mixed metrics49.39not applicable

Sources: Hugging Face model card for EmbeddingGemma 2; Google AI for Developers model card for EmbeddingGemma 1. The rows use different metrics, so they cannot be read as one scale: a video Hit@1 of 50.67 and an audio MRR@10 of 69.54 do not mean video retrieval is worse than audio retrieval. Compare each row only with its own predecessor or with other models on the same benchmark. Three readings. The code gain is the one Google leads with, 9.92 points. Multilingual text is flat, which Google describes as matching its predecessor. English text is slightly lower than EmbeddingGemma 1, a 1.2-point drop that Google's announcement does not mention. The multimodal rows have no predecessor to compare against; Google's claim that they are best in class among sub-1 billion-parameter multimodal embedders, and that the model beats some specialist models more than twice its size, is the company's reading of the leaderboards, not an independent result.

Where it runs on day one

Google lists weights on Hugging Face and Kaggle, with Model Garden on Google Cloud to follow, and names these integrations: Hugging Face Transformers and sentence-transformers (version 6.1.0 or later for the multimodal inputs), vLLM, SGLang, MLX, llama.cpp, Ollama, and LM Studio for serving; LiteRT and MediaPipe for on-device apps; transformers.js and WebGPU for the browser; Qdrant for vector storage; Unsloth for fine-tuning. Being listed does not mean every runtime accepts every input type. Ollama's model page, for example, offers the 270M, 440M, 570M, and 740M variants but lists text and image inputs only, with no audio or video. Check modality support and model format for your chosen runtime before planning a multimodal index around it.

Who should use EmbeddingGemma 2

  • Developers indexing a codebase locally or building coding-agent retrieval: the code score is the main gain, and the 270M text-only setup is the lightest way to get it.
  • Teams building private, offline search over mixed media: photos, screenshots, slide decks, call recordings, and documents in one index on a laptop or phone is the use case the model was built for.
  • Anyone already on EmbeddingGemma 1 for English text: the English score is marginally lower, so test on your own corpus before migrating; the gains are in code and in modalities the old model did not have.
  • Teams that need the strongest text embeddings regardless of size: this is a sub-1B on-device model. Larger cloud embedding models score higher on text; the trade is privacy, latency, and cost.

What is not yet known

QuestionStatus on October 7, 2026
Independent benchmark runsNone published; all scores are Google's
Scores for the quantized on-device variantsMemory figures given; quality at quantization not in the model card
Comparison against other small embedders such as Qwen3 Embedding 0.6B or BGENot published by Google; check the MTEB leaderboard
Google Cloud Model Garden availability"Coming soon"
Video and audio limits in practiceDefaults documented (1 fps, 16 kHz mono); quality at other settings not reported

What this means for small business owners

Embedding models sit underneath the tools people use rather than in front of them. They are what lets a search find the right document, photo, or recording by meaning rather than by exact words, and a small model that runs on a phone keeps that search private and fast. Owners of small service businesses usually meet this technology inside finished products, not as a model to set up. Empacer is a voice-first AI business assistant for small service business owners. You say what needs doing, review the finished work, and decide whether it goes out.

FAQ

What is EmbeddingGemma 2?

EmbeddingGemma 2 is an open embedding model from Google DeepMind, released on October 6, 2026. It turns text, code, images, video, and audio into vectors in one shared 768-dimensional space, so content of different types can be searched together. It has 740 million parameters in its full form, 270 million for text alone, and runs on phones and laptops.

Is EmbeddingGemma 2 free for commercial use?

Apache 2.0 permits commercial use and modification. Google also requires deployments to follow the separately stated Gemma Prohibited Use Policy, so review it for your case. EmbeddingGemma 1 shipped under the Gemma Terms of Use instead.

Can I keep an index built with EmbeddingGemma 1?

No. The compatibility Google describes is between the 270M, 440M, 570M, and 740M configurations of EmbeddingGemma 2, which share one vector space. EmbeddingGemma 1 vectors are not in that space, so moving an existing index to EmbeddingGemma 2 means re-embedding the corpus.

Do I need the full 740M model?

Only if you embed audio and video as well as images. Text and code need 270M, text plus images 440M, text plus audio 570M. All four setups load from the same checkpoint and share one vector space, so indexes built with one setup remain compatible with another.

How much quality do I lose by truncating to 128 dimensions?

A measurable amount for text and code and a large amount for media. MTEB code falls from 78.68 to 71.41, about 7 points; Google's MMEB overall score drops from 59.01 to 45.65 and audio retrieval from 69.54 to 56.71. Google positions 128 dimensions for text-heavy indexes and first-stage shortlisting before re-ranking, recommends validating it on your own data, and treats 256 as the floor for multimodal indexes.

Share: