Google DeepMind launches EmbeddingGemma 2
Google DeepMind launched EmbeddingGemma 2, an open-weight multimodal embedding model, on Oct. 6, 2026. Google says it maps text, code, images, video and audio into a shared 768-dimensional space.

Key takeaways · 4
- 01
The full multimodal configuration has 740 million parameters; Google gives its size as about 567 MB on a Google Pixel 11 Pro.
- 02
The text-only weights can run with about 191 MB of active RAM, according to Google.
- 03
Google says developers can reduce embeddings from 768 to 128 dimensions; 256 dimensions retain about 95% of original quality for image, video and speech retrieval.
- 04
Google plans to offer the model as an Android service through ML Kit in the coming weeks, with NPU acceleration.
One model for five data types
Google says EmbeddingGemma 2 maps text, code, images, video and audio into a shared 768-dimensional space.[2] Its listed configurations range from 270 million parameters for text and code to 740 million for all modalities; the other options are 440 million for text and vision, and 570 million for text and audio.[2] The model is based on Gemma 4, has fewer than 1 billion parameters and is released under Apache 2.0.[2] Google calls it “best-in-class for its size,” a company characterization rather than an independently established result in the available information.[1]
Choose encoders and embedding size
Developers can selectively load only the encoders they need at runtime, according to Google.[2] That means a project can choose among the listed modality configurations rather than defaulting to the full multimodal model.[2] Google gives the full model’s size as about 567 MB on a Google Pixel 11 Pro, while its text-only weights can run with about 191 MB of active RAM.[1] Google also says Matryoshka Representation Learning can reduce embeddings from 768 dimensions to 128; at 256 dimensions, image, video and speech retrieval retains about 95% of original quality.[2]
Local use and Android plans
Google says the model is designed for local, privacy-first applications.[1] For retrieval, the company says using it can reduce latency and memory overhead compared with chaining separate image-captioning, speech-to-text and text-embedding models.[1] Google also says the model can route intents in milliseconds without training data or fine-tuning.[1] Google AI Edge Gallery on mobile and Google AI Edge Foresight on Mac are named in the announcement.[1] The company plans to make EmbeddingGemma 2 available as an Android service through ML Kit in the coming weeks, using NPU acceleration to optimize performance across a broad range of devices.[1]
Teams evaluating retrieval or intent routing can compare a single shared embedding model with pipelines that chain separate modality-specific systems. The configuration choices, reported memory figures and planned Android service give developers concrete options to assess against their deployment constraints.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
6 October 2026
Google DeepMind launches EmbeddingGemma 2
Sources
- Bring multimodal semantic search to the edge with EmbeddingGemma 2 - Google Developers Blogdevelopers.googleblog.com
- EmbeddingGemma 2: The Developer Guide - Google Developers Blogdevelopers.googleblog.com