Skip to main content

Google DeepMind launches EmbeddingGemma 2

6 OCTOBER 2026·2 MIN READ·2 SOURCES

Google DeepMind launched EmbeddingGemma 2, an open-weight multimodal embedding model, on Oct. 6, 2026. Google says it maps text, code, images, video and audio into a shared 768-dimensional space.

Google DeepMind launches EmbeddingGemma 2

Key takeaways · 4

  • 01

    The full multimodal configuration has 740 million parameters; Google gives its size as about 567 MB on a Google Pixel 11 Pro.

  • 02

    The text-only weights can run with about 191 MB of active RAM, according to Google.

  • 03

    Google says developers can reduce embeddings from 768 to 128 dimensions; 256 dimensions retain about 95% of original quality for image, video and speech retrieval.

  • 04

    Google plans to offer the model as an Android service through ML Kit in the coming weeks, with NPU acceleration.

One model for five data types

Google says EmbeddingGemma 2 maps text, code, images, video and audio into a shared 768-dimensional space.[2] Its listed configurations range from 270 million parameters for text and code to 740 million for all modalities; the other options are 440 million for text and vision, and 570 million for text and audio.[2] The model is based on Gemma 4, has fewer than 1 billion parameters and is released under Apache 2.0.[2] Google calls it “best-in-class for its size,” a company characterization rather than an independently established result in the available information.[1]

Choose encoders and embedding size

Developers can selectively load only the encoders they need at runtime, according to Google.[2] That means a project can choose among the listed modality configurations rather than defaulting to the full multimodal model.[2] Google gives the full model’s size as about 567 MB on a Google Pixel 11 Pro, while its text-only weights can run with about 191 MB of active RAM.[1] Google also says Matryoshka Representation Learning can reduce embeddings from 768 dimensions to 128; at 256 dimensions, image, video and speech retrieval retains about 95% of original quality.[2]

Local use and Android plans

Google says the model is designed for local, privacy-first applications.[1] For retrieval, the company says using it can reduce latency and memory overhead compared with chaining separate image-captioning, speech-to-text and text-embedding models.[1] Google also says the model can route intents in milliseconds without training data or fine-tuning.[1] Google AI Edge Gallery on mobile and Google AI Edge Foresight on Mac are named in the announcement.[1] The company plans to make EmbeddingGemma 2 available as an Android service through ML Kit in the coming weeks, using NPU acceleration to optimize performance across a broad range of devices.[1]

Teams evaluating retrieval or intent routing can compare a single shared embedding model with pipelines that chain separate modality-specific systems. The configuration choices, reported memory figures and planned Android service give developers concrete options to assess against their deployment constraints.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 6 October 2026

    Google DeepMind launches EmbeddingGemma 2

Sources

AI fluency, one session a day, built for your work.