rkj dev

Google DeepMind Releases EmbeddingGemma 2 Open Multimodal Model

The sub-1B parameter open-weight model maps text, code, vision, and audio into a unified vector space for privacy-first, on-device semantic search.

Conceptual illustration of multimodal embeddings connecting text, audio, images, and video on a mobile device
Illustration: EmbeddingGemma 2 unites multiple data modalities directly on consumer hardware.AI-generated illustration

Key takeaways

  • EmbeddingGemma 2 features a 740M parameter footprint designed to map text, code, images, video, and audio into a single 768-dimensional vector space.
  • The model is released under a permissive Apache 2.0 license and supports modular loading, operating as a 270M parameter base for text and code workloads.
  • Quantized builds require approximately 191MB of active RAM for text weights and 567MB for full multimodal capabilities on a Google Pixel 11 Pro.
  • Matryoshka Representation Learning allows vector truncation from 768 dimensions down to 128 dimensions, enabling up to a 6x storage reduction.

On October 6, 2026, Google DeepMind launched EmbeddingGemma 2, an open-weight multimodal embedding model designed to run on consumer hardware such as smartphones and laptops. Building on the Gemma 4 architecture released earlier in the year, the model natively maps text, code, visual documents, images, video frames, and audio into a unified 768-dimensional vector space. Released under a commercially permissive Apache 2.0 license, the sub-1-billion parameter release expands on the original text-only EmbeddingGemma introduced in September 2025.

The initial EmbeddingGemma model achieved more than 20 million downloads, according to Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera. By expanding to vision and audio while keeping resource footprints low, EmbeddingGemma 2 aims to eliminate the latency and memory overhead typically caused by chaining separate speech-to-text, vision captioning, and text embedding models.

Modular Architecture and On-Device Efficiency

EmbeddingGemma 2 contains 740 million total parameters structured around a modular backbone. The base text and code model consists of 270 million parameters (130 million backbone parameters plus a 140 million embedder), augmented by an optional 170 million parameter vision encoder and a 300 million parameter audio encoder. Developers can selectively load only the required modalities at runtime, as outlined in Google's developer guide.

Illustration of a software developer configuring modular AI encoders
Illustration: The modular architecture allows loading only the required modality encoders to minimize memory usage.AI-generated illustration

This modularity translates to low resource consumption on edge devices. According to the Google AI Edge team, running quantized builds on a Google Pixel 11 Pro consumes approximately 191MB of active RAM for text-only weights and roughly 567MB for the full multimodal model. Inference speed measurements on a MacBook M5 Pro GPU demonstrate visual embedding latencies as low as 37.3 ms per image (about 26.9 images per second) at a budget of 70 vision tokens.

The model features an 8,192-token context window—a fourfold increase over its predecessor. This expanded capacity accommodates up to 5.5 minutes of 16 kHz mono audio (at 25 tokens per second), 29 images (at default 280 tokens per image), or 58 video frames (at 1 frame per second and 140 tokens per frame), as well as interleaved multimodal sequences containing text and media placeholder tokens.

Benchmark Performance and Matryoshka Compression

In standard evaluations detailed in the EmbeddingGemma 2 model card, the model posted major gains in code search, jumping 9.92 points from 68.76 on EmbeddingGemma 1 to 78.68 on the Massive Text Embedding Benchmark (MTEB Code v1, NDCG@10). Its multilingual text performance held steady at 61.36 on MTEB multilingual v2, matching the previous generation's 61.15 across more than 100 languages.

For non-text modalities, EmbeddingGemma 2 scored 64.64 on the Massive Image Embedding Benchmark (MIEB lite), 57.28 Hit@1 on MMEB v2 Image, 67.84 NDCG@5 on visual documents, 50.67 Hit@1 on video, 69.54 MRR@10 on the Massive Sound Embedding Benchmark (MSEB Retrieval), and 49.39 on the Massive Audio Embedding Benchmark (MAEB).

Abstract digital graphic representing Matryoshka representation learning vector dimension truncation
Illustration: Matryoshka dimension truncation enables developers to compress vector storage by up to six times.AI-generated illustration

To manage storage footprint in local vector databases, EmbeddingGemma 2 supports Matryoshka Representation Learning (MRL). Developers can dynamically truncate the default 768-dimensional vectors to 512, 256, or 128 dimensions followed by L2 re-normalization. At 256 dimensions, storage needs shrink by 3x while retaining around 95% of full quality on image, video, and speech retrieval. Truncating to 128 dimensions yields a 6x compression ratio (reducing a 1-million vector index from roughly 1.5 GB to 250 MB in bfloat16), though Google notes that multimodal recall drops to around 75% at 128 dimensions.

Developer Tooling and Applications

To showcase practical edge capabilities, Google updated the Google AI Edge Gallery app on Android and iOS with two interactive showcases: Instant Media Search, which performs search-as-you-type indexing in local SQLite databases, and Video Moments Finder, which identifies timestamps for actions like "kids laughing" without generating intermediate text transcripts, according to developers.googleblog.com.

On macOS, Google launched an experimental app called Google AI Edge Foresight. Foresight acts as an offline meeting companion that pairs EmbeddingGemma 2 with Gemma 4 models. Because EmbeddingGemma 2 shares its text tokenizer and audio encoder architecture with Gemma 4, running both models together reduces overall memory overhead. The app indexes live system microphone audio and meeting transcripts locally to turn shorthand notes into comprehensive summaries without transmitting voice data to cloud servers.

For cross-platform software engineering, MediaPipe Tasks provides turnkey Universal Embedder, Semantic Retriever, and Decision Task components for Web, iOS, macOS, Windows, and Linux. Google also confirmed that Android developers will gain native access through ML Kit in coming weeks, incorporating NPU hardware acceleration.

Availability and Technical Guidelines

EmbeddingGemma 2 model weights are available immediately under the Apache 2.0 license across Hugging Face and Kaggle, with deployment planned for the Gemini Enterprise Agent Platform Model Garden. The model is supported by serving frameworks including sentence-transformers (v6.1.0+), Hugging Face Transformers, LiteRT, vLLM, llama.cpp, Ollama, SGLang, and MLX, with fine-tuning support available via Unsloth and vector indexing via Qdrant.

Google advises developers to run inference in bfloat16 or float32, cautioning that float16 dynamic range limitations can produce NaN values or silent quality degradation. Additionally, the pre-trained embedding model did not undergo post-training safety alignment or output moderation; safety mitigations were applied exclusively during pre-training data curation, requiring developers to implement application-level filtering downstream.

Frequently asked questions

What modalities are supported by EmbeddingGemma 2?

EmbeddingGemma 2 natively maps text, source code, images, video frames, and audio into a shared 768-dimensional embedding space.

How much memory does EmbeddingGemma 2 require on mobile devices?

When quantized on a Google Pixel 11 Pro, the 270M parameter text core requires roughly 191MB of active RAM, while the full 740M parameter multimodal model uses about 567MB.

What license is EmbeddingGemma 2 released under?

EmbeddingGemma 2 is released under a commercially permissive Apache 2.0 license with model weights hosted on Hugging Face and Kaggle.

How does Matryoshka Representation Learning help with vector storage?

MRL allows output vectors to be sliced from 768 dimensions down to 512, 256, or 128 dimensions. Truncating to 128 dimensions delivers up to a 6x reduction in vector database storage.

Sources

  1. Bring multimodal semantic search to the edge with EmbeddingGemma 2developers.googleblog.com · Oct 6, 2026 · Official
  2. EmbeddingGemma 2: The Developer Guidedevelopers.googleblog.com · Oct 6, 2026 · Official
  3. EmbeddingGemma 2: an open, lightweight multimodal embedding modelGoogle · Oct 6, 2026 · Official
  4. EmbeddingGemma 2 model card | Google AI for Developersai.google.dev · Oct 6, 2026 · Official
  5. Google expands EmbeddingGemma beyond text to images, audio and videoSiliconANGLE · Oct 6, 2026
  6. Google launches EmbeddingGemma 2, an open multimodal embedding model for devicesTNW | Artificial-intelligence · Oct 6, 2026
  7. Google AI Edge Foresight for Mac turns scribbles into full notes using offline recordings9to5Google · Oct 6, 2026

How this story was made: the newsroom picked it up from developers.googleblog.com, Techmeme and Google News, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (47 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.

#Google DeepMind #EmbeddingGemma 2 #Multimodal AI #On-Device AI #Open Source #Gemma 4

Published October 7, 2026 at 02:06 UTC