Google's EmbeddingGemma 2: On-Device Multimodal Embeddings for Edge AI
AI News

Google's EmbeddingGemma 2: On-Device Multimodal Embeddings for Edge AI

4 min
10/7/2026
EmbeddingGemma 2Google DeepMindmultimodal embeddingson-device AI

Google DeepMind launches EmbeddingGemma 2: a unified multimodal embedding model for edge AI

Google DeepMind has officially unveiled EmbeddingGemma 2, the successor to its popular EmbeddingGemma model, designed to natively map text, code, images, video, and audio into a single, unified embedding space. The release marks a significant step forward for on-device AI, enabling developers to build cross-modal search, retrieval, and RAG pipelines that operate entirely offline. With over 20 million downloads of the original EmbeddingGemma, the new model builds on that momentum while expanding far beyond text.

EmbeddingGemma 2 is built on the Gemma 4 architecture and is released under a commercially permissive Apache 2.0 license. The model packs 740 million parameters, but its modular design allows developers to use only the components they need—dropping to as little as 270M parameters for text-only workloads, with optional vision (170M) and audio (300M) encoders for full multimodal support. This flexibility makes it an attractive option for resource-constrained environments.

Why EmbeddingGemma 2 matters for developers

The original EmbeddingGemma was already a go-to for privacy-first RAG pipelines and on-device search, but it was limited to text. EmbeddingGemma 2 removes that limitation. By unifying multiple modalities into a shared embedding space, developers can now perform semantic searches across media libraries, find specific moments in video using text or audio queries, and even route decisions based on multimodal context—all without sending data to the cloud.

This is a critical development for industries where privacy is paramount, such as healthcare, finance, and enterprise search. Running embeddings locally reduces latency, cuts bandwidth costs, and ensures sensitive data never leaves the device. The model's efficiency is impressive: with quantization, it requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model on a Google Pixel 11 Pro.

continue reading below...

Technical highlights and performance benchmarks

EmbeddingGemma 2 delivers best-in-class performance for its size. Here are the key specs and improvements:

  • Best-in-class for sub-1B models: Leads benchmarks like MTEB Code and MAEB, with a 9.92-point improvement on code performance (from 68.76 to 78.68 on MTEB Code).
  • Modular architecture: Text-only mode uses 270M parameters; optional vision and audio encoders bring it to 740M total.
  • Storage-efficient: Uses Matryoshka Representation Learning (MRL) to truncate output vectors from 768 dimensions down to 512, 256, or 128, offering up to 6x storage reduction for local vector databases.
  • Extended context window: 8K tokens (4x larger than EmbeddingGemma 1), allowing processing of up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations.
  • On-device optimized: Runs efficiently within tight constraints, making it ideal for mobile and edge hardware.

In testing, EmbeddingGemma 2 not only matches the multilingual text performance of its predecessor but also outperforms some specialist models more than twice its size in vision and audio tasks. This quality-per-parameter ratio is a major selling point for developers who need high accuracy without the computational overhead of larger models.

Integration and ecosystem support

Google has ensured that EmbeddingGemma 2 works seamlessly with existing tools. The model shares its text tokenizer and audio encoder with Gemma 4, meaning developers can run both models in a unified pipeline with a lower combined memory footprint. This is particularly useful for on-device RAG systems that pair EmbeddingGemma 2 for retrieval with Gemma 4 for contextual reasoning.

Partners and tools supporting EmbeddingGemma 2 from day one include:

  • Hugging Face and Kaggle: Model weights are available for direct download.
  • Google AI Edge: MediaPipe and LiteRT integrations for turnkey embedding, retrieval, and decision tasks.
  • Browser deployment: Transformers.js and WebGPU support for in-browser inference.
  • Popular frameworks: vLLM, llama.cpp, SGLang, Ollama, MLX, and sentence-transformers.
  • Fine-tuning: Guidance from Unsloth for custom use cases.

Google has also launched demo apps in its AI Edge Gallery, including Instant Media Search and Video Moments Finder, to showcase real-world use cases.

Availability and getting started

EmbeddingGemma 2 is available now under the Apache 2.0 license, which allows commercial use. Developers can download the models from Hugging Face and Kaggle. For on-device deployment, Google recommends using MediaPipe for turnkey tasks or LiteRT for custom integration. The company has also published a comprehensive developer guide and documentation to help teams get started quickly.

With its combination of performance, efficiency, and permissive licensing, EmbeddingGemma 2 is poised to become a cornerstone of on-device multimodal AI. As edge computing continues to grow, this model gives developers the tools to build smarter, more private, and more responsive applications.