EmbeddingGemma 2: an open, lightweight multimodal embedding model
2026-10-11 · Google DeepMind
EmbeddingGemma 2: an open, lightweight multimodal embedding model
Overview
Google DeepMind launched EmbeddingGemma 2 on October 6, 2026, the latest version of its on-device multimodal embedding model. The original EmbeddingGemma, introduced last year, has surpassed 20 million downloads, powering smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines. EmbeddingGemma 2 expands beyond text to unify code, images, video, and audio in a shared embedding space.
Key Features
Built on the Gemma 4 architecture and released under the commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference. Its key features include:
- Best-in-class for its size: Achieves leading scores among sub-1B multimodal embedders on benchmarks like MTEB Code and MAEB, while matching or outperforming many larger models across text, vision, and audio tasks.
- Modular by design: Requires as little as 270M parameters for text-only workloads, with optional vision (170M) and audio (300M) encoders for full multimodal support.
- Storage-efficient: Using Matryoshka Representation Learning (MRL), developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions, providing up to 6x storage reduction.
- Optimized for on-device performance: With quantization on a Google Pixel 11 Pro, requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.
- Extended context ready: Features an 8K token context window (4x larger than EmbeddingGemma 1), processing up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations.
Performance Improvements
EmbeddingGemma 2 maintains the strong multilingual text performance of its predecessor while delivering a significant 9.92-point improvement on code performance (MTEB Code: 68.76 to 78.68), making it well-suited for local codebase indexing, semantic code search, and coding agent retrieval. Across image, video, document, and audio tasks, it sets a new standard in quality-per-parameter for sub-1B models, even outperforming some specialist models more than twice its size.
Use Cases
EmbeddingGemma 2 brings robust capabilities directly to edge hardware. Generating embeddings locally ensures data privacy, reduces pipeline latency, and supports fully offline cross-modal search and retrieval. Specific applications include:
- Instant Media Search: Use text or images to find top matches in media libraries based on semantic similarity.
- Video Moments Finder: Locate specific moments in video using text or audio queries.
- Local file retrieval with contextual reasoning: Pair EmbeddingGemma 2 for local file retrieval with Gemma 4 for contextual reasoning.
- Real-time decision engines: Leverage multimodal context for classification, routing, and predictive capabilities via the MediaPipe Decision Task API.
Availability and Deployment
Model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming soon. Developers can:
- Deploy on-device: Build cross-platform apps with Google AI Edge MediaPipe or LiteRT, or run in browsers with transformers.js or WebGPU.
- Use favorite development tools: Serve the model with transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio. Store embedding vectors with Qdrant.
- Fine-tune: Follow guidance by Unsloth for fine-tuning EmbeddingGemma 2 for specific use cases.