Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for EmbeddingGemma 2: On-Device Search for Text, Images, Audio

EmbeddingGemma 2: On-Device Search for Text, Images, Audio

EmbeddingGemma 2 is a 740M open embedding model that searches text, images, audio and video on a phone using as little as 191MB of RAM. See how it works.

Dr. Nova Chen
Dr. Nova Chen★Oct 6, 2026★3 min read

Google released EmbeddingGemma 2 on October 6, 2026, and it is a small model with a big job: turning text, code, images, audio and video into one shared set of searchable vectors, right on your phone. EmbeddingGemma 2 is open under Apache 2.0 and built on Gemma 4, so developers can ship private semantic search and retrieval-augmented generation without sending user data to a server.

  • Size: 740 million parameters, split into a 270M text core with optional 170M vision and 300M audio encoders.
  • Context: 8K tokens, enough for about 5.5 minutes of audio, 29 images or 58 video frames.
  • Memory: roughly 191MB of active RAM for text-only use and about 567MB fully multimodal on a Pixel 11 Pro.
  • License and access: Apache 2.0, with downloads on Hugging Face and Kaggle.

What Is an Embedding Model, and Why Does It Matter?

An embedding model converts content into a list of numbers that captures its meaning. Similar ideas land close together, so a search for "beach sunset" can find a photo, a voice memo or a video clip about the same thing even if none of them contain those words. Embeddings are the retrieval layer behind most RAG systems, which is why a better on-device embedding model quietly improves every local AI assistant built on top of it.

What Is New in EmbeddingGemma 2?

The headline change is multimodality. The original EmbeddingGemma handled text; EmbeddingGemma 2 maps text, images, audio and video into a single embedding space. The model is modular, so an app that only needs text search can load the 270M core and skip the vision and audio encoders entirely. Google also quadrupled the context window to 8K tokens and reports a 9.92-point jump on the MTEB Code benchmark, from 68.76 to 78.68, which makes it stronger for local codebase search.

Storage gets smarter too. The default vectors have 768 dimensions, but Matryoshka Representation Learning lets developers truncate them to 512, 256 or 128 dimensions, which Google says can cut storage by up to six times with a modest trade in accuracy.

How Can Developers Run EmbeddingGemma 2?

Google lists support across Transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, with optimized LiteRT versions on Hugging Face for mobile. That wide toolchain support means the on-device AI model can slot into an existing pipeline with little rework. Google suggests use cases such as on-device semantic search, local code indexing, searching a personal media library and finding a moment in a video with a text or audio query.

Why On-Device AI Search Is Worth Watching

Running embeddings locally keeps personal photos, recordings and documents on the device, and it removes per-query cloud costs. The Gemma family has already proven popular with local builders, as our look at Gemma 4 QAT in Ollama showed. A multimodal embedding model under 1GB of RAM is a practical building block for the next wave of private assistants. More in our AI section.

Sources: Google — October 6, 2026; Investing.com — October 6, 2026.

More AI Stories