Trending: On-device modelsSearch
iHeartGeek
iTECH

Google's EmbeddingGemma 2 puts multimodal search on the device

The 740-million-parameter model maps text, code, images, audio and video into one shared embedding space and runs locally, so searching a media library no longer needs the cloud.

The EmbeddingGemma 2 wordmark in white and blue on a dark geometric background

Google has released EmbeddingGemma 2, an on-device embedding model that goes beyond the text-only original by mapping code, images, video and audio into a single shared vector space. It has 740 million parameters, sits under the Gemma 4 architecture and ships under the Apache 2.0 licence.

The first EmbeddingGemma was downloaded more than 20 million times for on-device search and privacy-first retrieval-augmented generation. The sequel keeps that job but widens it: a single model can now find a specific video clip from a voice memo, or search hours of recorded audio from a text query, without sending either off the device.

What changed in the second generation

Google says the model leads sub-1B multimodal embedders on benchmarks including MTEB Code and the Massive Audio Embedding Benchmark, and that it matches or beats some specialist models more than twice its size. Multilingual text performance carries over from the first version, while code retrieval improves by 9.92 points on MTEB Code, from 68.76 to 78.68.

Multi-modal support is modular rather than monolithic. Text-only work can run on as little as 270 million parameters, with optional vision and audio encoders adding 170 million and 300 million more. Using Matryoshka Representation Learning, developers can truncate output vectors from 768 dimensions down to 512, 256 or 128, which Google pitches as up to a sixfold saving in local vector storage.

How little memory it needs

With quantisation, Google quotes roughly 191MB of active RAM for the text-only weights on a Pixel 11 Pro and about 567MB for the full multimodal model. The context window has quadrupled to 8,000 tokens, enough for around five and a half minutes of audio, 29 images, 58 video frames, or interleaved combinations of all three.

Weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability promised later. Google names MediaPipe and LiteRT for on-device deployment, transformers.js and WebGPU for the browser, and the usual serving stacks including llama.cpp, Ollama and vLLM, with Qdrant for vector storage.

Our opinion

The interesting number here is not the benchmark delta, it is the 191MB. Embeddings are the cheapest useful thing a model does, and until now the practical way to get good ones was to call a hosted endpoint, which quietly made every local search feature a network feature. A sub-gigabyte embedder that handles audio and images changes what a phone, a laptop or a Raspberry Pi can index on its own.

Google also has an obvious interest in making on-device embedding good: it keeps Google's models on the device next to Gemma 4, and the shared tokenizer means the two run as one pipeline with a smaller combined memory bill. That is a sensible engineering choice and also a competitive one, since the same weights are Apache 2.0 licensed and can be dropped into a stack that owes nothing to Google.