Skip to content
The AI-First Web

Google's EmbeddingGemma 2 puts video and audio search on the device

Google says the full model needs about 567MB of RAM when quantized on a Pixel 11 Pro. Its index is data to govern.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Google's EmbeddingGemma 2 puts video and audio search on the device

Photo: Ayush Bali / Pexels

In brief
  • Google DeepMind released EmbeddingGemma 2, an open model that searches text, code, images, audio and video, and that it says can run on consumer devices.
  • Local search creates an index of a device's media and files. That index is a new store of data needing an owner and rules.
  • Ask which devices index content locally, how the vectors are protected, and what happens to the index when a device is lost or an employee leaves.

Picture using a voice memo to find a specific video clip, or typing a text query to search hours of audio recordings. Google DeepMind says its new model can do both on consumer hardware, and that this search can work entirely offline. That shifts a question for leaders: who holds the index of everything a device can find?

Our view is that on-device search turns the index into an asset. When search moves to the device, the search index becomes a new store of data about your content. It needs an owner, a policy and a review, like any other store.

What Google released

On October 6, 2026, Google DeepMind launched EmbeddingGemma 2. It is an embedding model, which turns content into lists of numbers that capture meaning. DeepMind says it maps text, code, images, audio and video into one shared space.

The model shares its foundations with Gemma 4. Under the Apache 2.0 license, which DeepMind calls commercially permissive, companies can build products with it. It has 740 million parameters. DeepMind says the first EmbeddingGemma, released last year, passed 20 million downloads.

20 million+
Downloads of the first EmbeddingGemma
Source: Google DeepMind (October 6, 2026)

How it works

An embedding places each item as a point in a space. Items with similar meaning land close together. A shared space means a text question and a video clip can be compared directly. That is how a spoken query can find a moment in a video.

DeepMind describes the model as modular. Text-only work needs 270 million parameters. A 170 million parameter vision encoder and a 300 million parameter audio encoder are optional additions for full multimodal use.

The model also shrinks its output. Using a technique called Matryoshka Representation Learning, developers can cut each vector from 768 numbers to 512, 256 or 128. DeepMind says this saves up to 6x on storage for local vector databases.

Up to 6x
Storage reduction from shorter vectors
Source: Google DeepMind (October 6, 2026)

On a Google Pixel 11 Pro, with quantization, DeepMind reports about 191MB of active RAM for text-only weights. The full multimodal model needs about 567MB. Quantization stores the model's numbers at lower precision to save memory.

~567MB
Active RAM, full multimodal model, quantized, Pixel 11 Pro
Source: Google DeepMind (October 6, 2026)

What the numbers show, and what they do not

DeepMind reports a 9.92-point gain on MTEB Code, a code-search benchmark. The score rose from 68.76 to 78.68. The company says multilingual text performance matches the first model.

68.76 to 78.68
MTEB Code score, EmbeddingGemma to EmbeddingGemma 2
Source: Google DeepMind (October 6, 2026)

These are the vendor's own results. DeepMind adds that on image, video, document and audio tasks, it beats some specialist models with more than double its parameters. The full evaluation sits in a model card that this story has not reviewed. Test the model on your own files before you plan around it.

The index is the new asset

DeepMind says generating embeddings locally helps ensure data privacy and reduces pipeline latency. That is a real benefit. Content does not have to leave the device to become searchable.

It also creates something new. Every device that runs this kind of search holds an index of its own media, files and recordings. The source does not say how those indexes are protected. That gap is where leaders should ask questions.

The comparison is a library's card catalog. The catalog is smaller than the books, but it tells a reader what the library holds and where to look. A vector index plays that role for a phone or laptop.

What to ask your team

First, which devices or apps in your organisation will index audio, video or files locally, and who approved it? Second, where are the vectors stored, and are they covered by the same encryption and retention rules as the source files? Third, if an employee leaves or a device is lost, what happens to the index?

Fourth, have engineers test the model on your own content and languages, not only on published benchmarks. Finally, check the Apache 2.0 terms against your legal team's open-source policy.

Private search is only as private as the index you keep. Treat that index as data, not as a feature.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Google DeepMind.

Share this insight