AI news, with the context that matters.

Columns · ·

Search photos and audio by meaning: what EmbeddingGemma 2 brings to local retrieval

EmbeddingGemma 2 puts text, photos, audio and video into a shared representation for search. A guide to its role alongside generative AI, the steps in a local index, its model footprint and the trade-offs in shorter vectors.

You want to find a photo taken at a station on a rainy day, but cannot remember its filename or date. Photos, recordings and notes may live in different apps. Search by meaning can connect the words you remember with relevant material even when you have forgotten how it was filed.

Google announced EmbeddingGemma 2 on October 6, 2026, as a model for building capabilities such as on-device semantic search. Its role differs from that of a conversational generator: it turns text, code, images, video and audio into numerical representations that can be compared. Representing a query and stored material in the same space makes retrieval possible.

Key takeaways

  1. EmbeddingGemma 2 maps different kinds of material into a shared 768-dimensional representation for retrieval. A generative LLM can be added separately when an application needs a written answer.
  2. The 740M-parameter model comprises 270M for text, 170M for vision and 300M for audio. Unused encoders can be omitted to match the material being processed.
  3. Vectors can be shortened to 512, 256 or 128 dimensions, but multimodal quality falls substantially at 128. Retrieval quality and speed still need testing with the actual material and device configuration.

Retrieval and generation do different jobs

An embedding represents characteristics of content as a numerical vector. EmbeddingGemma 2 normally outputs 768 values. An application can encode a query and a photo, compare their vector directions, and rank similar candidates higher. People do not need to read the individual values; the search software uses them.

FIGURE 01

Make different material comparable

To search photos with words, encode both the query and the material.

Text and code

Photos and page images

Audio recordings

Video moments

Shared 768-dimensional vectors

Retrieve similar candidates

Open the original image, audio or document

Figure 1. The dots are illustrative, not model outputs or measurements. Actual vectors have 768 dimensions; search does not use the two-dimensional positions shown here.

The immediate output is a set of likely matches. Finding and opening the point in a recording where a deadline was discussed can be useful without generating a long answer. To turn retrieved material into an explanation, an application can pass it to a separate generative LLM. That retrieval-plus-generation arrangement is the basis of RAG.

High similarity does not establish that a source is true. A relevant passage might contain an old decision or a rejected proposal. Showing the original material, its date and the matching location gives people a way to check the result.

Source: EmbeddingGemma 2 model card

Search across photos, recordings and code

A useful starting point is a kind of material you often lose track of: a scene absent from a filename, a topic inside a recording, or a function whose name you do not know. Multimodal retrieval lets the query and stored item take different forms.

MaterialExample queryUseful information in a result
PhotosA person holding an umbrella on a station platformOriginal image and capture information
Meeting audioThe part where the deadline was changedPlayback position and surrounding discussion
VideoThe part demonstrating a procedureTimestamp and original video
CodeCode that removes expired sessionsFilename and nearby code

The queries in the table illustrate use cases; they are not LATENT accuracy tests. Google’s AI Edge Gallery examples include searching local media with language or images and locating moments in videos. One search interface across types of material can reduce the need to look separately in several apps.

Source: Google AI Edge retrieval demos and implementation explanation

Prepare an index, then search it

Putting the model on a device is only one part of a search application. The app must read the selected material, divide it into useful units, create embeddings and store them. Paragraphs or sections can work for documents; timed segments can work for recordings. The units should let a result point back to its original location.

FIGURE 02

Separate indexing from each search

Store document vectors ahead of time, then encode the query when a search is made.

Before searching

Split material and retain its location

Encode each part

Store vectors and locations

At search time

Encode the search input

Compare with stored vectors

Show matching source material

Optional: a separate generative LLM

Pass retrieved material to a generator for an answer. Retrieval alone does not require this step.

Figure 2. Local embedding and retrieval do not make a later cloud-generation step local. Check the full application to understand which data remains on the device.

Google’s example stores vectors in a local SQLite database and retrieves items using cosine similarity, a comparison of vector direction. Search does not inherently require a large cloud backend. Storage and retrieval methods can be chosen to fit the collection and required response time.

Initial indexing and searching an existing index are different workloads. New files need new index entries; deleted material needs to disappear from the stored index as well. A small search model still needs an application that maintains the collection over time.

Load the input encoders you need

740M describes the approximate total parameter count: 270M for text, 170M for vision and 300M for audio. Unused modality encoders can be omitted. A text-only search workload does not need to load the vision and audio components.

FIGURE 03

Configure the model for the material

Published parameter counts. M means million; these figures are not memory or speed measurements.

270M

Text

Core for text and code

170M

Vision

Include for visual input

300M

Audio

Include for audio input

Text only: 270M / Text + vision: 440M / Text + audio: 570M / Full model: 740M

Figure 3. How to omit an encoder depends on the library. A reduction in parameter count does not imply the same proportional reduction in total application memory.

Source: component sizes on the official model distribution page

Google also reports active RAM figures of about 191MB for text-only weights and 567MB for the full model in a quantized configuration on Pixel 11 Pro. Those are figures for a particular published configuration, not total application memory including image loading, video decoding, the search database and the app itself.

To judge whether a device is suitable, consider the material processed at once and the index size as well as model weights. Initial video indexing time, query response and battery use deserve separate checks. This article does not report device speed or power measurements.

Source: Google announcement and reported Pixel 11 Pro memory figures

What actually becomes one-sixth the size?

Matryoshka Representation Learning, or MRL, allows the 768-dimensional output to be shortened to 512, 256 or 128 dimensions. At the same numeric data type, 128 values take one-sixth the raw space of 768 values. That reduction concerns vector elements, not the entire model or the original photos.

FIGURE 04

Shorter stored vectors

Element count relative to 768 dimensions. Simple arithmetic for storage at the same numeric data type.

768

100%

512

66.7%

256

33.3%

128

16.7%

At 128 dimensions, multimodal retrieval quality declines substantially. Validate quality after shortening the vectors.

Figure 4. This is not a speed or accuracy benchmark. Database overhead and source material remain separate; total application storage need not fall by these ratios.

The official documentation positions 128 dimensions mainly for text workloads and warns about a substantial multimodal quality loss. For cross-modal search, a practical approach is to establish a 768-dimensional baseline and compare 512 or 256 using the same queries. The smallest supported representation should not automatically become the default.

Check whether the wanted source remains near the top and whether plausible but wrong alternatives displace it. Saving storage does not help the experience if people have to repeat their searches. The appropriate dimension depends on the collection and the kinds of matches the application must retain.

Source: truncation evaluation and guidance on 128 dimensions

Keep the retrieval pipeline consistent

For text retrieval, queries and stored documents use short task-specific prefixes. The official guide distinguishes search queries, documents and code retrieval, among other tasks. These prefixes identify the role of a text input; the same text prefixes are not prepended to image or audio input.

Stored material and queries need compatible embeddings from the same model and configuration, with matching dimensions. A 768-dimensional query cannot be compared directly with a 128-dimensional index. After truncation, L2-normalize the vector again before similarity scoring; normalization before slicing does not preserve its length after slicing.

For inference with the original checkpoint in tools such as Sentence Transformers, the official documentation specifies bfloat16 or float32 and warns against float16, which can produce NaNs or degraded embeddings. Precision should not be selected simply because it is described as half precision. This also needs to be distinguished from separately distributed quantized on-device configurations.

Source: official Sentence Transformers inference guide

An 8,192-token context does not allow unlimited audio or images in one request. Google describes interleaved inputs, but their components share the same context budget. Dividing material according to the location precision needed in search results makes the source easier to verify.

Choose a useful first scope

Start with a small collection and queries whose expected results you already know. Identify the photo or recording segment you actually want to retrieve. Add alternative phrasings and similar-but-wrong candidates to test usefulness beyond a persuasive demonstration.

For retrieval without sending material outside the device, keep embedding generation, index storage and similarity calculation local. If summaries become necessary, decide separately where a generative LLM should run. When users can read the retrieved material themselves, a large conversational model is not required at the outset.

Availability also needs to be separated from planned delivery. The model is distributed through official channels including Hugging Face. The AI Edge article describes an Android service accessible through ML Kit as coming in the following weeks. As of October 8, that announcement should not be treated as an already available service.

The model is labeled Apache 2.0, while its model card also says deployments must adhere to the Gemma Prohibited Use Policy. The word “open” alone is not a complete statement of the published conditions; check both the license text and the conditions referenced by the model card when adopting it.

Source: Apache License 2.0 text

Source: model card usage conditions and limitations

EmbeddingGemma 2 makes meaning-based retrieval over personal material more approachable. Choosing what to find and where a result should take the user makes the role of a small local model concrete.

Sources and scope

Explanation based on primary sources checked as of October 8, 2026. Parameter counts, supported dimensions and device active RAM are Google’s published figures, not LATENT measurements. The dots illustrate a concept; vector bars show element-count arithmetic. No model download, installation or device benchmark was performed.

Google launch announcement, October 6, 2026

EmbeddingGemma 2 model card

Google’s official model distribution page

Google AI Edge implementation and availability plans

Sentence Transformers inference guide