Columns · · Team LATENT
Search photos and audio by meaning: what EmbeddingGemma 2 brings to local retrieval
EmbeddingGemma 2 puts text, photos, audio and video into a shared representation for search. A guide to its role alongside generative AI, the steps in a local index, its model footprint and the trade-offs in shorter vectors.
You want to find a photo taken at a station on a rainy day, but cannot remember its filename or date. Photos, recordings and notes may live in different apps. Search by meaning can connect the words you remember with relevant material even when you have forgotten how it was filed.
Google announced EmbeddingGemma 2 on October 6, 2026, as a model for building capabilities such as on-device semantic search. Its role differs from that of a conversational generator: it turns text, code, images, video and audio into numerical representations that can be compared. Representing a query and stored material in the same space makes retrieval possible.
Key takeaways
- EmbeddingGemma 2 maps different kinds of material into a shared 768-dimensional representation for retrieval. A generative LLM can be added separately when an application needs a written answer.
- The 740M-parameter model comprises 270M for text, 170M for vision and 300M for audio. Unused encoders can be omitted to match the material being processed.
- Vectors can be shortened to 512, 256 or 128 dimensions, but multimodal quality falls substantially at 128. Retrieval quality and speed still need testing with the actual material and device configuration.
Retrieval and generation do different jobs
An embedding represents characteristics of content as a numerical vector. EmbeddingGemma 2 normally outputs 768 values. An application can encode a query and a photo, compare their vector directions, and rank similar candidates higher. People do not need to read the individual values; the search software uses them.
FIGURE 01
Make different material comparable
To search photos with words, encode both the query and the material.
Text and code
Photos and page images
Audio recordings
Video moments
Shared 768-dimensional vectors
Retrieve similar candidates
Open the original image, audio or document
The immediate output is a set of likely matches. Finding and opening the point in a recording where a deadline was discussed can be useful without generating a long answer. To turn retrieved material into an explanation, an application can pass it to a separate generative LLM. That retrieval-plus-generation arrangement is the basis of RAG.
High similarity does not establish that a source is true. A relevant passage might contain an old decision or a rejected proposal. Showing the original material, its date and the matching location gives people a way to check the result.
Source: EmbeddingGemma 2 model card
Search across photos, recordings and code
A useful starting point is a kind of material you often lose track of: a scene absent from a filename, a topic inside a recording, or a function whose name you do not know. Multimodal retrieval lets the query and stored item take different forms.
| Material | Example query | Useful information in a result |
|---|---|---|
| Photos | A person holding an umbrella on a station platform | Original image and capture information |
| Meeting audio | The part where the deadline was changed | Playback position and surrounding discussion |
| Video | The part demonstrating a procedure | Timestamp and original video |
| Code | Code that removes expired sessions | Filename and nearby code |
The queries in the table illustrate use cases; they are not LATENT accuracy tests. Google’s AI Edge Gallery examples include searching local media with language or images and locating moments in videos. One search interface across types of material can reduce the need to look separately in several apps.
Source: Google AI Edge retrieval demos and implementation explanation
Prepare an index, then search it
Putting the model on a device is only one part of a search application. The app must read the selected material, divide it into useful units, create embeddings and store them. Paragraphs or sections can work for documents; timed segments can work for recordings. The units should let a result point back to its original location.
FIGURE 02
Separate indexing from each search
Store document vectors ahead of time, then encode the query when a search is made.
Before searching
Split material and retain its location
Encode each part
Store vectors and locations
At search time
Encode the search input
Compare with stored vectors
Show matching source material
Optional: a separate generative LLM
Pass retrieved material to a generator for an answer. Retrieval alone does not require this step.
Google’s example stores vectors in a local SQLite database and retrieves items using cosine similarity, a comparison of vector direction. Search does not inherently require a large cloud backend. Storage and retrieval methods can be chosen to fit the collection and required response time.
Initial indexing and searching an existing index are different workloads. New files need new index entries; deleted material needs to disappear from the stored index as well. A small search model still needs an application that maintains the collection over time.
Load the input encoders you need
740M describes the approximate total parameter count: 270M for text, 170M for vision and 300M for audio. Unused modality encoders can be omitted. A text-only search workload does not need to load the vision and audio components.
FIGURE 03
Configure the model for the material
Published parameter counts. M means million; these figures are not memory or speed measurements.
270M
Text
Core for text and code
170M
Vision
Include for visual input
300M
Audio
Include for audio input
Text only: 270M / Text + vision: 440M / Text + audio: 570M / Full model: 740M
Source: component sizes on the official model distribution page
Google also reports active RAM figures of about 191MB for text-only weights and 567MB for the full model in a quantized configuration on Pixel 11 Pro. Those are figures for a particular published configuration, not total application memory including image loading, video decoding, the search database and the app itself.
To judge whether a device is suitable, consider the material processed at once and the index size as well as model weights. Initial video indexing time, query response and battery use deserve separate checks. This article does not report device speed or power measurements.
Source: Google announcement and reported Pixel 11 Pro memory figures
What actually becomes one-sixth the size?
Matryoshka Representation Learning, or MRL, allows the 768-dimensional output to be shortened to 512, 256 or 128 dimensions. At the same numeric data type, 128 values take one-sixth the raw space of 768 values. That reduction concerns vector elements, not the entire model or the original photos.
FIGURE 04
Shorter stored vectors
Element count relative to 768 dimensions. Simple arithmetic for storage at the same numeric data type.
768
100%
512
66.7%
256
33.3%
128
16.7%
At 128 dimensions, multimodal retrieval quality declines substantially. Validate quality after shortening the vectors.
The official documentation positions 128 dimensions mainly for text workloads and warns about a substantial multimodal quality loss. For cross-modal search, a practical approach is to establish a 768-dimensional baseline and compare 512 or 256 using the same queries. The smallest supported representation should not automatically become the default.
Check whether the wanted source remains near the top and whether plausible but wrong alternatives displace it. Saving storage does not help the experience if people have to repeat their searches. The appropriate dimension depends on the collection and the kinds of matches the application must retain.
Source: truncation evaluation and guidance on 128 dimensions
Keep the retrieval pipeline consistent
For text retrieval, queries and stored documents use short task-specific prefixes. The official guide distinguishes search queries, documents and code retrieval, among other tasks. These prefixes identify the role of a text input; the same text prefixes are not prepended to image or audio input.
Stored material and queries need compatible embeddings from the same model and configuration, with matching dimensions. A 768-dimensional query cannot be compared directly with a 128-dimensional index. After truncation, L2-normalize the vector again before similarity scoring; normalization before slicing does not preserve its length after slicing.
For inference with the original checkpoint in tools such as Sentence Transformers, the official documentation specifies bfloat16 or float32 and warns against float16, which can produce NaNs or degraded embeddings. Precision should not be selected simply because it is described as half precision. This also needs to be distinguished from separately distributed quantized on-device configurations.
Source: official Sentence Transformers inference guide
An 8,192-token context does not allow unlimited audio or images in one request. Google describes interleaved inputs, but their components share the same context budget. Dividing material according to the location precision needed in search results makes the source easier to verify.
Choose a useful first scope
Start with a small collection and queries whose expected results you already know. Identify the photo or recording segment you actually want to retrieve. Add alternative phrasings and similar-but-wrong candidates to test usefulness beyond a persuasive demonstration.
For retrieval without sending material outside the device, keep embedding generation, index storage and similarity calculation local. If summaries become necessary, decide separately where a generative LLM should run. When users can read the retrieved material themselves, a large conversational model is not required at the outset.
Availability also needs to be separated from planned delivery. The model is distributed through official channels including Hugging Face. The AI Edge article describes an Android service accessible through ML Kit as coming in the following weeks. As of October 8, that announcement should not be treated as an already available service.
The model is labeled Apache 2.0, while its model card also says deployments must adhere to the Gemma Prohibited Use Policy. The word “open” alone is not a complete statement of the published conditions; check both the license text and the conditions referenced by the model card when adopting it.
Source: Apache License 2.0 text
Source: model card usage conditions and limitations
EmbeddingGemma 2 makes meaning-based retrieval over personal material more approachable. Choosing what to find and where a result should take the user makes the role of a small local model concrete.
Sources and scope
Explanation based on primary sources checked as of October 8, 2026. Parameter counts, supported dimensions and device active RAM are Google’s published figures, not LATENT measurements. The dots illustrate a concept; vector bars show element-count arithmetic. No model download, installation or device benchmark was performed.
Google launch announcement, October 6, 2026
Google’s official model distribution page