Most people reading this have an archive problem. Years of work — images, project files, video exports, audio stems, sketches, notes — organised by folder and date, searchable only by filename. You know the piece exists. You cannot find it.
The reason this stays unsolved is that embedding models are per-modality. CLIP does images and text. Whisper-derived models do audio. A sentence transformer does text. Each produces vectors in its own space, so you cannot ask which of my recordings sounds like this photograph feels, because the two vectors are not comparable.
EmbeddingGemma 2, from Google DeepMind, puts them in one space.
What it is
An open multimodal embedding model that converts text (including code), images, video and audio — plus combinations — into a unified 768-dimensional vector space, for semantic search, RAG, classification, clustering, similarity and fact verification.
740M total parameters, and the architecture is the interesting part:
| Component | Size |
|---|---|
| Text backbone | 270M |
| Vision encoder | 170M, modular |
| Audio encoder | 300M, modular |
24 layers, 512 model dimension, GQA/MQA attention. 8,192 tokens of context shared across all modalities. 100+ languages. Apache 2.0.
Reported benchmarks: 61.36 on MTEB multilingual, 78.68 on code retrieval (NDCG@10), with results across vision, video and audio benchmarks.
Selective encoder loading is the feature that makes it practical
The vision and audio encoders are modular and load only if you need them.
That matters more than it sounds. A text-only task loads 270M parameters, not 740M. Add images and you are at 440M. The full stack is only paid for when you are actually embedding sound.
For anyone deploying on real hardware — a laptop, a Jetson, an installation machine — this is the difference between a model you can run and one you plan around. It is the same economy as laya’s 421M: the useful models for this work are the ones that fit.
Matryoshka, which you should use
Matryoshka Representation Learning means the 768-dimensional output can be truncated to 512, 256 or 128 dimensions and still work. Google quotes up to 6x vector storage reduction.
The trick is in the training: the model is optimised so that the first N dimensions of the vector are themselves a valid embedding. You are not compressing afterwards — you are reading a prefix.
Why that is a bigger deal than a storage saving: vector search cost scales with dimensionality. A 128-dimensional index is roughly six times cheaper to store and substantially faster to query than a 768-dimensional one. For an archive of a hundred thousand items on a modest machine, that is the difference between instant and sluggish.
The standard pattern is two-stage retrieval: search the 128-dimensional index to get a few hundred candidates fast, then re-rank those candidates with the full 768-dimensional vectors. You get cheap recall and accurate precision.
What you would actually build
Cross-modal archive search. Embed everything you have made into one index. Then query it with anything — a reference image, a hummed melody, a text description, a video clip — and get back the nearest items regardless of their type. This is the thing that has not been possible without stitching several models together and giving up on comparing across them.
Finding your own work again. Embed a folder of ten thousand sketches and query with a rough drawing. No tagging, no filenames, no metadata discipline required retroactively.
Clustering an archive to see its shape. Embed, reduce to 2D with UMAP, and look at what groups. This is a genuinely good way to discover structure in your own output that you did not know was there — recurring motifs, periods, abandoned directions.
Live similarity in an installation. 740M is small enough to embed a camera frame or a microphone buffer in real time, and compare it against a precomputed index. A piece that responds to what the room resembles from a corpus you chose is a different proposition from one that responds to motion.
And retrieval for generation. The unglamorous but most immediately useful case: a local RAG index over your own notes, references and documentation.
The practical notes
Embeddings are not interchangeable between models or versions. If you index an archive with EmbeddingGemma 2 and later switch models, you re-embed everything. Choose with that in mind, and keep the source files.
Normalise and pick your metric deliberately. Cosine similarity on L2-normalised vectors is the usual default and is what most benchmarks assume.
Test the cross-modal claim on your own material before trusting it. A unified space means the model can compare an image to a sound; whether its notion of similarity matches yours for your particular corpus is an empirical question. Embed fifty items you know well and look at the nearest neighbours.
And sentence-transformers support is in the tags, which means the integration path is the short one.
If embeddings are new to you, our semantic search primer goes from a folder of files to a working index.