AI & Creative Tools

Semantic Search Over Your Own Archive, From Zero

Embed a folder of files into vectors, put them in an index, and query by meaning instead of filename — locally, in about forty lines.

You have years of work on a drive and no way to find anything in it. Filenames are inconsistent, tags were never applied, and full-text search does not help because the thing you are looking for is an image, or a sound, or a description you can only paraphrase.

Embeddings fix this, and the core idea is simpler than the terminology suggests.

AssemblyAI’s four-minute explanation of embeddings and indexes.

The idea

An embedding is a list of numbers that represents a thing’s meaning, produced by a model trained so that similar things get similar numbers.

That is it. A 768-dimensional embedding is a point in 768-dimensional space, and “find things like this” becomes “find the nearest points” — which is a geometry problem computers are extremely good at.

The consequence that matters: you search by meaning, not by string. Query a text index with “that sketch of a bridge at night” and it returns relevant items even if none of them contains the words bridge or night.

The minimum working version

pip install sentence-transformers numpy
from sentence_transformers import SentenceTransformer
import numpy as np, pathlib, json

model = SentenceTransformer("google/embeddinggemma-2")   # or any embedding model

# 1. gather your documents
paths = list(pathlib.Path("notes").rglob("*.md"))
docs  = [p.read_text(errors="ignore") for p in paths]

# 2. embed them once
vecs = model.encode(docs, normalize_embeddings=True, batch_size=16,
                    show_progress_bar=True)
np.save("vecs.npy", vecs)
json.dump([str(p) for p in paths], open("paths.json", "w"))

# 3. query
def search(q, k=5):
    qv = model.encode([q], normalize_embeddings=True)[0]
    scores = vecs @ qv                 # cosine similarity, because normalised
    for i in np.argsort(-scores)[:k]:
        print(f"{scores[i]:.3f}  {paths[i]}")

search("the piece about tidal data and sound")

That is a working semantic search engine. vecs @ qv is a matrix multiply; with normalised vectors it is cosine similarity, and for anything under roughly 100,000 items it is fast enough that you do not need a database.

The chunking decision, which matters more than the model

Do not embed whole documents. This is the single most common mistake.

An embedding is a fixed-size vector, so embedding a 10,000-word document compresses everything it says into one point — and that point ends up in the average of all its topics, matching nothing specifically. A document about three subjects becomes a vector near none of them.

Chunk instead. Split into passages of roughly 200–500 words with a little overlap, embed each chunk, and keep a map from chunk back to source:

def chunk(text, size=400, overlap=60):
    words = text.split()
    return [" ".join(words[i:i+size])
            for i in range(0, max(1, len(words)-overlap), size-overlap)]

Then a search returns the passage and you can show the file it came from. The overlap stops a relevant sentence being split across a boundary and lost.

For images, audio and video there is no chunking question for stills, and there is for time-based media. Embed a video as a handful of keyframes plus audio segments rather than as one vector, for exactly the same reason.

Picking a model

For text: anything in the sentence-transformers ecosystem works. all-MiniLM-L6-v2 is tiny, fast and a perfectly good baseline. The current generation is substantially better at multilingual and code.

For mixed media: you need a model with a shared space across modalities, or you cannot compare an image to a text query. EmbeddingGemma 2 covers text, code, images, video and audio in one 768-dimensional space at 740M parameters, which is the current practical answer for an archive that is not all one type. CLIP remains the standard for image-and-text only.

Three things to check before you commit, because switching later means re-embedding everything:

  • Does it cover your modalities?
  • Does it cover your languages?
  • Does it support Matryoshka truncation? If it does, you can store 128-dimensional vectors instead of 768 and search six times cheaper. Use it.

When to add a real index

Brute force is O(n) per query. It stops being instant somewhere around 100,000 to a million vectors depending on your machine.

Past that, use an approximate nearest neighbour index:

  • faiss — the standard, from Meta. IndexHNSWFlat is a good default.
  • hnswlib — smaller dependency, excellent, does one thing.
  • chromadb or lancedb — add metadata filtering and persistence, which you will want.
  • sqlite-vec — a vector index inside SQLite, which for a local tool is a very pleasant answer.

You do not need a hosted vector database for a personal archive, and a surprising amount of writing on this subject assumes you do.

Five things that will catch you out

1. Embeddings are model-specific and version-specific. Vectors from two different models are not comparable, at all. Record which model and which version produced your index, in the index.

2. Normalise, and then use dot product. If your vectors are L2-normalised, dot product is cosine similarity and is faster. Mixing normalised and unnormalised vectors silently produces nonsense.

3. Similarity is not relevance. The nearest neighbour to “red” may be “blue” — both are colours, and the model’s notion of similarity is relatedness, not sameness. Look at your top-10 results for twenty known queries before trusting the index for anything.

4. Short queries behave differently from long ones. A two-word query and a paragraph land in different regions of the space. If users will type short queries, test with short queries.

5. Keep the source files and the chunk offsets. An index of vectors with no way back to the original text is a lovely mathematical object and useless. Store the path and the character range.

What this unlocks

Finding your own work. The honest primary use, and it is worth the afternoon on its own.

Clustering to see the shape of an archive. Embed everything, reduce to 2D with UMAP, and plot it. The groups that appear are frequently not the ones you would have chosen as folders, and that is the interesting part.

Local RAG. Retrieve the top chunks for a question and pass them to a local language model. The retrieval is the hard part and you have just built it.

And live similarity in a piece. Precompute an index over a corpus, embed a camera frame or an audio buffer at runtime, and respond to what the room most resembles.