Embeddings

An embedding is a list of numbers that stands in for a piece of text, arranged so that text with similar meaning gets similar numbers. Once meaning has a position, finding related text becomes a matter of finding nearby points.

reset my password forgot my login refund policy get my money back
Same meaning lands together, even with almost no shared words. Real embeddings have hundreds or thousands of dimensions — separate numbers describing each piece of text — and left/right and up/down here stand in for all of them.

Meaning as a position

Feed "reset my password" to an embedding model and you get back something like [0.021, -0.118, 0.334, …] — anywhere from a few hundred to a few thousand numbers. On its own that list tells you nothing. It only means something compared against other lists.

Feed in "I forgot my login" and you get a different list that turns out to sit very close to the first. Feed in "what's your refund policy" and you get one that's far away. The model was trained on enormous amounts of text to produce exactly that arrangement: things people use in similar contexts end up in similar places. The two password phrases share almost no words, so keyword search would treat them as unrelated.

How closeness gets measured

Nearness has to become a number. The usual measure is cosine similarity — the cosine of the angle between two vectors, meaning the two lists of numbers treated as points in space. It runs from 1, pointing the same direction, through 0, unrelated, down to −1, pointing opposite. In practice unrelated text from a real model rarely scores near 0; what you care about is the ranking, not the absolute figure.

Which measure you use matters less than people expect. Cosine similarity, dot product and straight-line distance rank things in much the same order, provided the vectors are normalised — scaled to a standard length, which most embedding services do for you. Unnormalised, dot product favours longer vectors and can rank quite differently. What matters far more than the metric is which model produced the numbers.

What they're used for

The big one is search that works on meaning rather than wording — the retrieval step inside retrieval-augmented generation, or RAG. The same positioning also drives grouping, near-duplicate detection, categorising and recommendations. Embed every incoming support ticket, for instance, and you can flag any new one that lands almost on top of a ticket you've already answered.

Where embeddings mislead you

Similar isn't relevant

Embeddings capture what a passage is about. "I love this product" and "I hate this product" sit close together, because they're on the same topic with opposite sentiment. A retrieval system ranking purely by similarity will hand back passages that are clearly on-topic and answer nothing — and it will do it confidently, because by its own measure those passages scored well.

Exact strings are a weak spot

Order numbers, SKUs, error codes, rare surnames, unusual acronyms. A model trained on meaning has little to work with when a string doesn't really carry any — a part number is just characters, and nothing about it sits semantically near anything else.

Long inputs lose their detail

Two things bite here. Embedding models cap their input, commonly at somewhere between a few hundred and a few thousand tokens — chunks of text, usually a word or part of one — and anything past the cap is simply cut off rather than summarised. Below the cap, the output vector is a fixed size whether you embed one sentence or ten pages, so a long passage produces its overall gist and a specific fact buried inside it stops standing out. Both are reasons focused, smaller chunks retrieve better than whole documents.

Use one model on both sides

Vectors from two different embedding models cannot be compared. They're coordinates in different spaces, and the numbers mean different things in each. Index your documents with one model — embed them and store the vectors — then embed queries with another, and you don't get slightly worse results. You get meaningless ones.

The consequence is operational. Switching embedding models means re-embedding everything already stored, which on a large collection costs real time and money. Worth weighing before you choose one rather than after.

Dimensions — how many numbers each vector holds — vary by model, from a few hundred to a few thousand. More isn't automatically better: higher dimensions mean more storage and slower search, and several current models are deliberately trained so you can cut the vector short and lose very little accuracy, a trick usually called Matryoshka representation learning.

None of this arrived with modern chatbots. Placing whole sentences in a shared space goes back at least to Sentence-BERT in 2019, which made high-quality sentence vectors cheap enough to compare at scale, and to sentence encoders before that. What changed is that the models got good enough, and cheap enough, to build products on.

In this guide
  1. Meaning as a position
  2. How closeness gets measured
  3. What they're used for
  4. Where embeddings mislead you
  5. Use one model on both sides
  6. FAQ

FAQ

Is an embedding model the same as the model that writes the answers?

No. An embedding model only converts text into numbers — it can't hold a conversation or write anything. It's usually much smaller and cheaper to run than a generative model, which is what makes it practical to run over every document you own, though some of the strongest embedding models are themselves built from large ones.

Practice interview questions on Embeddings →