Embeddings Explained, With Code
What an embedding vector actually is, how to compute two and compare them in Python, and which model and dimension to pick.
On this page3 sections
An embedding is a list of floats that stands in for a piece of text’s meaning. Texts that mean similar things get vectors pointing in similar directions. That’s the whole idea — everything else is arithmetic.
Compute two and compare them
import numpy as np
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY
def embed(text, model="text-embedding-3-small"):
r = client.embeddings.create(input=text, model=model)
return np.array(r.data[0].embedding)
def cosine(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
a = embed("How do I reset my password?")
b = embed("I forgot my login details.")
c = embed("The mitochondria is the powerhouse of the cell.")
print(cosine(a, b)) # high
print(cosine(a, c)) # low
Two API calls and a dot product. Cosine similarity ranges from -1 to 1 and ignores vector length, which is why it beats raw distance for text.
Read the number as a ranking, not a score. There is no universal “0.8 means related” threshold — it shifts per model and per corpus. Embed 20 pairs you have opinions about, look at where the line actually falls for your data, and set the cutoff from that.
Which model, which dimension
| Model | Default dims | Max input | Reach for it when |
|---|---|---|---|
text-embedding-3-small | 1536 | 8192 tokens | Default. Cheap, good enough for most search. |
text-embedding-3-large | 3072 | 8192 tokens | Recall matters and storage doesn’t. |
voyage-4-lite | 1024 | 32,000 tokens | Latency and cost, long chunks. |
voyage-4-large | 1024 | 32,000 tokens | Best general-purpose retrieval quality. |
voyage-code-4 | 1024 | 32,000 tokens | Code and repo search. |
Both families let you shorten the vector: OpenAI takes a dimensions argument, Voyage takes output_dimension (256, 512, 1024, or 2048). Halving dimensions halves your index size and speeds up search, at some cost to recall — measure it on your own queries rather than guessing.
Two rules that save re-indexing later: never mix models in one index (vectors from different models are not comparable), and store the model name and dimension alongside the vectors so you know what you’d be re-running.
What embeddings are bad at
Negation (“refund allowed” vs “refund not allowed” score high together), exact identifiers like SKUs or error codes, and freshness — the vector has no idea which document is newer. Pair semantic search with a keyword filter for anything with IDs in it.