Retrieval & data
Cosine similarity
Definition
Cosine similarity measures how closely two vectors point in the same direction, on a scale from -1 to 1. It is the standard way to compare embeddings, because it captures semantic similarity while ignoring text length.
The metric is the cosine of the angle between two vectors. A value of 1 means identical direction, 0 means unrelated, and -1 means opposite. Because it depends only on direction and not magnitude, a short paragraph and a long article on the same topic score as similar.
That length-insensitivity is exactly why it suits text retrieval: you want to match meaning, not word count.
In practice a vector database computes this for you. It matters when interpreting scores — typical "relevant" thresholds sit around 0.7-0.85 depending on the embedding model, and thresholds are not transferable between models.
Related terms
Embedding
An embedding is a list of numbers representing the meaning of a piece of text, such that semantically similar texts have mathematically similar vectors. Embeddings make meaning-based search possible.
Semantic search
Semantic search finds results by meaning rather than exact keywords, using embeddings to compare concepts. It matches "how do I get my money back" to a document titled "Refund Policy".
Vector database
A vector database stores embeddings and finds the most similar ones to a query vector quickly. It is the retrieval layer of most RAG systems. Examples include Pinecone, Weaviate, Qdrant and pgvector.
RAG (Retrieval-Augmented Generation)
RAG retrieves relevant passages from your own documents and inserts them into the prompt before the model answers. It grounds responses in your data, cuts hallucination, and needs no retraining.
Chunking
Chunking splits documents into smaller passages before embedding them for retrieval. Chunk size is a key quality lever: too large dilutes relevance, too small loses the context needed to make sense.
Reranking
Reranking takes an initial set of retrieved candidates and reorders them with a more accurate but slower model. It is one of the cheapest ways to materially improve RAG quality.
Put this into practice
Understanding the term is step one. Our free courses and tools let you actually use it.