Retrieval & data
Chunking
Definition
Chunking splits documents into smaller passages before embedding them for retrieval. Chunk size is a key quality lever: too large dilutes relevance, too small loses the context needed to make sense.
A reasonable starting point is 300-800 tokens per chunk with 10-20% overlap between adjacent chunks. Overlap prevents a relevant passage being split across a boundary and retrieved without its surrounding context.
Split on natural boundaries — headings, paragraphs, sections — rather than fixed character counts. A chunk that starts mid-sentence retrieves poorly because its embedding represents a fragment rather than an idea.
The highest-leverage refinement is prepending a short context header to each chunk: the document title and section it came from. Retrieval accuracy on ambiguous queries improves markedly, because the embedding now carries the passage's place in the whole.
Example
Prepend to each chunk: "From: 2026 Refund Policy > International Orders" before the chunk text.
Related terms
RAG (Retrieval-Augmented Generation)
RAG retrieves relevant passages from your own documents and inserts them into the prompt before the model answers. It grounds responses in your data, cuts hallucination, and needs no retraining.
Embedding
An embedding is a list of numbers representing the meaning of a piece of text, such that semantically similar texts have mathematically similar vectors. Embeddings make meaning-based search possible.
Context window
The context window is the maximum number of tokens a model can consider at once — your prompt, any attached documents, the conversation history, and the response it generates. Exceed it and the earliest content gets dropped.
Vector database
A vector database stores embeddings and finds the most similar ones to a query vector quickly. It is the retrieval layer of most RAG systems. Examples include Pinecone, Weaviate, Qdrant and pgvector.
Semantic search
Semantic search finds results by meaning rather than exact keywords, using embeddings to compare concepts. It matches "how do I get my money back" to a document titled "Refund Policy".
Reranking
Reranking takes an initial set of retrieved candidates and reorders them with a more accurate but slower model. It is one of the cheapest ways to materially improve RAG quality.
Put this into practice
Understanding the term is step one. Our free courses and tools let you actually use it.