AI Infrastructure

Chunking

Definition
Chunking is the technique of splitting longer documents into individual segments (chunks) so a language model can process them. Each chunk becomes its own embedding in a vector database. That is the foundation of every RAG pipeline: the model receives only the text passages relevant to a question, instead of the entire corpus inside the context window. Common strategies are fixed-length splitting by token count, recursive splitting along structural boundaries such as headings, paragraphs, and list items, and semantic chunking, which uses embedding similarity to find natural break points. Overlapping chunks are the proven default: 10–20 % shared content between neighbours prevents information loss at the cut lines. Chunk size is a trade-off between precision and coherence. Small chunks are found more precisely and are cheaper to embed, but they tear terms out of context. Large chunks retain more meaning but blur similarity search and consume tokens in the context window. Typical values for LLM pipelines range from 256 to 1024 tokens. Newer approaches such as late chunking embed the full document first and derive the chunk vectors from the intermediate layers, improving the context fit of the results. Chunking directly determines the answer quality of any RAG application: clean segmentation reduces hallucinations, speeds up retrieval, and lowers token costs. For teams with their own document base, it is the first and most effective tuning parameter — ahead of the model choice itself.
Category
AI Infrastructure

Deep Dive: Chunking

Chunking is the technique of splitting longer documents into individual segments (chunks) so a language model can process them. Each chunk becomes its own embedding in a vector database. That is the foundation of every RAG pipeline: the model receives only the text passages relevant to a question, instead of the entire corpus inside the context window. Common strategies are fixed-length splitting by token count, recursive splitting along structural boundaries such as headings, paragraphs, and list items, and semantic chunking, which uses embedding similarity to find natural break points. Overlapping chunks are the proven default: 10–20 % shared content between neighbours prevents information loss at the cut lines. Chunk size is a trade-off between precision and coherence. Small chunks are found more precisely and are cheaper to embed, but they tear terms out of context. Large chunks retain more meaning but blur similarity search and consume tokens in the context window. Typical values for LLM pipelines range from 256 to 1024 tokens. Newer approaches such as late chunking embed the full document first and derive the chunk vectors from the intermediate layers, improving the context fit of the results. Chunking directly determines the answer quality of any RAG application: clean segmentation reduces hallucinations, speeds up retrieval, and lowers token costs. For teams with their own document base, it is the first and most effective tuning parameter — ahead of the model choice itself.

Production-Ready Guardrails