Document Chunking: Optimizing RAG Retrieval Pipelines

Michael BrenndoerferJanuary 23, 202663 min read

Part of Language AI Handbook

Covers document chunking for RAG systems. Examines fixed-size, recursive, and semantic strategies to balance retrieval precision with context window limits.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Document Chunking

In the previous chapters, we built up the core RAG pipeline: retrieving relevant documents and using them to ground an LLM's responses. But we glossed over a critical question: what exactly constitutes a "document" in retrieval? A single PDF might contain hundreds of pages. A webpage might span thousands of words. An internal knowledge base might hold gigabytes of technical documentation, legal contracts, and customer support transcripts. If you embed an entire document as one vector, the resulting embedding must compress an enormous amount of information into a single point in vector space. Important details get diluted, and retrieval precision suffers drastically.

Document chunking solves this problem by splitting large documents into smaller, self-contained pieces before embedding and indexing them. Think of it as deciding how to cut a long book into individual pages before filing them in a library catalog. The finer and more sensibly you cut, the easier it is for someone to find exactly the passage they need. The choice of chunking strategy, chunk size, and overlap between chunks has a surprisingly large impact on RAG quality. A poorly chunked document can bury relevant information inside irrelevant context, while a well-chunked document makes the needed passage easy to retrieve.

The importance of this decision is easy to underestimate. Many practitioners treat chunking as a preprocessing detail and move on quickly to more glamorous components like embedding models or vector databases. But chunking is foundational. It defines the fundamental unit of retrieval: what size and shape of knowledge can the system find and return? Get this wrong and no amount of fine-tuning the downstream components will fully compensate.

Consider a realistic scenario. You are building a RAG system over a collection of scientific papers. A user asks: "What was the sample size in the Jensen et al. study on working memory?" If you have chunked each paper into a single vector, the answer to that highly specific question is buried inside an embedding that also represents the abstract, the literature review, the discussion section, and the conclusions. The probability of retrieving the exact paper is reasonable, but the probability of retrieving the exact sentence containing the sample size is low. If instead you chunk at the paragraph level, the specific methodology paragraph containing the sample size becomes a distinct, findable unit in your index.

This chapter examines chunking strategies, from simple fixed-size splits to semantically aware approaches that respect the natural structure of text. We will implement each strategy, examine how chunk size affects retrieval, build intuition for when to use each approach, and work through a concrete numerical example that illustrates the mechanics of each method. Along the way, we will cover chunk overlap, metadata enrichment, and the special challenges posed by different document types.

Historical Context

The concept of breaking documents into smaller units for retrieval predates neural RAG systems by decades. Early information retrieval systems in the 1970s and 1980s faced similar challenges when indexing encyclopedias and legal databases for keyword search. Librarians and information scientists developed manual chunking conventions based on abstracts or paragraphs, with longer passages used where appropriate. These became standard retrieval units in systems like STAIRS (Storage and Information Retrieval System) and WESTLAW. When statistical IR methods like BM25 and TF-IDF emerged in the late 1980s, these passage-level units carried over naturally. The key insight from that era, that a focused topical passage retrieves better than a sprawling document, remains just as true for dense vector retrieval today. What changed is that modern semantic embedding models can exploit topical coherence in ways that keyword systems could not, making the quality of chunk boundaries even more consequential.

Why Chunking Matters

To understand why chunking is so important, consider what happens during the retrieval stage of RAG. As we discussed in the Dense Retrieval chapter, we encode both queries and documents into dense vectors and retrieve documents whose embeddings are closest to the query embedding. The quality of this retrieval depends directly on how well each embedding captures the meaning of its text. If the embedding is a faithful representation of a focused, coherent passage, it will match queries about that passage with high confidence. If the embedding is a blurred summary of many disparate topics, it becomes a mediocre match for all of them and a strong match for none.

Think of an embedding as a GPS coordinate in semantic space. A chunk about "ocean temperature trends in the Arctic" gets a precise coordinate in the region of oceanography and climate science. A chunk that covers ocean temperatures, carbon policy, economic impacts, and indigenous communities gets a coordinate that averages all of those regions, landing somewhere in the fuzzy middle ground between them. When a query about Arctic ocean temperatures approaches that coordinate, it is farther from the document-level average embedding than from a focused chunk embedding. Precision degrades not because the information is absent, but because it is diluted.

Embedding models have two fundamental constraints that make chunking necessary. Understanding both helps you reason about what will go wrong if you ignore them.

The first constraint is the context window limit. Most embedding models have a maximum input length, typically 512 tokens for models like those in the BERT family and up to 8,192 tokens for newer models. Text beyond this limit is simply truncated and lost. This means that if you feed a 10,000-token document into a 512-token embedding model, nearly 95% of the document is silently discarded. The resulting embedding reflects only the opening passage, leaving the vast majority of the document completely unrepresented in your index. An entire chapter on quantum computing, starting after the 512th token of a textbook, would simply not exist in your retrieval system.

The second constraint is information density. Even within the context window, longer texts produce embeddings that represent a blend of all topics covered. A 5,000-word article about climate change might discuss ocean temperatures, carbon emissions, policy proposals, and economic impacts. A single embedding for this article would be a vague average of all these topics, matching none of them precisely. Think of this like trying to describe an entire meal with a single adjective: "savory" might loosely apply to the soup, the steak, and the roasted vegetables, but it captures the distinctive character of none of them. The embedding loses the distinctiveness that makes retrieval useful.

Chunking addresses both constraints simultaneously. By splitting documents into smaller pieces, each chunk fits within the model's context window and covers a focused topic. When you ask about ocean temperature trends, the chunk specifically discussing that topic will produce a much stronger similarity match than the full article's embedding would.

Retrieval quality depends on having the right information represented in a way that makes it discoverable. Chunking creates this representation. Each chunk becomes a distinct, findable unit in vector space, carrying a focused semantic signal that can align with a matching query. The rest of this chapter explains how to create those units well.

The Chunking-Retrieval Connection

Chunking defines the fundamental unit of retrieval. When you chunk a document, you are deciding the granularity at which information can be found and returned. Too coarse, and relevant details are buried in noise. Too fine, and context is lost. This decision echoes throughout the entire pipeline: it shapes what the embedding model encodes, what the similarity search can find, and what the LLM ultimately sees in its context window.

Fixed-Size Chunking

The simplest chunking strategy splits text into pieces of a fixed number of characters or tokens, regardless of where sentences or paragraphs begin or end. Despite its simplicity, fixed-size chunking is common in production RAG systems because it runs quickly and produces predictable boundaries that are easy to reason about. When you need to process millions of documents quickly and you want deterministic, reproducible behavior, fixed-size chunking is often the pragmatic choice.

Think of fixed-size chunking as using a cookie cutter on dough. The cutter produces uniform shapes, and you can process them quickly on an assembly line. The downside is that sometimes you cut through the chocolate chip or split the raisin, producing shapes that lack the distinctive feature you were hoping to capture. The uniformity is an asset for engineering; it is a liability for meaning.

Its uniformity also simplifies downstream engineering: every chunk consumes roughly the same amount of storage, embedding compute, and retrieval bandwidth. You can predict exactly how many chunks a document will produce, how much index storage you need, and how long embedding will take. For large-scale systems where these properties matter, this predictability is valuable. Many production systems use fixed-size chunking not because it is optimal, but because it is reliable enough and far easier to operate at scale than semantically aware alternatives.

Before diving into implementations, it is worth developing intuition for what "fixed-size" means in practice. Two natural choices are character counts and token counts, and they behave quite differently.

Character-Based Splitting

The most basic approach splits text every nn characters. This produces chunks of uniform length but can cut words and sentences in the middle.

The mechanism is straightforward: we define a window of size nn and slide it along the text, recording each window as a chunk. When overlap is specified, adjacent chunks share some content. Formally, if we denote the full text as a sequence of characters T=t1,t2,…,tLT = t_1, t_2, \ldots, t_L, and we set chunk size nn and overlap oo, then chunk kk contains characters at positions:

Ck=T[k(n−o):k(n−o)+n]C_k = T[k(n-o) : k(n-o) + n]

where:

  • kk: the zero-indexed chunk number
  • nn: the chunk size in characters
  • oo: the overlap in characters (between 0 and nn)
  • T[a:b]T[a:b]: the substring of TT from position aa to bb

The stride between consecutive chunks is n−on - o characters. With zero overlap, o=0o = 0, and chunks are perfectly non-overlapping; each position in the text appears in exactly one chunk. With positive overlap, positions in the overlap zone appear in two consecutive chunks.

In[3]:
Code
!uv pip install tiktoken spacy sentence-transformers matplotlib numpy
!uv pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl

def chunk_by_characters(text, chunk_size=200, overlap=0):
    """Split text into fixed-size character chunks."""
    chunks = []
    start = 0
    while start < len(text):
        end = start + chunk_size
        chunks.append(text[start:end])
        start += chunk_size - overlap
    return chunks

sample_text = (
    "The Amazon rainforest produces about 20 percent of the world's oxygen. "
    "It spans across nine countries in South America. The forest is home to "
    "approximately 10 percent of all species on Earth. Deforestation has reduced "
    "its area significantly over the past decades. Scientists estimate that about "
    "17 percent of the Amazon has been destroyed in the last 50 years. Conservation "
    "efforts are critical to preserving this vital ecosystem. Many indigenous "
    "communities depend on the rainforest for their livelihoods. The Amazon River, "
    "which flows through the forest, is the largest river by volume in the world."
)
In[4]:
Code
chunks = chunk_by_characters(sample_text, chunk_size=150)
Out[5]:
Console
Chunk 0: [150 chars] 'The Amazon rainforest produces about 20 percent of the world's oxygen. It spans across nine countries in South America. The forest is home to approxim'
Chunk 1: [150 chars] 'ately 10 percent of all species on Earth. Deforestation has reduced its area significantly over the past decades. Scientists estimate that about 17 pe'
Chunk 2: [150 chars] 'rcent of the Amazon has been destroyed in the last 50 years. Conservation efforts are critical to preserving this vital ecosystem. Many indigenous com'
Chunk 3: [150 chars] 'munities depend on the rainforest for their livelihoods. The Amazon River, which flows through the forest, is the largest river by volume in the world'
Chunk 4: [1 chars] '.'

Notice how chunks cut through words and sentences without regard for meaning. Chunk 1 might start in the middle of a word, making it difficult for an embedding model to capture the intended meaning. A phrase like "approximately 10 percent of all species" could be severed into "approxi" at the end of one chunk and "mately 10 percent of all species" at the beginning of the next, rendering both fragments less semantically coherent than the original sentence.

This is the fundamental weakness of character-based splitting: it optimizes for uniform size at the expense of coherence. The embedding model must then try to extract meaning from text that may begin or end mid-thought, producing embeddings that are noisier and less representative of any single idea. For use cases where speed and simplicity matter more than embedding quality, character-based splitting is a reasonable starting point. For use cases where retrieval precision is critical, it almost always warrants improvement.

Token-Based Splitting

A better variant splits by token count rather than character count. Since embedding models operate on tokens, not raw characters, this ensures each chunk uses the model's capacity efficiently. A character-based chunk of 200 characters might translate to anywhere from 30 to 60 tokens depending on word length and vocabulary, creating unpredictable utilization of the embedding model's context window. Token-based splitting eliminates this variability by directly controlling the unit that matters.

The key insight here is alignment: you want to control the input size in the same units the model uses internally. Controlling characters while the model measures tokens is like setting a cooking timer in Fahrenheit when the recipe is in Celsius. You can make it work, but you are adding unnecessary conversion error. Token-based splitting removes that friction.

We can use the tiktoken library, which implements the tokenizers used by OpenAI models, to perform this conversion accurately. The cl100k_base encoding is the tokenizer used by GPT-4 and GPT-3.5-turbo, and it is widely used as a proxy tokenizer even for non-OpenAI models because it has well-understood behavior on English text.

In[6]:
Code
import tiktoken


def chunk_by_tokens(
    text, chunk_size=50, overlap=0, encoding_name="cl100k_base"
):
    """Split text into chunks of a fixed number of tokens."""
    enc = tiktoken.get_encoding(encoding_name)
    tokens = enc.encode(text)

    chunks = []
    start = 0
    while start < len(tokens):
        end = min(start + chunk_size, len(tokens))
        chunk_tokens = tokens[start:end]
        chunk_text = enc.decode(chunk_tokens)
        chunks.append(chunk_text)
        start += chunk_size - overlap
    return chunks
In[7]:
Code
token_chunks = chunk_by_tokens(sample_text, chunk_size=40, overlap=0)
Out[8]:
Console
Chunk 0: [40 tokens] 'The Amazon rainforest produces about 20 percent of the world's oxygen. It spans across nine countries in South America. The forest is home to approximately 10 percent of all species on Earth. Def'
Chunk 1: [40 tokens] 'orestation has reduced its area significantly over the past decades. Scientists estimate that about 17 percent of the Amazon has been destroyed in the last 50 years. Conservation efforts are critical to preserving this vital ecosystem'
Chunk 2: [34 tokens] '. Many indigenous communities depend on the rainforest for their livelihoods. The Amazon River, which flows through the forest, is the largest river by volume in the world.'
Out[9]:
Visualization
Histogram showing the distribution of character lengths for fixed 50-token chunks, with a red dashed vertical line marking the mean character length.
Distribution of character lengths for complete 50-token chunks, generated by repeating the sample text. Although every chunk contains exactly 50 tokens, character lengths vary because some words tokenize to more tokens than others. The red dashed line marks the mean character length.

Token-based splitting guarantees each chunk uses a predictable number of tokens, but it still does not respect sentence boundaries. A sentence split across two chunks loses coherence in both. The first chunk ends with an incomplete thought, and the second begins without the context established earlier in the sentence. For many retrieval scenarios, this partial-sentence problem degrades embedding quality enough to motivate the sentence-aware approaches we explore next.

Key Parameters

The key parameters for fixed-size chunking are:

  • chunk_size: The target size of each chunk (in characters or tokens). Smaller chunks are more precise, while larger chunks provide more context. This parameter directly controls the trade-off between retrieval granularity and the amount of information each chunk carries.
  • overlap: The number of units (characters or tokens) repeated between adjacent chunks to prevent information loss at boundaries. Overlap acts as a safety net. This keeps content near a cut point is fully represented in at least one chunk. We explore this concept in depth in the Chunk Overlap section below.

Sentence-Based Chunking

A more linguistically motivated approach uses sentence boundaries as the atomic unit. Rather than slicing text at arbitrary positions, this strategy first identifies where sentences end and then groups consecutive sentences together until a size limit is reached. The result is chunks that always begin and end at natural linguistic boundaries.

As we covered in the Sentence Segmentation chapter, identifying sentence boundaries is itself a non-trivial task that must handle abbreviations, decimal numbers, and other edge cases. A naive approach of splitting on periods fails for "Dr. Smith increased the dosage to 2.5 mg." Here, we use spaCy's sentence segmentation for reliable boundary detection. SpaCy's model combines statistical patterns with rule-based heuristics that handle most edge cases correctly.

The idea is simple but powerful: split the text into individual sentences first, then group consecutive sentences together into chunks until adding the next sentence would exceed the desired size limit. Think of sentences as Lego bricks. Each brick is a complete, self-contained unit. You can stack them together to form larger structures, but you never break a brick in the middle. This guarantees that no chunk will ever contain a partial sentence, and every chunk begins at the start of a sentence and ends at the conclusion of one.

This sentence-preservation property matters because embedding models are trained on complete, well-formed text. Most modern sentence embedding models fine-tune their weights on datasets of complete sentences or paragraphs. When you feed them truncated or fragmented text, the resulting embeddings are noisier and less semantically reliable. Preserving sentence boundaries aligns your chunks with the distribution of text the embedding model was trained on, producing higher-quality representations.

In[10]:
Code
import spacy

nlp = spacy.load("en_core_web_sm")


def chunk_by_sentences(text, max_chunk_size=200, overlap_sentences=0):
    """Group sentences into chunks that respect sentence boundaries."""
    doc = nlp(text)
    sentences = [sent.text.strip() for sent in doc.sents]

    chunks = []
    current_chunk = []
    current_length = 0

    for sent in sentences:
        sent_length = len(sent)
        # If adding this sentence would exceed the limit, finalize current chunk
        if current_chunk and current_length + sent_length + 1 > max_chunk_size:
            chunks.append(" ".join(current_chunk))
            # Keep the last `overlap_sentences` for context continuity
            if (
                overlap_sentences > 0
                and len(current_chunk) >= overlap_sentences
            ):
                current_chunk = current_chunk[-overlap_sentences:]
                current_length = (
                    sum(len(s) for s in current_chunk) + len(current_chunk) - 1
                )
            else:
                current_chunk = []
                current_length = 0
        current_chunk.append(sent)
        current_length += sent_length + (1 if current_length > 0 else 0)

    if current_chunk:
        chunks.append(" ".join(current_chunk))

    return chunks
In[11]:
Code
sent_chunks = chunk_by_sentences(
    sample_text, max_chunk_size=200, overlap_sentences=0
)
Out[12]:
Console
Chunk 0: [191 chars]
  'The Amazon rainforest produces about 20 percent of the world's oxygen. It spans across nine countries in South America. The forest is home to approximately 10 percent of all species on Earth.'

Chunk 1: [168 chars]
  'Deforestation has reduced its area significantly over the past decades. Scientists estimate that about 17 percent of the Amazon has been destroyed in the last 50 years.'

Chunk 2: [145 chars]
  'Conservation efforts are critical to preserving this vital ecosystem. Many indigenous communities depend on the rainforest for their livelihoods.'

Chunk 3: [94 chars]
  'The Amazon River, which flows through the forest, is the largest river by volume in the world.'

Each chunk now contains complete sentences. The text within each chunk is coherent, making it much easier for an embedding model to capture its meaning. When the model processes a sentence-aligned chunk, it encounters grammatically well-formed text with clear subjects, verbs, and objects, exactly the kind of input it was trained on. This alignment between the structure of training data and the structure of your chunks tends to produce higher-quality embeddings.

The chunk sizes are no longer perfectly uniform, since they depend on sentence lengths, but this is a worthwhile trade-off. A chunk that is 180 characters of coherent prose will almost always produce a better embedding than a chunk that is exactly 200 characters but begins with the fragment "tion efforts are critical to preserving." The embedding model extracts meaning from the semantics of the text, not from its length. Feeding it coherent input produces a coherent, useful embedding.

One practical consideration with sentence-based chunking is that it can produce chunks of highly variable size when sentence lengths vary dramatically. A document alternating between very short sentences ("DNA mutates.") and very long ones could produce chunks that range from 10 characters to 400 characters for the same max_chunk_size setting. This variability rarely causes problems for retrieval quality, but it is worth monitoring if you care about predictable indexing time or storage usage.

Key Parameters

The key parameters for sentence-based chunking are:

  • max_chunk_size: The maximum length (in characters or tokens) allowed for a chunk before a split is forced. Sentences are accumulated into the current chunk as long as this limit is not exceeded, so actual chunk sizes will vary depending on where the sentence boundaries fall relative to this ceiling.
  • overlap_sentences: The number of full sentences to repeat at the beginning of the next chunk to preserve context. Unlike character or token overlap, sentence overlap guarantees that the shared content is always a complete, meaningful unit. This is particularly valuable when consecutive sentences contain co-references ("This process..." or "These results...") that would be difficult to interpret without the preceding sentence.

Recursive Chunking

Real documents have hierarchical structure: sections, subsections, paragraphs, sentences, and words. Recursive chunking exploits this structure by attempting to split at the most meaningful boundary possible. The strategy embodies a simple but effective principle: always prefer splitting at the highest-level boundary that still produces chunks within the size limit.

Think of recursive chunking as the way you might physically organize a large collection of documents. You first try to separate them by major category (sections). If a category still has too many documents, you separate by subcategory (paragraphs). If a subcategory is still too large, you split by individual pages (sentences). Only as a last resort do you cut mid-page (characters). At each level, you prefer the coarser division, resorting to finer divisions only when necessary.

The algorithm tries a sequence of separators in order of decreasing granularity. If splitting by double newlines (paragraph boundaries) produces chunks that are small enough, it stops there. If not, it falls back to single newlines, then sentences, then words, then characters. This layered approach means the algorithm naturally adapts to different parts of a document. A section with short paragraphs will be split along paragraph boundaries, preserving each paragraph as a self-contained unit. A section containing a single long paragraph will be split at sentence boundaries within that paragraph.

This approach is popularized by LangChain's RecursiveCharacterTextSplitter and is one of the most widely used strategies in practice. Its popularity stems from the fact that it produces reasonable results across a wide variety of document types without requiring any domain-specific configuration beyond the choice of separators. For many teams building their first RAG system, recursive chunking is the right default choice: it respects document structure when that structure exists, and degrades gracefully to simpler splits when it does not.

In[13]:
Code
def recursive_chunk(text, chunk_size=200, overlap=50, separators=None):
    """Recursively split text using a hierarchy of separators."""
    if separators is None:
        separators = ["\n\n", "\n", ". ", " ", ""]

    # Base case: text fits in one chunk
    if len(text) <= chunk_size:
        return [text]

    # Find the best separator that produces a split
    chosen_sep = separators[-1]  # fallback to character-level
    for sep in separators:
        if sep in text:
            chosen_sep = sep
            break

    # Split text using the chosen separator
    parts = text.split(chosen_sep) if chosen_sep else list(text)

    # Merge parts into chunks that respect the size limit
    chunks = []
    current_chunk = ""

    for part in parts:
        # Reconstruct with separator
        candidate = current_chunk + chosen_sep + part if current_chunk else part

        if len(candidate) <= chunk_size:
            current_chunk = candidate
        else:
            if current_chunk:
                chunks.append(current_chunk)
            # If a single part exceeds chunk_size, recurse with finer separators
            if len(part) > chunk_size:
                remaining_seps = separators[separators.index(chosen_sep) + 1 :]
                if remaining_seps:
                    sub_chunks = recursive_chunk(
                        part, chunk_size, overlap, remaining_seps
                    )
                    chunks.extend(sub_chunks)
                    current_chunk = ""
                else:
                    current_chunk = part
            else:
                current_chunk = part

    if current_chunk:
        chunks.append(current_chunk)

    return chunks

Let's test it on a structured document with paragraphs:

In[14]:
Code
structured_text = """Climate Change Overview

Global temperatures have risen by approximately 1.1 degrees Celsius since the pre-industrial era. This warming is primarily driven by greenhouse gas emissions from human activities, including burning fossil fuels and deforestation.

Impact on Ecosystems

Rising temperatures affect biodiversity across the planet. Coral reefs are bleaching at unprecedented rates. Arctic sea ice is declining, threatening polar bear habitats. Migration patterns of birds and marine species are shifting northward.

Mitigation Strategies

Renewable energy sources like solar and wind power are expanding rapidly. Many countries have committed to net-zero emissions targets by 2050. Carbon capture technology is being developed but remains expensive. Individual actions like reducing meat consumption and flying less also contribute to emission reductions."""
In[15]:
Code
rec_chunks = recursive_chunk(structured_text, chunk_size=250, overlap=0)
Out[16]:
Console
Chunk 0: [23 chars]
  'Climate Change Overview'

Chunk 1: [231 chars]
  'Global temperatures have risen by approximately 1.1 degrees Celsius since the pre-industrial era. This warming is primarily driven by greenhouse gas emissions from human activities, including burning fossil fuels and deforestation.'

Chunk 2: [20 chars]
  'Impact on Ecosystems'

Chunk 3: [241 chars]
  'Rising temperatures affect biodiversity across the planet. Coral reefs are bleaching at unprecedented rates. Arctic sea ice is declining, threatening polar bear habitats. Migration patterns of birds and marine species are shifting northward.'

Chunk 4: [21 chars]
  'Mitigation Strategies'

Chunk 5: [209 chars]
  'Renewable energy sources like solar and wind power are expanding rapidly. Many countries have committed to net-zero emissions targets by 2050. Carbon capture technology is being developed but remains expensive'

Chunk 6: [105 chars]
  'Individual actions like reducing meat consumption and flying less also contribute to emission reductions.'

The recursive approach respects paragraph boundaries when possible, falling back to finer-grained splits only when a paragraph exceeds the size limit. This preserves the document's logical structure in the chunks. The "Climate Change Overview" heading and its accompanying paragraph naturally form one chunk, while the "Impact on Ecosystems" section forms another.

The algorithm arrives at these natural divisions not because it understands the content, but because the paragraph boundaries encoded by double newlines happen to align with topical boundaries. This is a pattern that holds reliably across well-structured documents, whether those documents are Wikipedia articles, academic papers, or technical documentation. The recursive chunker exploits this alignment silently and automatically, which is why it works so well in practice without domain-specific tuning.

Recursive chunking differs from simple separator-based splitting. A simple paragraph splitter would split every double newline, even when two short consecutive paragraphs together fit within the chunk size limit. Recursive chunking instead merges adjacent short paragraphs into a single chunk, only splitting when the merged size would exceed the limit. This behavior is generally preferable because it reduces the total number of chunks and ensures each chunk carries enough content to be useful.

Key Parameters

The key parameters for recursive chunking are:

  • chunk_size: The hard limit on chunk size. The algorithm attempts to keep chunks under this limit while using the largest possible separators. This means the algorithm always prefers a paragraph-level split over a sentence-level split, as long as both produce chunks that fit within this budget.
  • overlap: The number of characters to overlap between chunks. In recursive chunking, overlap is applied after the initial splitting pass, duplicating content from the end of one chunk into the beginning of the next.
  • separators: An ordered list of strings used to split the text (e.g., ["\n\n", "\n", " ", ""]). The algorithm tries them in sequence to find the best split point. The ordering encodes your preference for which types of boundaries should be preserved. By placing paragraph separators first, you ensure the algorithm only resorts to sentence or word breaks when paragraph-level splits are insufficient.

Chunk Overlap

When you split a document into non-overlapping chunks, information that spans a chunk boundary gets split across two chunks. A key fact might have its context in one chunk and its conclusion in the next. Neither chunk alone captures the complete idea, and retrieval may miss it entirely.

Consider a passage where one sentence introduces a concept and the very next sentence provides the critical detail you are searching for. If the boundary falls between these two sentences, the first chunk contains a setup without a payoff, and the second chunk contains an answer without its question. An embedding of either chunk alone may fail to match your query. Think of a jigsaw puzzle where the most important piece of the picture straddles two boxes: you cannot see what it depicts by looking at either box alone.

Chunk overlap solves this by including some text from the end of each chunk at the beginning of the next. If your chunk size is 200 tokens and your overlap is 50 tokens, then the last 50 tokens of chunk kk appear again as the first 50 tokens of chunk k+1k+1. This means any passage of 50 or fewer tokens near the boundary is guaranteed to appear in full within at least one chunk.

More formally, for a document with NN total tokens, a chunk size of nn, and an overlap of oo, chunk kk covers token positions:

Ck=tokens[k(n−o)  :  k(n−o)+n]C_k = \text{tokens}\left[k(n - o) \;:\; k(n - o) + n\right]

where:

  • kk: the zero-indexed chunk number
  • nn: the number of tokens per chunk
  • oo: the number of overlapping tokens between adjacent chunks
  • tokens[a:b]\text{tokens}[a:b]: the token subsequence from position aa to bb

The total number of chunks produced is approximately:

K≈⌈N−on−o⌉K \approx \left\lceil \frac{N - o}{n - o} \right\rceil

where:

  • NN: total number of tokens in the document
  • KK: total number of chunks
  • ⌈⋅⌉\lceil \cdot \rceil: the ceiling function (rounding up to the nearest integer)

Notice that increasing oo increases KK. With zero overlap, K=⌈N/n⌉K = \lceil N/n \rceil. With overlap equal to half the chunk size (o=n/2o = n/2), the number of chunks roughly doubles. This is the key cost of overlap: it increases storage and embedding compute proportionally.

Out[17]:
Visualization
Diagram showing three text chunks with overlapping regions highlighted in red.
Schematic of the chunk overlap mechanism. Three sequential chunks (blue, orange, and green) cover the document, with red shaded regions highlighting duplicated text at boundaries. Each overlap region ensures that information near a split point appears in full in at least one chunk's embedding.

The trade-off with overlap is straightforward:

  • More overlap means better boundary coverage, but it increases the total number of chunks and storage requirements. It also means the same text appears in multiple embeddings, which can inflate retrieval results with near-duplicate chunks. If your query matches the overlapping region, both adjacent chunks will score highly, potentially consuming two of your top-kk retrieval slots with content that is largely redundant.
  • Less overlap reduces redundancy, but risks losing context at boundaries. With zero overlap, any information that depends on content from both sides of a boundary is effectively invisible to retrieval.

A common rule of thumb is to set overlap to 10-20% of the chunk size. For a 500-token chunk, an overlap of 50-100 tokens usually provides sufficient boundary coverage without excessive duplication. This range represents a pragmatic balance: enough overlap to catch most boundary-spanning passages, but not so much that your index becomes bloated with near-identical chunks.

In practice, if you find that retrieval frequently returns two highly similar chunks from adjacent positions in the same document, you may be using too much overlap. If you find that relevant answers are being missed because critical context lands on the wrong side of a boundary, you may need more. The right value is empirical, not theoretical.

Let's see overlap in action with our token-based chunker:

In[18]:
Code
overlap_chunks = chunk_by_tokens(sample_text, chunk_size=40, overlap=10)
Out[19]:
Console
Chunk 0: 'The Amazon rainforest produces about 20 percent of the world's oxygen. It spans across nine countries in South America. The forest is home to approximately 10 percent of all species on Earth. Def'

Chunk 1: ' 10 percent of all species on Earth. Deforestation has reduced its area significantly over the past decades. Scientists estimate that about 17 percent of the Amazon has been destroyed in the last 50 years'

Chunk 2: ' Amazon has been destroyed in the last 50 years. Conservation efforts are critical to preserving this vital ecosystem. Many indigenous communities depend on the rainforest for their livelihoods. The Amazon River, which flows'

Chunk 3: ' their livelihoods. The Amazon River, which flows through the forest, is the largest river by volume in the world.'

Compare the end of each chunk with the beginning of the next. You should see shared text that bridges the boundary. This keeps any information near the cut point is captured in at least one chunk's embedding. This shared text is the overlap in action: it duplicates a small window of content so that the transition zone between chunks is always fully represented somewhere in the index.

Chunk Size Selection

Choosing the right chunk size is one of the most impactful decisions in a RAG pipeline. It involves a fundamental trade-off between precision and context, and the optimal balance depends on the nature of your data, the types of questions you ask, and the capabilities of your embedding model. Getting chunk size right can mean the difference between a system that reliably retrieves the right passage and one that returns vaguely related blocks of text.

Think of chunk size selection as choosing the magnification level on a microscope. A very high magnification gives you fine detail, but you lose the ability to see how structures relate to one another within the broader tissue. A low magnification shows you the big picture but reveals little about individual cells. The right magnification depends on what you are trying to see. Similarly, chunk size must be chosen based on the granularity of questions you expect users to ask and the granularity of information that documents contain.

Before examining the formal trade-off, it helps to understand the two failure modes at the extremes. When chunks are too small, retrieval becomes fragile: you can find the right words, but the surrounding context is missing, and the LLM cannot produce a useful answer. When chunks are too large, retrieval becomes imprecise: the embedding represents multiple topics, and the cosine similarity between query and chunk reflects a blend of relevance rather than a strong match. Understanding these failure modes helps you diagnose problems in your own system.

The Precision-Context Trade-off

The precision-context trade-off is the central tension in chunk size selection. At one extreme, you could make each chunk a single sentence, maximizing precision. At the other, you could embed entire documents, maximizing context. Neither extreme works well in practice, and understanding why illuminates the core challenge.

Small chunks (100-200 tokens) provide high retrieval precision. Each chunk covers a narrow topic, so when it matches a query, the match is likely relevant. The embedding vector represents a focused semantic concept, and the cosine similarity between query and chunk is a reliable indicator of topical alignment. However, small chunks may lack sufficient context for the LLM to generate a good answer. A chunk containing "The temperature was 42°C" is useless without knowing what system or location it refers to. The LLM receives a decontextualized fact and must either hallucinate the missing context or produce an unsatisfying response.

Large chunks (500-1,000 tokens) provide rich context. The LLM receives enough surrounding information to understand and synthesize an answer. A paragraph that introduces a concept, provides evidence, and draws a conclusion gives the LLM everything it needs to produce a coherent response. But large chunks produce less precise embeddings because they cover more topics, and they may include irrelevant information that confuses the model. When a 800-token chunk discusses three related but distinct subtopics, its embedding becomes a centroid in the semantic space that lies between all three topics. A query that is specifically about one of those subtopics may find a better cosine similarity match with a more focused chunk from a completely different document.

Out[20]:
Visualization
Line chart showing retrieval precision decreasing and context completeness increasing as chunk size grows, with an annotation marking the optimization zone where the two curves intersect.
Conceptual illustration of the precision-context trade-off as chunk size increases. Retrieval precision (red) decays as embeddings become less focused, while context completeness (green) grows as chunks contain more surrounding information. The optimization zone near the crossover point represents the chunk sizes that balance both concerns, typically between 200 and 600 tokens for most applications.

The optimal chunk size depends on several factors:

  • Query type: Factoid questions ("What is the boiling point of water?") benefit from small chunks, while analytical questions ("Explain the causes of the 2008 financial crisis") need larger chunks with more context. If your application primarily serves one type of query, you can tune chunk size accordingly. If it serves a mix, you may need to compromise or use multiple chunk sizes.
  • Document type: Technical documentation with short, self-contained sections works well with smaller chunks. Narrative text where ideas develop over paragraphs needs larger chunks. A software API reference, where each function's description is independent, naturally lends itself to small chunks. A legal brief, where arguments build across paragraphs, demands larger ones.
  • Embedding model capacity: Different models have different optimal input lengths. Some models are trained on short passages and perform best with 1-2 sentences, while others handle full paragraphs effectively. Feeding a paragraph-optimized model a single sentence wastes its capacity, while feeding a sentence-optimized model a full paragraph may degrade its embedding quality.

Empirical Comparison

Let's compare how different chunk sizes affect the content of chunks from the same document. By viewing the same text through different chunk size lenses, we can build intuition for how this parameter shapes the granularity and coherence of the resulting pieces.

In[21]:
Code
long_text = """
Machine learning is a subset of artificial intelligence that focuses on building systems
that learn from data. Unlike traditional programming where rules are explicitly coded,
machine learning algorithms identify patterns in data and make predictions based on
those patterns. The field has grown rapidly since the 2010s, driven by increases in
computational power and the availability of large datasets.

Supervised learning is the most common paradigm. In supervised learning, the algorithm
is trained on labeled examples, where each input is paired with its correct output.
Common supervised learning tasks include classification, where the goal is to assign
inputs to discrete categories, and regression, where the goal is to predict a continuous
value. Popular algorithms include linear regression, decision trees, and neural networks.

Unsupervised learning works with unlabeled data. The algorithm must discover structure
in the data without guidance. Clustering algorithms like K-means group similar data
points together. Dimensionality reduction techniques like PCA find compact representations
of high-dimensional data. Unsupervised learning is often used for exploratory data
analysis and feature extraction.

Reinforcement learning involves an agent that learns by interacting with an environment.
The agent takes actions and receives rewards or penalties based on the outcomes. Over time,
the agent learns a policy that maximizes cumulative reward. Reinforcement learning has
achieved notable successes in game playing, robotics, and most recently in aligning
large language models through RLHF, as discussed in earlier chapters of this book.
""".strip()
In[22]:
Code
size_results = {}
chunk_sizes = [100, 250, 500]

for size in chunk_sizes:
    size_results[size] = chunk_by_sentences(long_text, max_chunk_size=size)
Out[23]:
Console
=== Chunk size: 100 chars → 16 chunks ===
  Chunk 0: [110 chars] Machine learning is a subset of artificial intelligence that focuses on building...
  Chunk 1: [164 chars] Unlike traditional programming where rules are explicitly coded, machine learnin...
  Chunk 2: [127 chars] The field has grown rapidly since the 2010s, driven by increases in computationa...
  Chunk 3: [48 chars] Supervised learning is the most common paradigm....
  Chunk 4: [121 chars] In supervised learning, the algorithm is trained on labeled examples, where each...
  Chunk 5: [180 chars] Common supervised learning tasks include classification, where the goal is to as...
  Chunk 6: [82 chars] Popular algorithms include linear regression, decision trees, and neural network...
  Chunk 7: [48 chars] Unsupervised learning works with unlabeled data....
  Chunk 8: [67 chars] The algorithm must discover structure in the data without guidance....
  Chunk 9: [70 chars] Clustering algorithms like K-means group similar data points together....
  Chunk 10: [99 chars] Dimensionality reduction techniques like PCA find compact representations of hig...
  Chunk 11: [89 chars] Unsupervised learning is often used for exploratory data analysis and feature ex...
  Chunk 12: [88 chars] Reinforcement learning involves an agent that learns by interacting with an envi...
  Chunk 13: [80 chars] The agent takes actions and receives rewards or penalties based on the outcomes....
  Chunk 14: [70 chars] Over time, the agent learns a policy that maximizes cumulative reward....
  Chunk 15: [193 chars] Reinforcement learning has achieved notable successes in game playing, robotics,...

=== Chunk size: 250 chars → 10 chunks ===
  Chunk 0: [110 chars] Machine learning is a subset of artificial intelligence that focuses on building...
  Chunk 1: [164 chars] Unlike traditional programming where rules are explicitly coded, machine learnin...
  Chunk 2: [176 chars] The field has grown rapidly since the 2010s, driven by increases in computationa...
  Chunk 3: [121 chars] In supervised learning, the algorithm is trained on labeled examples, where each...
  Chunk 4: [180 chars] Common supervised learning tasks include classification, where the goal is to as...
  Chunk 5: [199 chars] Popular algorithms include linear regression, decision trees, and neural network...
  Chunk 6: [170 chars] Clustering algorithms like K-means group similar data points together. Dimension...
  Chunk 7: [178 chars] Unsupervised learning is often used for exploratory data analysis and feature ex...
  Chunk 8: [151 chars] The agent takes actions and receives rewards or penalties based on the outcomes....
  Chunk 9: [193 chars] Reinforcement learning has achieved notable successes in game playing, robotics,...

=== Chunk size: 500 chars → 4 chunks ===
  Chunk 0: [452 chars] Machine learning is a subset of artificial intelligence that focuses on building...
  Chunk 1: [434 chars] In supervised learning, the algorithm is trained on labeled examples, where each...
  Chunk 2: [498 chars] The algorithm must discover structure in the data without guidance. Clustering a...
  Chunk 3: [264 chars] Over time, the agent learns a policy that maximizes cumulative reward. Reinforce...

Smaller chunk sizes produce more chunks, each covering a narrower topic. With a 100-character limit, individual sentences or pairs of short sentences become the unit of retrieval, giving the system laser-like precision but very little surrounding context. Larger chunk sizes produce fewer chunks that span multiple topics: a 500-character chunk might combine the definition of supervised learning with examples of specific algorithms, creating a richer but less focused unit.

The key insight here is that neither precision nor context is a fixed quantity: both are determined by the relationship between chunk size and document structure. A 250-character chunk of a dense academic paper might be too small to convey a complete argument, while a 250-character chunk of a FAQ document might contain an entire question-answer pair. The same number means different things in different documents.

There is no universally correct size; the right choice depends on your use case and should ideally be determined through evaluation. Most practitioners start with 256-512 tokens as a reasonable default, then adjust based on observed retrieval quality. We will cover systematic evaluation in the RAG Evaluation chapter.

Structural Chunking

Many real-world documents have explicit structural markers: Markdown headers, HTML tags, LaTeX section commands, or table of contents entries. These markers are not arbitrary formatting; they represent deliberate authorial decisions about how information is organized. A section header signals a topical boundary, and the content beneath it forms a coherent unit that the author intended to be read together.

Structural chunking uses these markers to split documents along their natural boundaries. This keeps each chunk corresponds to a coherent section or subsection. This approach works particularly well for structured documents because it uses organizational signals that other chunking methods ignore. A fixed-size chunker treats a Markdown header as just another line of text, potentially grouping it with the tail end of the previous section. A structural chunker recognizes it as a boundary, keeping each section intact and associated with its heading.

Think of structural chunking as reading the document through the author's own table of contents. The author has already done the topical segmentation work; structural chunking simply follows their lead. This is especially valuable for product manuals and API references, as well as textbooks or Wikipedia articles, where authors invest significant effort in logical organization. Ignoring that organization and applying arbitrary cuts would throw away valuable editorial work.

There is an additional benefit beyond clean boundaries: the header path. By prepending the header hierarchy to each chunk, we give the embedding model information about what broader topic the chunk belongs to. "Classification" embedded in isolation could refer to library science, machine learning, biology, or many other fields. "Classification" embedded with the header path "[Machine Learning > Supervised Learning > Classification]" is unambiguously about the machine learning subfield. This disambiguation can significantly improve retrieval precision.

In[24]:
Code
import re


def chunk_by_markdown_headers(text, max_chunk_size=500):
    """Split a Markdown document by headers, preserving hierarchy."""
    # Split on lines that start with one or more # characters
    header_pattern = re.compile(r"^(#{1,6})\s+(.+)$", re.MULTILINE)

    sections = []
    last_end = 0
    headers_stack = []

    for match in header_pattern.finditer(text):
        # Save content before this header
        if last_end < match.start():
            content = text[last_end : match.start()].strip()
            if content and headers_stack:
                sections.append(
                    {"header": " > ".join(headers_stack), "content": content}
                )

        level = len(match.group(1))
        title = match.group(2)

        # Update header stack based on level
        while headers_stack and len(headers_stack) >= level:
            headers_stack.pop()
        headers_stack.append(title)

        last_end = match.end()

    # Capture remaining content
    remaining = text[last_end:].strip()
    if remaining and headers_stack:
        sections.append(
            {"header": " > ".join(headers_stack), "content": remaining}
        )

    # Combine header context with content into chunks
    chunks = []
    for section in sections:
        chunk_text = f"[{section['header']}]\n{section['content']}"
        if len(chunk_text) <= max_chunk_size:
            chunks.append(chunk_text)
        else:
            # Fall back to sentence-level splitting within the section
            sub_chunks = chunk_by_sentences(
                section["content"], max_chunk_size=max_chunk_size
            )
            for sc in sub_chunks:
                chunks.append(f"[{section['header']}]\n{sc}")

    return chunks
In[25]:
Code
markdown_doc = """# Machine Learning

Machine learning enables computers to learn from data without explicit programming.

## Supervised Learning

### Classification

Classification assigns inputs to discrete categories. Common algorithms include
logistic regression, support vector machines, and neural networks. The model learns
a decision boundary that separates different classes in the feature space.

### Regression

Regression predicts continuous values. Linear regression fits a straight line to
the data, while polynomial regression can capture nonlinear relationships. Neural
networks can learn arbitrarily complex regression functions.

## Unsupervised Learning

### Clustering

Clustering groups similar data points without labels. K-means is the most popular
clustering algorithm. It partitions data into K groups by minimizing within-cluster
variance. DBSCAN is an alternative that can find clusters of arbitrary shape.

### Dimensionality Reduction

PCA projects high-dimensional data onto its principal components. t-SNE and UMAP
create nonlinear 2D projections for visualization. Autoencoders learn compressed
representations through neural networks.
"""
In[26]:
Code
md_chunks = chunk_by_markdown_headers(markdown_doc, max_chunk_size=400)
Out[27]:
Console
Chunk 0:
[Machine Learning]
Machine learning enables computers to learn from data without explicit programming.
------------------------------
Chunk 1:
[Machine Learning > Supervised Learning > Classification]
Classification assigns inputs to discrete categories. Common algorithms include
logistic regression, support vector machines, and neural networks. The model learns
a decision boundary that separates different classes in the feature space.
------------------------------
Chunk 2:
[Machine Learning > Supervised Learning > Regression]
Regression predicts continuous values. Linear regression fits a straight line to
the data, while polynomial regression can capture nonlinear relationships. Neural
networks can learn arbitrarily complex regression functions.
------------------------------
Chunk 3:
[Machine Learning > Unsupervised Learning > Clustering]
Clustering groups similar data points without labels. K-means is the most popular
clustering algorithm. It partitions data into K groups by minimizing within-cluster
variance. DBSCAN is an alternative that can find clusters of arbitrary shape.
------------------------------
Chunk 4:
[Machine Learning > Unsupervised Learning > Dimensionality Reduction]
PCA projects high-dimensional data onto its principal components. t-SNE and UMAP
create nonlinear 2D projections for visualization. Autoencoders learn compressed
representations through neural networks.
------------------------------

Notice how each chunk includes its header context (e.g., [Supervised Learning > Classification]). This metadata helps both the embedding model and the LLM understand the chunk's position within the document's hierarchy. A chunk about "Classification" without the context that it falls under "Supervised Learning" within a "Machine Learning" document would be less informative. The header path acts as a breadcrumb trail that disambiguates the content and enriches the semantic signal available to the embedding model.

By prepending this contextual path, the resulting embedding encodes what the chunk says and where it sits within the larger knowledge structure. This is one of the few free improvements available in chunking: you are not changing the content, only adding context, and that added context improves retrieval precision with essentially no downside.

Key Parameters

The key parameters for structural chunking are:

  • max_chunk_size: The maximum size for the content of each structural section. If a section exceeds this, it is further split using a fallback strategy, typically sentence-based chunking within the section. This hybrid behavior is important because some documents contain sections that vary enormously in length. A brief introductory section might be a single sentence, while a detailed methodology section might span several pages. The fallback ensures that even oversized sections are chunked coherently, with each sub-chunk still carrying the parent section's header context.

Semantic Chunking

All the strategies we have discussed so far use surface-level cues: character counts, sentence boundaries, structural markers. None of them consider the meaning of the text. A paragraph break might occur in the middle of a sustained argument, or two paragraphs separated by a heading might discuss the same topic from different angles. Surface-level chunking strategies are blind to these semantic relationships.

Semantic chunking addresses this by using embeddings to detect where topics shift within a document, and placing chunk boundaries at those transition points. Rather than relying on formatting conventions that may or may not align with topical structure, semantic chunking directly measures the conceptual continuity of the text and uses that measurement to determine where to place boundaries.

Think of semantic chunking as listening to a lecture and noticing when the speaker changes subject. You do not rely on the speaker saying "next, let's talk about..." You simply feel the shift in topic from the content itself. Semantic chunking automates this intuition by measuring similarity between consecutive sentences: when similarity drops sharply, the topic is changing and a boundary should be placed.

This approach is more computationally expensive than the methods we have seen so far, because it requires embedding every sentence in the document during the chunking step itself. But for documents where topical boundaries do not align with formatting boundaries, it can produce substantially better chunks. The cost is paid at indexing time and amortized across every retrieval query that benefits from cleaner boundaries.

The core idea is elegant: if two consecutive sentences have similar embeddings, they likely discuss the same topic and should be in the same chunk. When embeddings diverge significantly between adjacent sentences, that point likely marks a topic shift and is a natural place to split.

Algorithm

The semantic chunking algorithm proceeds in four steps, each building naturally on the previous one.

  1. Segment text into sentences using standard sentence segmentation. Sentences serve as the finest-grained unit we consider, since splitting within a sentence would break grammatical coherence.
  2. Embed each sentence using a sentence embedding model. Each sentence is mapped to a dense vector that captures its semantic content. This gives us a sequence of vectors, one per sentence, representing the "meaning trajectory" of the document.
  3. Compute similarity between consecutive sentences using cosine similarity. For sentences ii and i+1i+1 with embedding vectors ei\mathbf{e}_i and ei+1\mathbf{e}_{i+1}, the similarity is:
si=ei⋅ei+1∥ei∥⋅∥ei+1∥s_i = \frac{\mathbf{e}_i \cdot \mathbf{e}_{i+1}}{\|\mathbf{e}_i\| \cdot \|\mathbf{e}_{i+1}\|}

where:

  • sis_i: the cosine similarity between sentences ii and i+1i+1, a scalar in [−1,1][-1, 1]
  • ei⋅ei+1\mathbf{e}_i \cdot \mathbf{e}_{i+1}: the dot product of the two embedding vectors, measuring how aligned they are in semantic space
  • ∥ei∥\|\mathbf{e}_i\| and ∥ei+1∥\|\mathbf{e}_{i+1}\|: the L2 norms of each embedding, used to normalize for vector magnitude

A value near 1.0 means the two sentences discuss nearly identical content. A value near 0.0 means they discuss unrelated topics. Values below 0.3 typically signal a significant topic shift.

  1. Identify breakpoints where similarity drops below a threshold. Specifically, we compute a threshold τ\tau at the pp-th percentile of all similarity values:
τ=percentile(s1,s2,…,sN−1,  p)\tau = \text{percentile}(s_1, s_2, \ldots, s_{N-1},\; p)

where:

  • τ\tau: the similarity threshold below which a boundary is placed
  • NN: the total number of sentences in the document
  • pp: the percentile value, typically between 10 and 40

A boundary is placed between sentences ii and i+1i+1 whenever si<τs_i < \tau. The text between consecutive boundaries forms a single chunk. The percentile-based threshold is adaptive: documents with overall high inter-sentence coherence will have a higher threshold than documents with naturally varied topics. This keeps the algorithm calibrates to the document's own baseline rather than applying a fixed global standard.

In[28]:
Code
import numpy as np
from sentence_transformers import SentenceTransformer

_embedding_models = {}


def semantic_chunk(
    text,
    model_name="all-MiniLM-L6-v2",
    threshold_percentile=25,
    min_chunk_size=2,
):
    """Split text into chunks based on semantic similarity between sentences."""
    # Step 1: Segment into sentences
    doc = nlp(text)
    sentences = [sent.text.strip() for sent in doc.sents if sent.text.strip()]

    if len(sentences) <= min_chunk_size:
        return [text], sentences, []

    # Step 2: Embed each sentence
    # Reuse the model when comparing several documents in one notebook run.
    if model_name not in _embedding_models:
        _embedding_models[model_name] = SentenceTransformer(model_name)
    model = _embedding_models[model_name]
    embeddings = model.encode(sentences)

    # Step 3: Compute cosine similarity between consecutive sentences
    similarities = []
    for i in range(len(embeddings) - 1):
        sim = np.dot(embeddings[i], embeddings[i + 1]) / (
            np.linalg.norm(embeddings[i]) * np.linalg.norm(embeddings[i + 1])
        )
        similarities.append(sim)

    # Step 4: Find breakpoints where similarity drops below threshold
    threshold = np.percentile(similarities, threshold_percentile)
    breakpoints = [
        i + 1 for i, sim in enumerate(similarities) if sim < threshold
    ]

    # Build chunks from breakpoints
    chunks = []
    start = 0
    for bp in breakpoints:
        chunk_sentences = sentences[start:bp]
        if len(chunk_sentences) >= min_chunk_size:
            chunks.append(" ".join(chunk_sentences))
            start = bp

    # Add remaining sentences
    if start < len(sentences):
        remaining = " ".join(sentences[start:])
        if chunks and len(sentences[start:]) < min_chunk_size:
            chunks[-1] += " " + remaining
        else:
            chunks.append(remaining)

    return chunks, sentences, similarities

Let's apply semantic chunking to a document that covers multiple distinct topics:

In[29]:
Code
multi_topic_text = """
The solar system consists of eight planets orbiting the Sun. Mercury is the closest
planet to the Sun and has no atmosphere. Venus is the hottest planet due to its thick
carbon dioxide atmosphere. Earth is the only known planet to support life.

Photosynthesis is the process by which plants convert sunlight into energy. Chlorophyll
in plant cells absorbs light energy. The process produces oxygen as a byproduct.
Plants use carbon dioxide and water as inputs for photosynthesis.

The stock market experienced significant volatility in 2023. Interest rate hikes by
central banks affected investor sentiment. Technology stocks showed strong recovery
in the second half of the year. Cryptocurrency markets also saw renewed interest from
institutional investors.

Returning to astronomy, the James Webb Space Telescope has revealed new details about
distant galaxies. Its infrared capabilities allow it to see through cosmic dust.
Scientists have discovered exoplanets in habitable zones using data from the telescope.
""".strip()
In[30]:
Code
result = semantic_chunk(multi_topic_text, threshold_percentile=30)
sem_chunks, sentences, similarities = result
Out[31]:
Visualization
Line chart of sentence-to-sentence cosine similarity with dips at topic transitions.
Cosine similarity scores between consecutive sentences in a multi-topic document covering astronomy, photosynthesis, finance, and space telescopes. Deep troughs in the similarity curve (below the dashed red threshold) mark topic transitions, which the semantic chunker uses as split points. Notice the similarity recovery when the text returns to astronomy in the final paragraph.
Out[32]:
Console
Number of semantic chunks: 5

Chunk 0: 'The solar system consists of eight planets orbiting the Sun. Mercury is the closest planet to the Sun and has no atmosph...
'
Chunk 1: 'Photosynthesis is the process by which plants convert sunlight into energy. Chlorophyll in plant cells absorbs light ene...
'
Chunk 2: 'The stock market experienced significant volatility in 2023. Interest rate hikes by central banks affected investor sent...
'
Chunk 3: 'Technology stocks showed strong recovery in the second half of the year. Cryptocurrency markets also saw renewed interes...
'
Chunk 4: 'Returning to astronomy, the James Webb Space Telescope has revealed new details about distant galaxies. Its infrared cap...
'

The semantic chunker detects topic transitions, placing boundaries between the astronomy, biology, and finance sections. Unlike fixed-size chunking, it produces chunks of varying length, each covering a coherent topic. Notice in particular that the document deliberately returns to astronomy in its final paragraph after the finance section. A structural chunker relying on paragraph breaks alone would separate all four paragraphs into distinct chunks without recognizing the thematic relationship between the first and last paragraphs. The semantic chunker, by contrast, detects the low similarity at the topic transitions, regardless of how the document is formatted.

Choosing the Threshold

The threshold parameter controls how aggressively the chunker splits. It determines the sensitivity of the algorithm to topic transitions: a lenient threshold ignores all but the most dramatic shifts, while a strict threshold reacts to even subtle changes in subject matter. A lower percentile (e.g., 10th) means only the most dramatic topic shifts trigger splits, producing larger, fewer chunks. A higher percentile (e.g., 50th) creates more, smaller chunks at every moderate topic change.

There are three natural ways to set this threshold, each with different adaptivity properties:

  • Percentile-based thresholds (as above): Split at the pp-th percentile of similarity scores. This adapts to the document's overall coherence. A document where all consecutive sentences are highly related will have a higher threshold than one with naturally varied topics. This adaptivity is the key advantage: the algorithm calibrates itself to the document's baseline level of topical coherence, only splitting where coherence drops significantly relative to the document's own norm.
  • Absolute thresholds: Split whenever similarity drops below a fixed value (e.g., 0.3). This is simpler but does not adapt to documents with generally high or low inter-sentence similarity. A technical paper with consistently high inter-sentence similarity might never trigger splits with an absolute threshold of 0.3, while a casual blog post with naturally varied topics might be split into tiny fragments.
  • Standard deviation-based: Split when similarity drops more than one standard deviation below the mean. This captures statistically significant topic shifts. The logic mirrors outlier detection: if most sentence pairs have similarity around 0.7 with a standard deviation of 0.1, a pair with similarity 0.5 represents a meaningful departure from the norm and likely signals a topic transition.

Key Parameters

The key parameters for the semantic chunking implementation are:

  • threshold_percentile: Controls the sensitivity of split detection. Lower values (e.g., 10) trigger fewer splits, producing large chunks that may span loosely related subtopics. Higher values (e.g., 50) trigger more splits, producing smaller chunks that isolate finer-grained topics at the risk of fragmenting coherent discussions.
  • min_chunk_size: Minimum number of sentences per chunk. Prevents creating tiny fragments from transient topic shifts. Without this constraint, a single transitional sentence that happens to differ from both its neighbors could be isolated into its own chunk, producing a fragment too small to be useful for retrieval or generation.
  • model_name: The SentenceTransformer model used for embedding. Models with better semantic understanding produce more accurate boundaries. The choice of model matters because the similarity scores that drive splitting are only as good as the embeddings that produce them. A model with weak topical discrimination will produce noisy similarity curves that lead to poorly placed boundaries.

Worked Example: Chunking a Real Document

To make the mechanics of each strategy concrete, let's walk through a single short passage and trace exactly what each method produces. Consider this five-sentence passage about climate science:

"The global mean surface temperature has increased by approximately 1.1 degrees Celsius since the pre-industrial period. This warming is primarily driven by increased concentrations of greenhouse gases. Carbon dioxide levels have risen from 280 parts per million to over 420 parts per million since industrialization. Methane is the second most important greenhouse gas, with concentrations more than doubling since 1750. These changes are causing widespread disruptions to weather patterns, sea levels, and ecosystems worldwide."

We will call the five sentences S1 through S5 for reference. The total character count is approximately 570 characters. Now let's trace through each method.

Character-based chunking (chunk_size=200, overlap=0): The algorithm ignores sentence boundaries entirely. Chunk 0 covers characters 0-199, which falls mid-sentence in S2. Chunk 1 covers 200-399, starting mid-sentence and ending mid-sentence. Chunk 2 covers 400-569 (the remainder). The word "Methane" is cut into "Meth" at the end of one chunk and "ane is..." at the start of the next, and neither half carries the full meaning of the sentence. Result: 3 chunks, each covering incomplete thoughts.

Token-based chunking (chunk_size=40, overlap=0): The total passage contains roughly 90 tokens. The algorithm produces approximately 3 chunks of 40, 40, and 10 tokens respectively. Each chunk is well within the embedding model's context window, and token boundaries are respected, but sentence boundaries are not. S2 will likely be split between chunk 0 and chunk 1. Result: 3 chunks with predictable token counts but arbitrary sentence splits.

Sentence-based chunking (max_chunk_size=200): The algorithm accumulates sentences until the next sentence would push the chunk over 200 characters. S1 alone is 97 characters; adding S2 (73 characters) gives 170, within the limit. Adding S3 (80 characters) would give 250, over the limit. So chunk 0 = S1 + S2. We restart: chunk 1 starts with S3, and adding S4 (73 characters) gives 153, within the limit. Adding S5 (90 characters) would give 243, over the limit. So chunk 1 = S3 + S4. Chunk 2 = S5 alone. Result: 3 chunks, each containing complete sentences and focused on a coherent sub-topic (warming trend, greenhouse gas concentrations, consequences).

Recursive chunking (chunk_size=300): The algorithm first tries to split on double newlines. The passage has none, so it tries single newlines. The passage has none either, so it tries ". " (sentence-ending periods). The first period after position 0 falls at the end of S1, at character 96. It checks whether accumulating S1 and S2 fits within 300 characters: 170 characters, yes. It checks S1 + S2 + S3: 250 characters, still within 300. It checks S1 + S2 + S3 + S4: 323, over the limit. So chunk 0 = S1 + S2 + S3. Chunk 1 = S4 + S5. Result: 2 chunks, each larger and richer in context than the sentence-based approach.

Semantic chunking (threshold_percentile=25): The algorithm embeds each of the 5 sentences and computes 4 pairwise similarities: sim(S1,S2), sim(S2,S3), sim(S3,S4), sim(S4,S5). Sentences about temperature change (S1), greenhouse gases (S2), CO2 specifically (S3), and methane specifically (S4) are all highly related, so their similarity scores are high. S5 discusses consequences (weather, sea levels, ecosystems), a somewhat different topic, so sim(S4,S5) is moderately lower. The 25th percentile threshold might fall just below sim(S4,S5), placing a boundary before S5. Result: chunk 0 = S1+S2+S3+S4 (all about greenhouse gas causes), chunk 1 = S5 (consequences).

This comparison illustrates that different strategies serve different purposes. Character-based chunking is fast but crude. Sentence-based chunking is reliable and respects linguistic structure. Recursive chunking produces larger, contextually richer units. Semantic chunking aligns boundaries with meaning, at the cost of requiring embeddings during chunking.

Comparing Chunking Strategies

Let's compare all our strategies on the same document to see how they differ in practice. Applying each method to identical input text makes the differences concrete and allows us to observe how each strategy's assumptions about text structure manifest in the resulting chunks.

In[33]:
Code
comparison_text = long_text  # Use the longer machine learning document for clearer differentiation

# Unpack tuple from semantic_chunk
sem_chunks, _, _ = semantic_chunk(comparison_text, threshold_percentile=30)

results = {
    "Fixed (200 chars)": chunk_by_characters(comparison_text, chunk_size=200),
    "Token (50 tokens)": chunk_by_tokens(comparison_text, chunk_size=50),
    "Sentence (200 chars)": chunk_by_sentences(
        comparison_text, max_chunk_size=200
    ),
    "Recursive (250 chars)": recursive_chunk(comparison_text, chunk_size=250),
    "Semantic": sem_chunks,
}
Out[34]:
Visualization
Bar chart comparing the number of chunks produced by five strategies: fixed-character, token, sentence-aware, recursive, and semantic chunking.
Number of chunks generated by each strategy applied to the same document. Sentence-aware and recursive splitting produce the most chunks in this example, followed by fixed-character, token, and semantic chunking.
Bar chart with error bars comparing average character length and one standard deviation across five chunking strategies.
Average chunk size and size variability (error bars show one standard deviation) for each strategy. Token-based chunks have the tightest character-length distribution here, while semantic chunks are larger and more variable because they follow topic boundaries.
Out[35]:
Console
Fixed (200 chars):
  Chunks: 9, Avg size: 184, Std: 46 chars
  First chunk preview: 'Machine learning is a subset of artificial intelligence that...
'
Token (50 tokens):
  Chunks: 6, Avg size: 276, Std: 21 chars
  First chunk preview: 'Machine learning is a subset of artificial intelligence that...
'
Sentence (200 chars):
  Chunks: 10, Avg size: 164, Std: 28 chars
  First chunk preview: 'Machine learning is a subset of artificial intelligence that...
'
Recursive (250 chars):
  Chunks: 10, Avg size: 164, Std: 43 chars
  First chunk preview: 'Machine learning is a subset of artificial intelligence that...
'
Semantic:
  Chunks: 5, Avg size: 329, Std: 154 chars
  First chunk preview: 'Machine learning is a subset of artificial intelligence that...
'

The results confirm our expectations. Fixed-size strategies produce uniform chunks but cut through text arbitrarily, while sentence-based and recursive strategies produce chunks of varying size that preserve linguistic coherence. The standard deviation of chunk sizes tells an important story: fixed-size methods have near-zero variance by construction, while linguistically aware methods trade that uniformity for meaningfulness.

The key insight from this comparison is that the strategies form a hierarchy of increasing sophistication and cost. Character-based splitting requires no external libraries and runs in microseconds. Token-based splitting adds a tokenization step. Sentence-based splitting adds a sentence segmentation step. Semantic chunking adds full sentence embedding during indexing. Each step up the hierarchy adds computational cost but also adds meaning-awareness to the resulting chunks. For most production systems, sentence-based or recursive chunking represents the best balance of quality and cost.

Metadata Enrichment

Raw chunks lose their document context. A chunk about "Section 3.2: Results" is much more useful to an LLM if it knows this section comes from "Annual Report 2023" by "Acme Corp." Enriching chunks with metadata improves both retrieval accuracy and the quality of generated answers. Metadata records the chunk's provenance and position, placing it back in its document context.

Think of metadata as the label on a file folder. Without the label, you cannot tell what is inside without opening it. With a label that says "2023 Q3 Revenue Analysis, Section 3, Page 12", you know exactly where it fits in the larger picture before you read a word. For retrieval systems, this contextual information serves two purposes: it helps the vector similarity search return more relevant results (because metadata can be used as filters), and it helps the LLM produce more grounded, citable responses.

Common metadata to attach to each chunk includes:

  • Source document: filename, URL, or document ID
  • Position: chunk index, page number, or section path
  • Structural context: parent headers, preceding and following chunk IDs
  • Document metadata: author, date, document type, language
In[36]:
Code
from dataclasses import dataclass, field

import tiktoken


@dataclass
class Chunk:
    text: str
    index: int
    source: str
    section: str = ""
    start_char: int = 0
    end_char: int = 0
    metadata: dict = field(default_factory=dict)

    @property
    def token_count(self):
        enc = tiktoken.get_encoding("cl100k_base")
        return len(enc.encode(self.text))


def create_enriched_chunks(text, source, chunk_size=200):
    """Create chunks with metadata."""
    raw_chunks = chunk_by_sentences(text, max_chunk_size=chunk_size)

    enriched = []
    char_offset = 0
    for i, chunk_text in enumerate(raw_chunks):
        start = text.find(chunk_text, char_offset)
        end = (
            start + len(chunk_text)
            if start >= 0
            else char_offset + len(chunk_text)
        )

        chunk = Chunk(
            text=chunk_text,
            index=i,
            source=source,
            start_char=max(start, 0),
            end_char=end,
            metadata={
                "total_chunks": len(raw_chunks),
                "has_previous": i > 0,
                "has_next": i < len(raw_chunks) - 1,
            },
        )
        enriched.append(chunk)
        char_offset = end

    return enriched
In[37]:
Code
enriched = create_enriched_chunks(
    sample_text, source="amazon_facts.txt", chunk_size=200
)
Out[38]:
Console
Chunk 0/3:
  Source: amazon_facts.txt
  Tokens: 39
  Chars [0:191]
  Text: 'The Amazon rainforest produces about 20 percent of the world's oxygen. It spans ...
'
Chunk 1/3:
  Source: amazon_facts.txt
  Tokens: 32
  Chars [192:360]
  Text: 'Deforestation has reduced its area significantly over the past decades. Scientis...
'
Chunk 2/3:
  Source: amazon_facts.txt
  Tokens: 24
  Chars [361:506]
  Text: 'Conservation efforts are critical to preserving this vital ecosystem. Many indig...
'

This metadata becomes invaluable during the retrieval and generation stages. The chunk index lets you retrieve neighboring chunks for additional context: if a retrieved chunk does not contain quite enough information to answer a question, the system can automatically pull in the preceding or following chunks to expand the context window. The source attribution enables citation in the generated response, allowing the LLM to tell you exactly where its information came from. The character offsets make it possible to highlight the relevant passage in the original document, connecting generated answers back to their source material.

Beyond these basic fields, production systems often enrich chunks with computed properties that support more sophisticated retrieval. A density score can flag chunks that are likely to be high-information (many named entities, technical terms, or citations). A freshness timestamp enables time-filtered retrieval, so queries about recent events can preferentially return newer chunks. A language tag allows multilingual systems to restrict retrieval to chunks in the user's language. The chunking step is the right moment to compute and attach all of these properties, because this is the last time you have easy access to the full document context before individual chunks enter the index.

We will see how this metadata integrates with vector databases in the upcoming chapters on Vector Similarity Search and the HNSW Index.

Special Document Types

Different document formats require specialized chunking approaches. The strategies we have discussed work well on well-formed prose, but many real-world documents contain structured content that breaks those assumptions entirely.

Code files present a particularly common challenge. A Python function that spans 80 lines cannot be split sensibly at the character level. The last 10 lines of a function contain variable assignments and return statements that only make sense in the context of the preceding 70 lines. The right unit for code chunking is the function or class definition, not an arbitrary size limit. If possible, parse the abstract syntax tree to identify natural boundaries. For Python, the ast module provides programmatic access to these boundaries. For languages without a Python AST library, regular expressions that match common function signatures often work well enough in practice.

Tables present a different challenge. A table is fundamentally relational: each row is meaningless without the column headers, and each cell value is meaningless without its row and column context. Splitting a table across chunks destroys this relational structure entirely. A chunk containing rows 15-20 of a table without the column headers from row 1 is essentially uninterpretable. The solution is to keep tables as single chunks, even if they exceed the normal chunk size limit, and to convert them to a text representation (like Markdown or CSV format) that preserves row-column relationships. If a table is too large for a single chunk, repeat the header row at the start of each sub-chunk.

The full list of special considerations by document type includes:

  • Code files: Chunk along function or class boundaries rather than by line count. Parse the abstract syntax tree if possible to identify natural boundaries.
  • Tables: Keep as single chunks even if they exceed the normal size limit. Convert to Markdown or CSV format that preserves row-column relationships.
  • Conversational data: Chat logs and dialogue transcripts should be chunked by conversational turns or topic shifts, not by arbitrary size. Each chunk should contain enough turns to understand the context of the conversation.
  • Legal and regulatory documents: These often have numbered sections and subsections with precise cross-references. Structural chunking that preserves section numbers and hierarchies is essential. A chunk referencing "as defined in Section 2.1(b)" is useless if you cannot trace that reference.

Limitations and Practical Considerations

Despite the variety of strategies available, document chunking remains more art than science. No single strategy works optimally across all document types, query patterns, and embedding models. The fundamental challenge is that chunking decisions must be made at indexing time, before you know what questions you will ask. A chunk boundary that perfectly separates two topics for one query may split the exact passage needed for another. This temporal mismatch between when you chunk and when you retrieve is a structural property of the problem, not a bug that better algorithms can fully eliminate.

This points to a deeper limitation: the optimal chunking strategy is query-dependent, but chunking happens before any query is known. In an ideal world, you would chunk the same document differently for different queries: coarser chunks for broad analytical questions, finer chunks for specific factoid lookups, topic-aligned chunks for thematic questions. Practical systems cannot afford this luxury, so they must pick a strategy and chunk size that works acceptably across the distribution of expected queries. This is inherently a compromise, and it means that any production RAG system will have some queries it handles well and some it does not. Understanding this limitation helps you set realistic expectations and motivates investing in RAG evaluation infrastructure.

Semantic chunking, while conceptually appealing, introduces its own challenges. It requires running an embedding model over every sentence during indexing, which significantly increases preprocessing time and cost. For a corpus of one million documents, each with an average of 100 sentences, semantic chunking requires 100 million sentence embeddings before a single query can be answered. This is feasible with GPU infrastructure but impractical for small teams doing quick experiments. The quality of the splits depends heavily on the embedding model's ability to capture topic coherence, and models trained primarily on sentence similarity may not always detect document-level topic transitions accurately. Semantic chunking can also produce highly variable chunk sizes: a long section on a single topic might produce a single enormous chunk, while a passage that quickly surveys several topics might be split into tiny fragments. Managing these extremes with minimum and maximum size constraints is possible, but those constraints partially undermine the semantic purity of the approach.

Another practical limitation is the interaction between chunk size and the downstream LLM's context window. If you retrieve five chunks of 500 tokens each, you consume 2,500 tokens of the LLM's context for retrieved content alone. With larger chunk sizes or more retrieved results, you may exhaust the context window before including your query and system instructions. This creates a system-level constraint: chunk size, number of retrieved results, prompt template length, and LLM context window all interact. A naive approach that optimizes chunk size in isolation, without considering how many chunks you plan to retrieve, can produce a system where the retrieval quality is excellent but the final LLM response is degraded by an overloaded context window.

Finally, evaluation of chunking quality is inherently tied to end-to-end RAG performance. You cannot evaluate chunking in isolation because the same chunks might work well with one embedding model and poorly with another, or might retrieve perfectly but confuse a particular LLM. A chunk that is semantically clean but starts with a pronoun referencing an entity from the previous chunk ("It was first discovered...") may retrieve correctly but confuse the LLM about what "it" refers to. These cross-chunk coherence issues are invisible to chunking-level metrics and only surface in end-to-end evaluation. We will address this challenge systematically in the RAG Evaluation chapter later in this part.

Summary

Document chunking turns large documents into retrieval-friendly pieces that each capture a focused topic. The key takeaways from this chapter are:

  • Fixed-size chunking (by characters or tokens) is simple and fast but ignores text structure, often cutting through sentences and ideas.
  • Sentence-based chunking uses linguistic boundaries to ensure each chunk contains complete sentences, producing more coherent embeddings at the cost of variable chunk sizes.
  • Recursive chunking respects document hierarchy by trying paragraph boundaries first, then falling back to finer-grained splits, combining structural awareness with size control.
  • Semantic chunking uses embeddings to detect topic shifts, placing boundaries where the text's meaning changes most dramatically.
  • Chunk overlap ensures information near boundaries appears in multiple chunks, reducing the risk of losing cross-boundary context.
  • Chunk size involves a precision-context trade-off: smaller chunks give more precise retrieval but less context, while larger chunks provide richer context but less focused embeddings. Typical values range from 200 to 1,000 tokens.
  • Metadata enrichment preserves document context (source, position, section hierarchy) that would otherwise be lost during chunking, enabling better retrieval and citation.
  • Special document types (code, tables, legal text) require customized chunking strategies that respect their internal structure rather than applying generic size-based splits.

In the next chapter, we will examine the embedding models that convert these chunks into the dense vectors used for retrieval, completing the connection between chunking decisions and retrieval quality.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about document chunking strategies and their impact on retrieval.

Document Chunking Quiz

Question 1 of 70 of 7 completed
Why does embedding an entire long document as a single vector often result in poor retrieval performance?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026documentchunking, author = {Michael Brenndoerfer}, title = {Document Chunking: Optimizing RAG Retrieval Pipelines}, year = {2026}, url = {https://mbrenndoerfer.com/writing/document-chunking-rag-strategies-retrieval}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Document Chunking: Optimizing RAG Retrieval Pipelines. Retrieved from https://mbrenndoerfer.com/writing/document-chunking-rag-strategies-retrieval
MLAAcademic
Michael Brenndoerfer. "Document Chunking: Optimizing RAG Retrieval Pipelines." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/document-chunking-rag-strategies-retrieval>.
CHICAGOAcademic
Michael Brenndoerfer. "Document Chunking: Optimizing RAG Retrieval Pipelines." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/document-chunking-rag-strategies-retrieval.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Document Chunking: Optimizing RAG Retrieval Pipelines'. Available at: https://mbrenndoerfer.com/writing/document-chunking-rag-strategies-retrieval (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Document Chunking: Optimizing RAG Retrieval Pipelines. https://mbrenndoerfer.com/writing/document-chunking-rag-strategies-retrieval

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.