The Decision Nobody Talks About
When I first built a RAG pipeline, I spent hours obsessing over which embedding model to use, which vector database to pick, and how to write the perfect retrieval prompt. I got all of those things reasonably right. And my system still gave terrible answers.
It took me an embarrassingly long time to figure out why. The problem wasn't the model. It wasn't the database. It was that I had split my documents into chunks that made no sense to a retrieval system. I was feeding it mangled paragraphs, sentences cut off mid-thought, and context stripped away from the very information that made it meaningful.
Here's the honest truth: retrieval quality is almost entirely capped by how you chunk your documents. You can have the best embedding model in the world, but if your chunks are bad, you're searching a broken index. Let's fix that.
What Is a Chunk, Really?
A chunk is just a piece of your document — a slice of text that gets turned into an embedding (a vector) and stored in your database. When a user asks a question, the system finds the chunks that are most semantically similar to the question and hands those to the LLM as context.
So your chunk is what the LLM actually reads. If the chunk has the answer in it, the LLM can answer. If it doesn't — because you split the document in the wrong place — the LLM has nothing to work with and will either guess or say it doesn't know.
That's why chunking isn't a minor configuration detail. It's the foundation everything else is built on.
The Three Chunking Decisions
Every chunking strategy boils down to three variables: chunk size, overlap, and where you draw the boundary. Get these wrong in different ways and you get different failure modes.
Chunk size is measured in tokens or characters. Small chunks (say, 100–200 tokens) are precise but lose context. Large chunks (1000+ tokens) preserve context but dilute the signal — the embedding has to represent too many ideas at once, so it doesn't represent any of them particularly well.
A reasonable starting point for most documents is somewhere between 256 and 512 tokens. Not too small that you lose context, not so large that your embedding becomes a vague mush.
The Goldilocks Problem
There's no universally correct chunk size. A legal contract needs different chunking than a FAQ page. Always test with your actual documents and actual queries before committing to a size.
Overlap is how many tokens you repeat between adjacent chunks. If chunk one ends at token 256 and you have 50 tokens of overlap, chunk two starts at token 206. Why bother? Because key information often lives right at the boundary between chunks. Without overlap, that information might never get retrieved cleanly.
Boundary strategy is where most beginners go wrong. This is deciding where to cut. Do you cut every 500 characters no matter what? Do you cut at paragraph breaks? Sentence breaks? Section headers? The answer dramatically changes retrieval quality.
Fixed-Size Chunking: The Quick and Dirty Approach
The simplest strategy is just chopping your document into equal-size pieces. Easy to implement, fast to run, and often surprisingly decent for unstructured text.
# Simple fixed-size chunking with overlap
def chunk_text(text, chunk_size=400, overlap=50):
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end])
start += chunk_size - overlap
return chunks
# The problem: this cuts mid-sentence constantly
# "The refund policy applies to all..." | "...purchases made after Jan 1"
# That ellipsis in the middle? Lost forever.The failure mode is obvious once you see it: the splitter doesn't care about sentences, paragraphs, or logical units. It just counts characters. So you end up with chunks that start or end mid-sentence, and the embedding for that chunk is trying to represent an incomplete thought.
Recursive Character Splitting: The Better Default
A big step up from fixed-size chunking. Instead of blindly chopping at a character count, recursive character splitting tries a hierarchy of separators. It first tries to split on double newlines (paragraph breaks). If a paragraph is still too big, it tries single newlines. Then sentences. Then words. Only as a last resort does it split mid-word.
LangChain's RecursiveCharacterTextSplitter is the most popular implementation of this, and for good reason. It's simple, predictable, and respects natural language boundaries.
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=500, # target size in characters
chunk_overlap=50, # overlap between chunks
separators=["
", "
", ". ", " ", ""]
)
chunks = splitter.split_text(my_document)
→ Chunks that respect paragraph and sentence boundariesSemantic Chunking: Splitting by Meaning
This is where things get genuinely interesting. Instead of splitting by character count or separators, semantic chunking embeds sentences individually and then looks for points where the meaning shifts significantly. When two adjacent sentences are semantically very different, that's where you draw the boundary.
The upside: your chunks actually represent coherent topics. A chunk about pricing won't have a sentence about shipping policies jammed into the end of it just because it happened to follow in the document.
The downside: it's slower and more expensive because you're embedding every sentence just to figure out where to split. For a one-time index build on a moderately sized document set, that's usually fine. For real-time chunking of user uploads? You might want to think twice.
When to Use Semantic Chunking
Use semantic chunking when your documents mix multiple topics in unstructured ways — like long blog posts, transcripts, or reports. For structured docs like FAQs or product manuals, simpler strategies often work just as well.
Document-Aware Chunking: Respect the Structure
Some documents come with structure built in — headers, sections, bullet points, tables. If you ignore that structure and apply a generic splitter, you're throwing away a huge signal.
For markdown documents, split on heading levels first. A section under an ## Pricing header should probably be its own chunk (or set of chunks), not mixed with the ## Refund Policy section.
Even better: when you chunk by section, prepend the section title to the chunk text. This means every chunk carries its own context about where it came from.
# Instead of just storing the chunk text:
"Refunds are processed within 5-7 business days..."
# Store it with its heading context prepended:
"Refund Policy: Refunds are processed within 5-7 business days..."
# Now the embedding captures the TOPIC, not just the sentence.How Overlap Actually Helps (And When It Hurts)
Overlap is your safety net for boundary problems. Let's say a document says:
More tutorials in this category, or explore the full field guide.