Chunk size gets all the attention because it's the one knob everyone knows. But plenty of RAG pipelines have their chunking right and still return garbage — because the failure is somewhere else in the retrieval path.
Chunk size is the first thing everyone tunes in a RAG pipeline, and rightly so — it's the most common single cause of bad retrieval. But it's also become a comfortable scapegoat. Teams tune chunk size, see some improvement, and stop, while their pipeline still returns confidently wrong answers. If you've got your chunking sensible and quality is still poor, the problem has moved downstream. Here are the failure modes that live in the rest of the retrieval path.
First, what RAG is doing
Keep the mechanism in view, because every failure below is a break in one link of it:
"Retrieval-augmented generation (RAG) is a technique that enables large language models (LLMs) to retrieve and incorporate new information from external data sources. With RAG, LLMs first refer to a specified set of documents, then respond to user queries."
— Wikipedia, "Retrieval-augmented generation" (CC BY-SA 4.0)
Retrieve, then respond. Chunk size affects what's retrievable. These other failures affect whether the right thing is actually retrieved and then actually used.
1. The query/document phrasing mismatch
Vector search matches on semantic similarity, and users don't phrase questions the way your documents phrase answers. A user asks "why won't it turn on"; the manual says "troubleshooting power supply failures." A human sees these as the same topic; two embeddings of them can sit surprisingly far apart. The chunk containing the answer exists, is well-sized, and still doesn't get retrieved because the query vector didn't land near it. Fixes live at the query end: rewriting or expanding the user's question before searching, or generating a hypothetical answer and searching with that instead of the raw question. Chunk size can't help here — the retrieval missed for a reason that has nothing to do with it.
2. Pure vector search misses exact terms
Embeddings are great at meaning and bad at literals. Search for an error code like E-4021, a specific product SKU, or a person's surname, and semantic similarity can shrug — it retrieves things that are "about" the topic while missing the chunk containing the exact string you need. This is why serious pipelines use hybrid search: run a keyword search (which nails exact terms) alongside the vector search (which catches meaning), then combine the results. If your RAG is bad at codes, names, and identifiers, it's not a chunking problem; it's a missing keyword-search leg.
3. No reranking of the top results
Vector search gives you the top-k chunks by raw similarity, and raw similarity is a rough sort. The genuinely most relevant passage might be your 8th result, not your 1st — but if you only pass the top 3 to the model, it never sees it. A reranking step fixes this: retrieve a generous set (say the top 20), then run a more expensive, more accurate model over just those to reorder them, and pass the true top few to the LLM. Skipping reranking means trusting the cheapest sort to be the final answer, and it often isn't. The right chunk was retrieved — it just wasn't ranked high enough to be used.
4. The answer isn't grounded
The last failure happens after retrieval works perfectly. The right chunks are in the context window, and the model still answers from its own training instead of the documents, or blends the two and states something the sources don't support. This is a grounding and prompting problem: the instruction has to insist the model answer only from the provided context, say "I don't know" when the context doesn't cover it, and ideally cite which chunk it used. A RAG pipeline that retrieves flawlessly and then hallucinates on top of good context has a prompt problem, not a retrieval one — and it's the most dangerous failure, because the answer looks well-sourced while being wrong.
Diagnose the path, not just the chunks
When RAG underperforms, resist the reflex to re-tune chunk size again. Start by confirming the chunks themselves are coherent in a RAG text chunker — then, if they are, move down the path. Use a token counter to check you're not silently truncating retrieved context before it reaches the model (a "grounding" failure that's really a budget failure). And when you're iterating on the generation prompt to force grounding and citations, a prompt diff shows exactly what changed between versions so you can tell which edit actually improved the answers. Chunk size is the first knob. Once it's right, the interesting failures — and the biggest wins — are everywhere else in the pipeline.
← All articles