← All posts

Tutorials · 12 min read

RAG that actually retrieves: chunking, ranking and testing

Most RAG problems are retrieval problems in disguise. A practical walkthrough of chunking, hybrid search, re-ranking and how to test retrieval separately from generation.

Subhash Nunna ·

When a retrieval-augmented generation (RAG) system gives a bad answer, teams usually blame the model. More often, the model never saw the right information. Fix retrieval first.

Chunk for meaning, not size

Splitting documents every N characters cuts sentences, tables and arguments in half. Better defaults:

  • split on the document’s own structure — headings, sections, list items
  • keep chunks small enough to be specific, but large enough to stand alone
  • attach metadata to every chunk: title, section path, date, source URL

Prepending the document title and section heading to each chunk often improves retrieval noticeably on its own.

Vector search is good at meaning; keyword search is good at exact terms like product codes, names and error messages. Combine both and merge the results. Hybrid search is a strong default for most enterprise content.

Re-rank the top results

Retrieve generously (say, the top 30–50 candidates), then use a re-ranking model to pick the handful that best answer the question. Re-rankers are slower but far more precise, and you only run them on a short list.

Test retrieval on its own

Build a small set of questions where you know which documents hold the answer. Measure whether those documents appear in your top results — a simple “hit rate at k”. This tells you whether a bad answer is a retrieval failure or a generation failure, and stops you tuning prompts to fix a search problem.

Then tune generation

Once the right context reliably arrives, ask the model to answer only from that context, cite its sources, and say clearly when the context doesn’t contain the answer. That last instruction does more for trust than almost anything else.