RAG Is Just a Search Problem Wearing a Trench Coat
Everyone building retrieval augmented generation is rediscovering, one painful sprint at a time, that the hard part was never the G. The model does its job. The R is where your quality lives and dies, and the R is a search engine, which means twenty five years of information retrieval lessons apply, and most teams skip them.
The symptoms are predictable. Answers that confidently miss because the relevant chunk never made it into the context. Chunks that are technically similar to the query but useless, embeddings love topical similarity and users need answers, not topics. Great performance on the demo questions and collapse on real ones, because the demo questions were written by people who knew what was in the corpus.
The fixes are search fixes, dressed differently. Hybrid retrieval, dense vectors plus plain keyword BM25, because embeddings whiff on exact identifiers, part numbers, error codes, names, the exact things enterprise users search for. A reranker on top of first stage retrieval, retrieve 50 cheap, rerank to 5 good, this single addition moved our answer quality more than any model upgrade. Chunking that respects document structure instead of slicing every 500 tokens mid sentence, headers kept with their sections, tables kept whole. And metadata filters, because "what changed in the March release" is a filter plus a query, not a pure vector lookup.
Above all, the discipline nobody wants: measure retrieval separately from generation. Build a set of real questions with known correct sources and track whether the right chunk lands in context at all. When a RAG answer is wrong, look at what was retrieved before blaming the model. In our incident review of bad answers, it was retrieval about four times out of five.
Take the trench coat off. Do search properly. The generation mostly takes care of itself.