back
RAG That Works
A four-part series on the RAG techniques that go beyond chunk-embed-retrieve — organised by where you intervene, with honest numbers on what each one actually buys you.
- 01Why Basic RAG Fails - and Whether You Still Need ItContext windows are now measured in millions of tokens, so is retrieval obsolete? No. A taxonomy of how basic RAG actually fails, the three places you can fix it, and why you have to measure before you change anything.8 min
- 02RAG at Ingestion Time - Chunking, Context and What Actually PaysStructure-aware chunking, Anthropic's contextual retrieval at a verified 35-67% failure reduction, late chunking's much smaller payoff, hierarchical parent-child retrieval, and when fine-tuned embeddings are worth it.9 min
- 03RAG at Query Time - Rewriting, Multi-Query, Reranking and Self-ReflectionThe query-time half of RAG, where Spring AI 2.0 does most of the work. Query transformers, MultiQueryExpander, DocumentPostProcessor reranking, self-reflective loops and agentic retrieval, with the latency and cost each one adds.8 min
- 04When Vectors Aren't Enough - Hybrid Search, Knowledge Graphs and ChoosingEmbeddings are bad at exact tokens and blind to relationships. Hybrid search with BM25, knowledge graphs for connected data, and a symptom-to-strategy table for picking what to actually build.7 min