Skip to content
Go back

RAG, Step by Step: What Each Part Does and Where It Broke

Edit page

RAG, retrieval-augmented generation, is how you get a language model to answer from your own documents instead of its memory. I built a small one for my eval gate project: a fictional product manual, a chunker, two kinds of search, a reranker and a model. This post walks through each part in the order a question passes through it: what it does, why it’s there, and where it broke.

Some of the breaks were the obvious kind. The most surprising one wasn’t. For half the test questions, search ranked a chunk opening with a document title, like “Nyx R7 Technical Specifications”, above the chunk with the actual answer, because the questions and the title share the product name. The numbers in this post come from re-running the retrieval step by step and checking where the right chunk landed.

How a question passes through my RAG: vector search and BM25 run side by side, results merged and deduplicated, reranked, top 4 sent to the LLM

What RAG solves

Ask a general model about the Nyx R7 manual and it can’t know the answer. Its knowledge stops at a training cutoff, and a product manual it never saw isn’t in there at all. Worse, it usually won’t say so. It will produce a confident, plausible answer that is simply made up.

RAG fixes that by handing the model the right pages at question time. Retrieval finds the pieces of the manual most related to the question. Those pieces get pasted into the prompt as plain text, next to the question. The model answers from what it was just given, not from memory. Vectors only matter for the search step. The model itself never sees one.

Here’s what the model actually receives for the weight question, apart from its instructions: the four chunks retrieval picked, then the question.

Context:
[1] ## Physical
The R7 weighs 1.15 kg with the battery installed and measures 180 x 95 x 62 mm. It carries an IP54 rating (dust-protected and splash-resistant) and operates in temperatures from -10 C to 45 C. The current firmware version is 3.2.1.

[2] # Nyx R7 Technical Specifications
## Ranging and accuracy
The Nyx R7 has a measurement range of 0.3 m to 120 m. Range accuracy is plus or minus 3 mm at 50 m. ...

[3] ## Kit and pricing
The R7 base kit sells for 8,900 EUR and ships with the scanner, one 5,200 mAh battery, a USB-C charger, and a hard case.

[4] # Nyx R7 Portable LiDAR Scanner
The Nyx R7 is a handheld LiDAR scanner made by Corvid Instruments, first released in 2025. ...

Question: How much does the Nyx R7 weigh?

Only the first chunk answers it. The other three ride along because retrieval always keeps four, and two of them are the title chunks this post keeps coming back to.

Why not paste the whole manual?

Why not paste the whole manual into every prompt? For a real document set, three reasons. Size: a product’s docs, tickets and wiki run far past what fits in a context window. Cost: every token in the prompt is paid for on every question, including the 95% that has nothing to do with it. Quality: the more irrelevant text the model reads, the easier it is to miss the one sentence that matters, or to answer from a similar-looking wrong one.

To be straight about my own demo: the Nyx R7 manual is four short files, about 1,200 tokens. It would fit in a prompt easily. It stands in for a real doc set, so the pipeline behaves the way it would at scale.

Chunking

Retrieval doesn’t search documents, it searches chunks, so how you cut the manual decides what can be found. Too big, and one chunk covers several topics: its embedding mixes all of them and points at none, so a question about weight loses to more focused chunks. Too small, and you get fragments like a bare title line that match every question and crowd out the real answer. I hit both, in that order (full story). The fix was to split on the document’s own structure, its section headings, instead of a character count.

Too big vs right size: one 800-character chunk mixing storage, export and weight loses the weight question; one topic per chunk finds it

Embeddings and similarity

An embedding turns text into a vector, a long list of numbers that works like coordinates on a map. Texts with similar meaning land close together. To find the right chunk, the question gets embedded too, and cosine similarity compares its direction with every chunk’s direction: pointing the same way means similar meaning. The chunks are sorted by that score. In the simplest version of RAG, the top few go straight to the model.

Text as points on a map: the question and the weight sentence point in almost the same direction (small angle), the price sentence points elsewhere (big angle)

Why vectors alone weren’t enough

That simplest version was my first one, and it was close but not right. Pure vector search found the right chunk for every question, inside the top four, but on 7 of 14 it ranked it below something else. This is where each question’s right chunk landed, first with vectors alone, then with each fix:

QuestionVectorsBM25Hybrid + reranker
Operating temperature411
Battery life321
Weight4171
IP rating322
Export formats211
Warranty313
Error code E03311
7 other questions111

In all seven misses, the winner was the same kind of chunk: the opening chunk of a document, which carries the document title, like ”# Nyx R7 Technical Specifications” or ”# Nyx R7 Setup Guide”. Every question also says “the Nyx R7”. So those chunks matched every question on the product name, not on the answer. The chunk that actually said “The R7 weighs 1.15 kg” had no title and lost, 0.52 against 0.63. The fix from my last post, folding a lone title into the block after it, removed the tiny title-only chunks but not their pull. I haven’t proven this by re-embedding without the titles. It’s what the rankings point to.

Here’s the weight question with vectors alone, the top four and their scores:

RankScoreChunk
10.626# Nyx R7 Technical Specifications / ## Ranging and accuracy
20.585# Nyx R7 Setup Guide / ## First use
30.553# Nyx R7 Portable LiDAR Scanner
40.518## Physical: “The R7 weighs 1.15 kg…”

Three document openers, none about weight, all ahead of the one chunk that answers.

Hybrid retrieval and reranking

BM25 is an old keyword-search method, from long before embeddings, and it covers exactly where they’re weak. It scores chunks by which words of the question they contain, and rare words count for more. “R7” is in 19 of the 21 chunks, so matching it says almost nothing. “IP54” is in one, so a match on it counts heavily. So I run both searches side by side: the top eight chunks by meaning, the top eight by keywords, merged into one candidate list with duplicates removed, since a chunk can rank high in both. Neither list goes straight to the model. A reranker reorders the merged list first and keeps the best four.

The merge is a few lines in hybridRetrieve():

const seen = new Set<string>();
const candidates: Chunk[] = [];
for (const c of [
  ...retrieve(queryEmbedding, chunks, CANDIDATE_K),
  ...bm25Search(query, bm25, CANDIDATE_K),
]) {
  if (seen.has(c.id)) continue;
  seen.add(c.id);
  candidates.push(c);
}

BM25 has its own hole, and the table shows it. It matches words exactly, with no stemming, so the question’s “weigh” never matched the chunk’s “weighs”, and the weight chunk fell to 17th.

Bi-encoder vs cross-encoder: the bi-encoder embeds question and chunk separately and compares vectors; the cross-encoder reads both together and outputs a relevance score

The two searches are fast because neither one reads the question and a chunk together. Every chunk was embedded once, ahead of time. At question time only the question gets embedded and compared. That’s a bi-encoder: quick enough to scan everything, but each side is squeezed into a vector without seeing the other. A cross-encoder is the opposite. It reads the question and one chunk as a single input and outputs a relevance score, so it can see how the words of each relate. It’s more accurate, but nothing can be precomputed: every question and chunk pair is a full model run. So it only judges the shortlist. The fast searches gather up to sixteen candidates, the reranker (bge-reranker-base, running locally) scores them, and the top four go to the model.

It rescued the weight question from BM25’s 17th place. It isn’t perfect either: BM25 had the warranty chunk first, and the reranker moved it to third, still inside the four the model sees. With BM25 and the reranker added, mean Contextual Precision went from about 0.65 to about 0.90, with recall unchanged. I didn’t measure the two separately, so I can’t say how much each one contributed.

Generation

The model answers under a strict prompt: answer only from the given context, and if the answer isn’t there, reply with one exact sentence, “I don’t know based on the provided documents.” The guards in code return that same sentence when they block a question before the model, or catch an answer leaking the context after it. So every refusal looks identical, whoever made it, and a test can check for it with a plain string compare. That’s also its weakness: the model sometimes rephrases the sentence, and the check goes red. I covered that trade in the last post.

Evaluating retrieval

Two metrics grade the retrieval step itself, before the answer even matters. Contextual Recall asks whether the retrieved chunks contain everything needed for the correct answer. Contextual Precision asks whether the right chunks sit at the top, above the useless ones. Pure vector search passed recall on all 14 questions, because the right chunk always made the top four. It failed precision on 7, most likely because the right chunk was there but not first, sitting behind the title chunks. That’s the gap BM25 and the reranker closed together: recall stayed the same, mean precision went from about 0.65 to about 0.90. Both are graded by an LLM judge against an expected answer, which has its own problems. More on that in the last post.

The rank table above is a different, simpler measurement: no judge, just where the chunk with the answer landed. I re-ran it on the current version of the corpus while writing this post, so it shows the same pattern as the judged run, not necessarily the same seven questions.

Where this stops working

This pipeline is small on purpose, and each shortcut has a limit. The index is a JSON file, and every question is compared against every chunk. At 21 chunks that’s instant. At millions it isn’t, and that’s where a vector database comes in: an approximate index that trades a little accuracy for speed. The corpus is four short, made-up files with 14 test questions, so the 0.65 to 0.90 jump is a direction, not a benchmark. And questions go to search exactly as typed. With no stemming, “weigh” missed “weighs”. Real systems usually fix that by normalising words, or by having a model rewrite the question before searching.


Edit page
Share this post on:

Next Post
Same Bot, Same Questions, Different Answer