RAG lab
When an assistant built on company documents gets an answer wrong, the model takes the blame, but often the right passage never reached it. Six exhibits take retrieval apart on a small handbook and measure every step, from cutting the text to evaluating the whole pipeline.
Everything runs in this tab, including an embedding and a reranker trained here. No language model runs. The reader in Exhibit VI quotes one sentence, so every number is a retrieval computation you can check. Ostler Instruments and the 80 labelled questions are fictional, and only your progress is saved.
Lab exploredThat's the whole retrieval pipeline, from the first cut to the last lost answer. For the other way to retrieve, following links instead of matching words, try the Graph lab.
Where the cuts land
A retriever ranks chunks, pieces small enough to fit in a model's context window. Every way of cutting has a cost. A fixed word count can split the one sentence that answers a question, and cutting at sentences or headings gives uneven pieces.
Below, we cut the handbook in this tab and check whether each test answer survives inside one chunk.
- Chunks
- not cut yet
- How many
- Press the button to cut the handbook.
The labelled questions: 80, split 50 train and 30 test
| Id | Split | Kind | Question | Answer passage |
|---|
The maths
chunks in a document of N words = ⌈(N − o) / (s − o)⌉ (one chunk when N ≤ s)
P(a span of L words is cut) ≈ (L − 1 − o) / (s − o) when o < L − 1, 0 when o ≥ L − 1
Why the count. Chunks start every s − o words and the last must reach the end, which gives ⌈(N − s) / (s − o)⌉ + 1 starts. Why the cut chance. The span's offset r from the last chunk start is spread evenly over a stride of s − o positions, and the span survives when r + L ≤ s. The rest, (L − 1 − o) / (s − o), is the chance of a cut, and an overlap of L − 1 words makes it zero.
Why only about. A span near the end of a document sits in the last, shorter chunk and is always safe, and a span longer than s is always cut. Sentence and heading chunkers never cut inside a sentence, so an answer that is one sentence long always survives them.
In practiceMany teams start with chunks of a few hundred tokens and a small overlap, then measure. Splitting at headings and sentences first, and by size only when a section is too long, keeps most answers whole without paying for overlap everywhere.
Watch outOverlap isn't free: it stores and ranks the same words twice, and near-identical neighbouring chunks crowd each other in the results. Measure the answer-intact rate on your own questions before settling on a size.
BM25 and the exact word
BM25 scores a chunk by the question's words it contains. Rare words count for more, each repeat counts for less, and long chunks are discounted. It needs no training and never misses an exact identifier, but it can't tell that two different words mean the same thing.
From here to Exhibit V we cut the handbook by heading, so every answer stays whole and only retrieval gets measured.
Search for something, or press one of the buttons.
Recall at 5 for every pair of k1 and b, measured on the training split. Click a cell to use it, or move the sliders.
- Train recall at 5
- not measured
- Best on the grid
- map the grid to see it
- Test recall at 5
- measured once you tune
The maths
BM25(q, d) = ∑t ∈ q IDF(t) × f (k1 + 1) / (f + k1 (1 − b + b |d| / avgdl))
IDF(t) = ln((N − nt + 0.5) / (nt + 0.5) + 1)
Why saturation. f is how often the word occurs in the chunk. The fraction f (k1 + 1) / (f + k1 …) is 1 for a single occurrence of an average chunk and can never pass k1 + 1, so the tenth repeat adds almost nothing, and with k1 = 0 only presence counts. Why length normalisation. A long chunk contains more words by chance. b scales the discount: with b = 1 a chunk twice the average length needs twice the occurrences for the same credit, with b = 0 length is ignored and long pages rise.
Why this IDF. N chunks, nt contain the word: a word in every chunk gets almost nothing, a word in one chunk gets the most, and the + 1 inside the logarithm keeps every weight positive. Words that appear nowhere contribute zero, which is why a paraphrase scores nothing.
In practiceBM25 grew out of Robertson's Okapi experiments in the 1990s and is the default ranking in Lucene, Elasticsearch and OpenSearch. It's the baseline every other retriever has to beat.
Watch outKeywords match word forms exactly: locks doesn't match lock, and a customer's word for a thing doesn't match the handbook's. Stemming, synonym lists and a second, semantic retriever are the usual fixes. Tune k1 and b on training questions, never on the ones you report.
Meaning from co-occurrence
An embedding puts text in a space where nearby points mean similar things. This one is learned in the tab by latent semantic analysis. The handbook is cut into 100 overlapping windows of 256 words, and a truncated SVD of their TF-IDF weights keeps the strongest patterns. Words that keep the same company end up close together.
Production systems use neural embedding models trained on far more text. The retrieval maths is the same: embed the question, embed every chunk, rank by cosine similarity.
The maths
cos(q, d) = q · d / (‖q‖ ‖d‖)
A ≈ Uk Sk VkT, q̂ = qT Uk Sk−1
The decomposition. A has one row per word and one column per training window. Keeping the k largest singular values gives the rank-k matrix closest to A in squared error (Eckart and Young), and the share of variance each dimension carries is si2 / ‖A‖2. The page computes it by randomised subspace iteration (multiply by a seeded random matrix, orthonormalise, repeat three times), then solves a 74 by 74 problem exactly. That recovers the leading singular values to a fraction of a percent.
Folding in. A question or a chunk is a column of TF-IDF weights like any window, so qT Uk Sk−1 gives it coordinates in the same frame as the rows of Vk. Similarity is the cosine after weighting each coordinate by its singular value, which is the cosine of the two projections UkTq and UkTd. Why co-occurrence. A word's vector is its row of UkSk. Words that share windows or neighbours load on the same dimensions, so mount and tripod head land near each other.
In practiceProduction systems use neural embedding models: every chunk is embedded once, and approximate nearest-neighbour indexes keep search fast at millions of chunks. Latent semantic analysis (Deerwester and colleagues, 1990) is one of their ancestors.
Watch outAn embedding smooths away exactly what makes an identifier useful. E-4107 and E-4170 keep the same company, so they land in the same place. Keep a keyword retriever for codes, part numbers and versions, and embed the corpus again when it changes.
Two rankers beat one
BM25 finds exact words and the embedding finds meaning, so they fail on different questions. Reciprocal rank fusion adds 1 / (60 + rank) from each list, so a chunk both rankers like rises. Score fusion rescales both scores to between 0 and 1 and takes a weighted average.
Everything here is measured on the 30 test questions, 26 of which have an answer in the handbook, with BM25 at k1 = 1.2 and b = 0.75 and the embedding at k = 48.
| Ranker | R@1 | R@3 | R@5 | R@10 | MRR | nDCG@10 |
|---|
Each square is a test question with three marks (BM25, the embedding and your fusion), lit when that ranker has the answer in its first k. Pick one to see its lists above.
The maths
RRF(d) = ∑r 1 / (60 + rankr(d)), fused(d) = w × mm(BM25) + (1 − w) × mm(cos), mm(x) = (x − min) / (max − min)
recall@k = answered in the first k / questions, MRR = mean of 1 / (rank of the first answer), nDCG@k = DCG / IDCG, DCG = ∑i ≤ k reli / log2(i + 1)
Why ranks. BM25 scores run from 0 to about 20 and cosines from −1 to 1, but ranks need no rescaling. The constant 60 keeps first place from dominating, so second place in both lists beats first place in one. Each ranker contributes its first 20 results here. Why min-max. Score fusion needs both scores on one scale before a weight means anything. Why the log. nDCG credits an answer less the lower it sits, by 1 / log2(i + 1), and divides by the best possible order so a perfect ranking scores 1.
In practiceReciprocal rank fusion comes from Cormack, Clarke and Buettcher (2009), who used the constant 60. Several search engines and vector databases build it in because it needs no tuning.
Watch outA fusion weight chosen by looking at the test questions flatters itself. Choose it on the training split, then measure once on the test split and report that number.
A second look at the top 20
Fusion gives you good candidates in a rough order. A reranker takes another look at the top 20, using features the first stage couldn't combine: BM25, the cosine, an identifier match, a title match, page age and position. It learns their weights by logistic regression on the training split.
Maximal marginal relevance then trades a little relevance for variety, so three copies of one policy don't fill the context.
Not trained yet.
| Split | Order | Top-1 accuracy | P@3 | R@5 | MRR |
|---|
The maths
p = σ(w · x + b), σ(z) = 1 / (1 + e−z)
L = (1/m) ∑ [ log(1 + ezi) − yi zi ] + (λ2/2) ‖w‖2, ∂L/∂w = (1/m) ∑ (pi − yi) xi + λ2 w
MMR: next = argmaxd λ sim(q, d) − (1 − λ) maxs ∈ picked sim(d, s)
Why this gradient. The derivative of the log loss in z is p − y, so the model gets pushed up on chunks that hold the answer and down on the rest, in proportion to how wrong it was. The features are standardised on the training rows first, so one learning rate suits them all. Why held-out questions. Scores on the training questions are optimistic. Only questions the model never saw say how it will do on new ones.
Why MMR works. sim(q, d) is the reranker's probability and sim(d, s) the TF-IDF cosine between chunks. The first pick is the most relevant. After that, a chunk nearly identical to one already picked pays for it, and with λ low enough an outdated copy falls out of the top three.
In practiceProduction rerankers are usually cross-encoders, neural models that read the question and the chunk together. They're slow, which is why they only see the first few dozen candidates from a fast first stage, like here.
Watch outOld and new versions of a policy look almost identical to every retriever. Keep dates as metadata, prefer the current version explicitly, and retire superseded pages instead of leaving both searchable.
Where the answers are lost
Run the whole pipeline on the 30 test questions: cut, retrieve, pack the best chunks into a fixed context window, and let a simple reader quote the sentence that overlaps the question most. Every miss gets a diagnosis, because an answer that never reached the context can't have come from any model. blame the retriever first lol.
- Accuracy
- not run
- 95% interval
- Bound
| Id | Question | Outcome | Rank |
|---|
| Run | Settings | Correct | 95% interval |
|---|
The maths
accuracy on answerable questions ≤ recall at the budget, kbudget = the most ranked chunks whose words fit the window
Wilson interval: (p̂ + z2/2n ± z √(p̂(1 − p̂)/n + z2/4n2)) / (1 + z2/n), z = 1.96
Why the bound. The reader only sees what was packed into the window, so a question can be answered only if an answer passage is in the first kbudget chunks. Out-of-corpus questions are right only when the reader says "not found", which a threshold on its word coverage decides. Why Wilson. The simple interval p̂ ± 1.96 √(p̂(1 − p̂)/n) breaks down on small samples and near 0 or 1, while Wilson's interval inverts the score test and stays inside 0 to 1. On 30 questions it is about ±17 points wide, so settings a few answers apart can't be told apart.
In practiceMeasure retrieval on its own first (recall at the context budget), then the reader. Keep a fixed labelled set that includes questions the corpus cannot answer, and report intervals with every estimate.
Watch outThirty questions give a wide interval, and two settings a few answers apart are usually indistinguishable. Compare settings on the same questions with a paired test, and grow the set before trusting a small difference.