Retrieval-augmented generation connects a language model to a searchable collection of documents. The useful engineering question is how to tell whether that connection is working. This is a proposed evaluation method, with an illustrative example—not a report of measured project results.
Explore this article visually3 figures
01
Work through one retrieval score
Suppose a question requires passages A and B, and the first three results are A, C and D. Recall@3 is 1/2, while hit rate for this question is 1: at least one relevant passage was found. The answer still lacks B. Now supply A and B directly to the generator. If the answer becomes complete, investigate retrieval; if it still omits B, inspect context use and answer construction. Keep this two-passage case separate from questions that need only one supporting passage.
02
The answer can be right for the wrong version
Figure 01Comparison
Two valid citations. Only one current answer.
Read every explanation
- Retired guide
- The cited passage really says 30 days, so a support-only evaluator may accept the answer. It still fails a question about the active policy.
- Active guide
- The active guide supplies the relevant answer for this example. Preserve its version and effective date so the reader can inspect why it was selected.
The cited passage really says 30 days, so a support-only evaluator may accept the answer. It still fails a question about the active policy.
Imagine an internal documentation assistant asked: ‘How long do we retain service logs?’ Its index contains a retired guide saying 30 days and an active guide saying 14 days. The assistant retrieves the retired guide, answers ‘30 days’ and attaches a perfectly functioning citation. The sentence is supported by the retrieved passage, but it does not answer the current question correctly.
That example separates three decisions: which documents were eligible, which passages were retrieved, and what the model concluded from them. Rewriting the generation prompt cannot reliably fix a missing active document. Increasing retrieval depth cannot settle which policy is authoritative unless version and status are available.
03
Build the test set around evidence
Start with a small reviewed set of real question types. For each case, record the question, the caller’s permitted scope, the relevant document version, the supporting passage and the expected behaviour. An expected answer alone is insufficient: a system could reproduce it from model memory while failing to retrieve any evidence.
| Case | Evidence condition | Expected behaviour |
|---|---|---|
| Direct lookup | One active passage answers the question | Answer and cite that passage |
| Version conflict | Current and retired guides disagree | Use the active version; explain the distinction if relevant |
| Two-document question | Each document supplies part of the answer | Retrieve both and support each claim |
| Missing answer | No permitted passage contains the fact | State the gap and ask a useful follow-up |
| Restricted source | The answer exists outside the caller’s scope | Do not reveal the restricted content or metadata |
Separate development cases from a held-out set. Keep paraphrases of one question in the same split, or tuning on one phrasing can leak into the test through another. Version the corpus snapshot as well as the questions: otherwise a score change might come from new documents rather than a better retrieval configuration.
04
Measure retrieval without the writer
Figure 02Data flow
The evidence path behind a RAG answer
Read every explanation
- Scope
- Resolve the caller’s access and the document version appropriate to the question. A relevant but forbidden passage is not an eligible result.
- Retrieve
- Retrieve and rank eligible passages. Preserve source IDs and effective dates so version conflicts remain visible downstream.
- Compose
- Build the answer from the supplied context. Log the exact context after truncation; retrieving a passage is not the same as showing it to the model.
- Check
- Check claim support separately from source currency. If the evidence cannot answer the question, expose the gap rather than inventing a bridge.
Resolve the caller’s access and the document version appropriate to the question. A relevant but forbidden passage is not an eligible result.
Run retrieval alone and inspect the passages delivered to generation, after filtering and reranking. Recall@k asks what fraction of the labelled relevant items appeared in the first k results. Hit rate asks whether at least one appeared. They answer different questions: one good passage can satisfy a direct lookup while leaving a two-document question incomplete.
def recall_at_k(retrieved_ids, relevant_ids, k):
if k <= 0:
raise ValueError("k must be positive")
relevant = set(relevant_ids)
if not relevant:
return None # evaluate unanswerable cases separately
found = set(retrieved_ids[:k]) & relevant
return len(found) / len(relevant)Freeze the passage IDs for a comparison. If chunking changes, remap labels to stable source spans; otherwise the metric penalises new identifiers rather than missing evidence. Inspect cases where a useful sentence loses its heading, table columns or effective date at a chunk boundary. More chunks do not help if they strip away the context that makes a passage interpretable.
Compare a lexical baseline with vector retrieval and a hybrid candidate under the same corpus, permissions and context budget. Exact error codes and product identifiers are useful test cases. Treat the choice as an experiment on your questions, not a universal ranking of retrieval methods.
05
Then test what the writer does with evidence
Figure 03Comparison
Change the evidence to locate the failure
Read every explanation
- Reviewed passages
- If generation fails even with the reviewed supporting passages, investigate instructions, answer construction or an ambiguous reference answer.
- Retrieved passages
- If reviewed context succeeds but retrieved context fails, inspect ranking, eligibility, missing passages and truncation before changing the model.
If generation fails even with the reviewed supporting passages, investigate instructions, answer construction or an ambiguous reference answer.
Give the generator the reviewed supporting passages directly. If it still fails, inspect the instructions and answer construction. If it succeeds with those passages but fails with retrieved context, investigate retrieval, ranking or truncation. This controlled comparison narrows the problem without changing several components at once.
Review each factual claim against its cited passage. A citation’s presence is not evidence of support, and support is not evidence that the source is current. Track answer correctness, claim support and citation coverage independently. Automated evaluators can help triage failures, but their judgements need checks against human labels, especially for partial support and contradictory sources.
06
Abstention is useful only if it is measured
A system that always declines produces few unsupported claims and little value. Report both unsupported answers on unanswerable questions and unnecessary refusals on answerable questions. Keep access-denied cases separate from genuinely missing evidence. The user-facing response should explain the available next step without revealing the existence or title of a restricted document.
Enforce access rules before content reaches the model, and include identity and scope in cache isolation. Retrieved documents remain data: text inside them cannot authorise tool calls or change application permissions. Add test documents containing instruction-like text to check this boundary.
07
Make the next change diagnosable
- Freeze the questions, source snapshot and access scopes; review the evidence labels.
- Save retrieved passage IDs, ranks, source versions and the exact context supplied to generation.
- Change one component, such as chunking or reranking, and compare failures by question type.
- Report retrieval recall, answer correctness, unsupported claims and unnecessary refusals alongside latency and cost.
- Inspect newly broken cases before accepting a better average; keep confirmed failures as regression cases.
The release decision is whether the system answers more of the intended questions with the right evidence, within its operating budget. A polished paragraph is the visible result. The test case should make the evidence behind it equally visible.
More articles
09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read01Building agentic AI beyond the demoAgentic AI · 9 min read02Designing AI that works without a networkEdge AI · 8 min read