Retrieval-augmented generation connects a language model to a searchable collection of documents. The useful engineering question is how to tell whether that connection is working. This is a proposed evaluation method, with an illustrative example—not a report of measured project results.

Explore this article visually3 figures

01

Work through one retrieval score

Suppose a question requires passages A and B, and the first three results are A, C and D. Recall@3 is 1/2, while hit rate for this question is 1: at least one relevant passage was found. The answer still lacks B. Now supply A and B directly to the generator. If the answer becomes complete, investigate retrieval; if it still omits B, inspect context use and answer construction. Keep this two-passage case separate from questions that need only one supporting passage.

02

The answer can be right for the wrong version

Figure 01Comparison

Two valid citations. Only one current answer.

Select an element to explore its role.
Read every explanation
Retired guide
The cited passage really says 30 days, so a support-only evaluator may accept the answer. It still fails a question about the active policy.
Active guide
The active guide supplies the relevant answer for this example. Preserve its version and effective date so the reader can inspect why it was selected.
01 / 02
Retired guide

The cited passage really says 30 days, so a support-only evaluator may accept the answer. It still fails a question about the active policy.

Invented retention periods from the article’s example. This illustrates a version conflict, not a retention recommendation.

Imagine an internal documentation assistant asked: ‘How long do we retain service logs?’ Its index contains a retired guide saying 30 days and an active guide saying 14 days. The assistant retrieves the retired guide, answers ‘30 days’ and attaches a perfectly functioning citation. The sentence is supported by the retrieved passage, but it does not answer the current question correctly.

That example separates three decisions: which documents were eligible, which passages were retrieved, and what the model concluded from them. Rewriting the generation prompt cannot reliably fix a missing active document. Increasing retrieval depth cannot settle which policy is authoritative unless version and status are available.

03

Build the test set around evidence

Start with a small reviewed set of real question types. For each case, record the question, the caller’s permitted scope, the relevant document version, the supporting passage and the expected behaviour. An expected answer alone is insufficient: a system could reproduce it from model memory while failing to retrieve any evidence.

CaseEvidence conditionExpected behaviour
Direct lookupOne active passage answers the questionAnswer and cite that passage
Version conflictCurrent and retired guides disagreeUse the active version; explain the distinction if relevant
Two-document questionEach document supplies part of the answerRetrieve both and support each claim
Missing answerNo permitted passage contains the factState the gap and ask a useful follow-up
Restricted sourceThe answer exists outside the caller’s scopeDo not reveal the restricted content or metadata

Separate development cases from a held-out set. Keep paraphrases of one question in the same split, or tuning on one phrasing can leak into the test through another. Version the corpus snapshot as well as the questions: otherwise a score change might come from new documents rather than a better retrieval configuration.

04

Measure retrieval without the writer

Figure 02Data flow

The evidence path behind a RAG answer

Select an element to explore its role.
Read every explanation
Scope
Resolve the caller’s access and the document version appropriate to the question. A relevant but forbidden passage is not an eligible result.
Retrieve
Retrieve and rank eligible passages. Preserve source IDs and effective dates so version conflicts remain visible downstream.
Compose
Build the answer from the supplied context. Log the exact context after truncation; retrieving a passage is not the same as showing it to the model.
Check
Check claim support separately from source currency. If the evidence cannot answer the question, expose the gap rather than inventing a bridge.
01 / 04
Scope

Resolve the caller’s access and the document version appropriate to the question. A relevant but forbidden passage is not an eligible result.

A proposed architecture. Scope and source eligibility are enforced before passages reach the generator.

Run retrieval alone and inspect the passages delivered to generation, after filtering and reranking. Recall@k asks what fraction of the labelled relevant items appeared in the first k results. Hit rate asks whether at least one appeared. They answer different questions: one good passage can satisfy a direct lookup while leaving a two-document question incomplete.

pythonA minimal recall calculation over labelled passage IDs; this does not score the answer.
def recall_at_k(retrieved_ids, relevant_ids, k):
    if k <= 0:
        raise ValueError("k must be positive")
    relevant = set(relevant_ids)
    if not relevant:
        return None  # evaluate unanswerable cases separately
    found = set(retrieved_ids[:k]) & relevant
    return len(found) / len(relevant)

Freeze the passage IDs for a comparison. If chunking changes, remap labels to stable source spans; otherwise the metric penalises new identifiers rather than missing evidence. Inspect cases where a useful sentence loses its heading, table columns or effective date at a chunk boundary. More chunks do not help if they strip away the context that makes a passage interpretable.

Compare a lexical baseline with vector retrieval and a hybrid candidate under the same corpus, permissions and context budget. Exact error codes and product identifiers are useful test cases. Treat the choice as an experiment on your questions, not a universal ranking of retrieval methods.

05

Then test what the writer does with evidence

Figure 03Comparison

Change the evidence to locate the failure

Select an element to explore its role.
Read every explanation
Reviewed passages
If generation fails even with the reviewed supporting passages, investigate instructions, answer construction or an ambiguous reference answer.
Retrieved passages
If reviewed context succeeds but retrieved context fails, inspect ranking, eligibility, missing passages and truncation before changing the model.
01 / 02
Reviewed passages

If generation fails even with the reviewed supporting passages, investigate instructions, answer construction or an ambiguous reference answer.

Controlled experiment: keep the question and generator fixed, then compare supplied evidence. Review both arms, not just the average score.

Give the generator the reviewed supporting passages directly. If it still fails, inspect the instructions and answer construction. If it succeeds with those passages but fails with retrieved context, investigate retrieval, ranking or truncation. This controlled comparison narrows the problem without changing several components at once.

Review each factual claim against its cited passage. A citation’s presence is not evidence of support, and support is not evidence that the source is current. Track answer correctness, claim support and citation coverage independently. Automated evaluators can help triage failures, but their judgements need checks against human labels, especially for partial support and contradictory sources.

06

Abstention is useful only if it is measured

A system that always declines produces few unsupported claims and little value. Report both unsupported answers on unanswerable questions and unnecessary refusals on answerable questions. Keep access-denied cases separate from genuinely missing evidence. The user-facing response should explain the available next step without revealing the existence or title of a restricted document.

Enforce access rules before content reaches the model, and include identity and scope in cache isolation. Retrieved documents remain data: text inside them cannot authorise tool calls or change application permissions. Add test documents containing instruction-like text to check this boundary.

07

Make the next change diagnosable

  1. Freeze the questions, source snapshot and access scopes; review the evidence labels.
  2. Save retrieved passage IDs, ranks, source versions and the exact context supplied to generation.
  3. Change one component, such as chunking or reranking, and compare failures by question type.
  4. Report retrieval recall, answer correctness, unsupported claims and unnecessary refusals alongside latency and cost.
  5. Inspect newly broken cases before accepting a better average; keep confirmed failures as regression cases.

The release decision is whether the system answers more of the intended questions with the right evidence, within its operating budget. A polished paragraph is the visible result. The test case should make the evidence behind it equally visible.

Discuss this article

More articles

09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read01Building agentic AI beyond the demoAgentic AI · 9 min read02Designing AI that works without a networkEdge AI · 8 min read