A RAG answer can sound convincing while using the wrong document, omitting an exception, or exposing information the caller cannot access. RAG evaluation needs to tell you which stage failed. A single score for the final answer cannot explain whether to fix ingestion, retrieval, permissions, or generation.
This guide builds a small Python evaluation harness, explains its output, and turns the results into a regression workflow. It assumes you already have a retrieval pipeline, such as the one in building a local AI application with RAG. The example uses fixed retrieval results so you can verify the evaluator without a model, API key, or database.
Measure the Boundaries Separately
Suppose a support assistant retrieves a refund policy and answers that every purchase is refundable. The policy also contains an exception for activated subscriptions. Finding the right document is a retrieval success; omitting the exception is an answer failure. If the exception was lost during chunking, the generator may never have received the evidence.
Faithfulness, also called groundedness, asks whether the answer's claims are supported by the supplied context. Correctness asks whether the answer satisfies the question against an independently reviewed reference. A faithful answer can still be wrong when its source is obsolete. A correct answer recalled from model training can still lack support in the retrieved evidence. Microsoft's RAG evaluator documentation similarly separates retrieval, groundedness, and response completeness.
Build Cases You Can Diagnose
Start with representative questions from approved, redacted support examples and domain reviewers. Include short keyword queries, paraphrases, multi-part questions, rare product names, outdated policies, and ambiguous requests. Synthetic questions can expand coverage, but review them: a question copied from a heading may make retrieval unrealistically easy.
Store a question ID, question text, caller or tenant context, corpus snapshot, expected behavior, required facts, and relevant evidence IDs. Keep source version and evidence spans alongside chunk IDs. Chunk IDs change when you change the splitter; stable source identifiers and spans let reviewers remap labels instead of silently comparing different units. The vector database and embeddings guide provides the retrieval background.
Label actual supporting passages, not every chunk from a relevant document. Record alternative valid evidence sets when several passages independently answer the question. Distinguish an unreviewed candidate from a reviewed irrelevant candidate: incomplete labels can make a valid result look like a retrieval failure.
Split cases before tuning. Use development cases to adjust chunking, prompts, and thresholds; reserve a held-out set for release comparison. Keep near-duplicate questions and paraphrases in the same split. Where appropriate, group by document family or customer scenario. Pin the corpus snapshot so a score change does not secretly reflect a different source collection. This follows the separation of development and test collections described in Stanford's information retrieval evaluation chapter.
Define Recall and Rank Before Writing Code
For this harness, Recall@k is the number of distinct labeled relevant chunks found in the first k results divided by the number of labeled relevant chunks. Duplicate results do not earn extra credit. The denominator means this measures recall against your labels, not against unknowable perfect coverage of the whole corpus.
Reciprocal rank at k is 1 divided by the position of the first relevant result within k, or zero when none appears. MRR@k is its mean across evaluated questions, consistent with Elastic's ranking evaluation definitions. A first-position hit gets full reciprocal-rank credit even when a second essential passage is missing. Use recall and a required-evidence check for questions that need several facts.
Neither metric measures whether the generator used the evidence correctly. Increasing k may improve measured recall while adding distracting text, tokens, and latency. Measure the final context after reranking and truncation as well as the retriever's original candidates.
Run a Reproducible Retrieval Evaluation
Save the following as rag_eval.py and run python rag_eval.py with Python 3. It uses only the standard library. These four hand-written cases deliberately contain failures; their scores are fixture output, not a benchmark for any retrieval product.
from statistics import mean
def retrieval_metrics(relevant, retrieved, k):
if k < 1:
raise ValueError("k must be positive")
top = retrieved[:k]
if not relevant:
return None # No relevant evidence: evaluate abstention separately.
recall = len(set(top) & relevant) / len(relevant)
rr = next((1 / rank for rank, item in enumerate(top, 1)
if item in relevant), 0.0)
return recall, rr
# IDs represent versioned chunks in one frozen corpus.
cases = [
{"id": "refund", "relevant": {"refund-v2", "exception-v2"},
"retrieved": ["refund-v2", "refund-v2", "shipping-v1"],
"forbidden": set()},
{"id": "restore", "relevant": {"restore-v3"},
"retrieved": ["restore-v1", "restore-v3", "shipping-v1"],
"forbidden": set()},
{"id": "sla", "relevant": {"sla-v2"},
"retrieved": ["shipping-v1", "restore-v1"],
"forbidden": set()},
{"id": "private-payroll", "relevant": set(),
"retrieved": ["payroll-private"],
"forbidden": {"payroll-private"}},
]
k = 3
scores = []
duplicates = leaks = unanswerable = 0
for case in cases:
top = case["retrieved"][:k]
duplicates += len(top) - len(set(top))
leaks += bool(set(top) & case["forbidden"])
result = retrieval_metrics(case["relevant"], top, k)
if result is None:
unanswerable += 1
continue
scores.append(result)
print(f'{case["id"]}: recall={result[0]:.3f} rr={result[1]:.3f}')
print(f"answerable cases: {len(scores)}")
if scores:
print(f"mean recall@{k}: {mean(s[0] for s in scores):.3f}")
print(f"MRR@{k}: {mean(s[1] for s in scores):.3f}")
print(f"duplicate slots: {duplicates}")
print(f"cases with forbidden evidence: {leaks}")
print(f"unanswerable cases requiring answer review: {unanswerable}")
# Sanity checks for the evaluator, not acceptance gates for a RAG system.
assert retrieval_metrics({"a"}, [], 3) == (0.0, 0.0)
assert retrieval_metrics({"a", "b"}, ["a", "a", "x"], 3) == (0.5, 1.0)
assert retrieval_metrics(set(), [], 3) is None
assert (duplicates, leaks, unanswerable) == (1, 1, 1)
The output is:
refund: recall=0.500 rr=1.000
restore: recall=1.000 rr=0.500
sla: recall=0.000 rr=0.000
answerable cases: 3
mean recall@3: 0.500
MRR@3: 0.500
duplicate slots: 1
cases with forbidden evidence: 1
unanswerable cases requiring answer review: 1
The refund case exposes why MRR alone is insufficient: the first hit is relevant, yet half the labeled evidence is missing. Repeated chunks occupy real result slots; the harness keeps those positions rather than removing duplicates before ranking. The payroll case is excluded from retrieval averages because its relevant set is empty, but its permission failure remains visible.
To use this with your pipeline, replace each fixed retrieved list with IDs returned under the recorded caller identity. Keep question text and labels in a versioned dataset. Capture intermediate candidates and the final model context separately. The small forbidden-ID sets here are test assertions; production authorization must check every returned document against the actual permission policy.
Test Unanswerable and Unauthorized Requests
A question can have no answer in the authorized corpus even when similar documents exist. Include nonexistent features, missing dates, misleading premises, and requests requiring private evidence. Require an appropriate refusal, clarification, or statement of insufficient evidence. Also measure unnecessary abstention on answerable questions; an assistant that refuses everything should not pass.
Run the same question as an allowed user, a different tenant, and a user whose access was revoked. Check retrieval, assembled context, citations, and cached responses. Microsoft's document access control guidance places permission enforcement at retrieval time. A generator declining to quote a forbidden passage does not repair the earlier access violation. Treat any unauthorized evidence entering the model context as a release blocker.
Review Answers with an Explicit Rubric
For each answer, check required facts, unsupported claims, contradictions, and whether each citation supports its associated claim. An existing URL is not enough. Preserve the actual context passed to the generator so reviewers can distinguish a missing passage from an ignored passage. Apply access controls and retention limits to these traces too.
A model judge can help triage answers, but calibrate it against independent human labels. Give reviewers the same rubric and resolve disagreements before treating their labels as references. Ask the judge for a verdict with evidence spans; compare false passes and false failures by scenario. Recheck calibration after changing its model or prompt.
For pairwise comparisons, hide system names and swap answer order to detect position sensitivity. Include concise correct answers and verbose incorrect answers in calibration. The MT-Bench study of LLM judges documents position and verbosity biases. Keep deterministic permission checks outside the judge, and have people review consequential failures.
Turn Failures into a Release Decision
| Observed failure | Inspect next |
|---|---|
| Evidence absent from candidates | Ingestion, labels, query matching, and authorization filters |
| Evidence retrieved but absent from context | Reranking, duplicate removal, and token truncation |
| Evidence present but answer wrong | Instructions, source conflicts, and generation behavior |
| Answer accurate but unauthorized | Identity propagation, permission synchronization, and cache isolation |
Compare one change at a time against the same baseline cases. Record corpus, embedding, retriever, prompt, generator, and judge versions. Inspect per-question differences and scenario groups before the average. Repeat nondeterministic generation runs and report variation; a small aggregate movement on a tiny dataset is not reliable evidence of improvement.
Track end-to-end and stage latency, including p95, alongside input/output tokens and measured cost per request. A reranker or larger context needs to earn its additional work through better answers. Choose release thresholds from your product's failure tolerance and baseline, not a universal quality score. Require zero observed authorization leaks in the suite, while recognizing that passing finite tests cannot prove the absence of every leak.
Keep newly discovered production failures as regression cases and refresh the representative sample as usage changes. Next, work through RAG evaluation and quality engineering to connect this harness to a broader quality program.