RAG Evaluation: Measure Retrieval and Answer Quality

Build a small Python RAG evaluation harness, separate retrieval failures from answer errors, and test missing evidence and document permissions.

RAG Evaluation: Measure Retrieval and Answer Quality illustration
On this page7 sections

A RAG answer can sound convincing while using the wrong document, omitting an exception, or exposing information the caller cannot access. RAG evaluation needs to tell you which stage failed. A single score for the final answer cannot explain whether to fix ingestion, retrieval, permissions, or generation.

This guide builds a small Python evaluation harness, explains its output, and turns the results into a regression workflow. It assumes you already have a retrieval pipeline, such as the one in building a local AI application with RAG. The example uses fixed retrieval results so you can verify the evaluator without a model, API key, or database.

Measure the Boundaries Separately

Suppose a support assistant retrieves a refund policy and answers that every purchase is refundable. The policy also contains an exception for activated subscriptions. Finding the right document is a retrieval success; omitting the exception is an answer failure. If the exception was lost during chunking, the generator may never have received the evidence.

Faithfulness, also called groundedness, asks whether the answer's claims are supported by the supplied context. Correctness asks whether the answer satisfies the question against an independently reviewed reference. A faithful answer can still be wrong when its source is obsolete. A correct answer recalled from model training can still lack support in the retrieved evidence. Microsoft's RAG evaluator documentation similarly separates retrieval, groundedness, and response completeness.

Build Cases You Can Diagnose

Start with representative questions from approved, redacted support examples and domain reviewers. Include short keyword queries, paraphrases, multi-part questions, rare product names, outdated policies, and ambiguous requests. Synthetic questions can expand coverage, but review them: a question copied from a heading may make retrieval unrealistically easy.

Store a question ID, question text, caller or tenant context, corpus snapshot, expected behavior, required facts, and relevant evidence IDs. Keep source version and evidence spans alongside chunk IDs. Chunk IDs change when you change the splitter; stable source identifiers and spans let reviewers remap labels instead of silently comparing different units. The vector database and embeddings guide provides the retrieval background.

Label actual supporting passages, not every chunk from a relevant document. Record alternative valid evidence sets when several passages independently answer the question. Distinguish an unreviewed candidate from a reviewed irrelevant candidate: incomplete labels can make a valid result look like a retrieval failure.

Split cases before tuning. Use development cases to adjust chunking, prompts, and thresholds; reserve a held-out set for release comparison. Keep near-duplicate questions and paraphrases in the same split. Where appropriate, group by document family or customer scenario. Pin the corpus snapshot so a score change does not secretly reflect a different source collection. This follows the separation of development and test collections described in Stanford's information retrieval evaluation chapter.

Define Recall and Rank Before Writing Code

For this harness, Recall@k is the number of distinct labeled relevant chunks found in the first k results divided by the number of labeled relevant chunks. Duplicate results do not earn extra credit. The denominator means this measures recall against your labels, not against unknowable perfect coverage of the whole corpus.

Reciprocal rank at k is 1 divided by the position of the first relevant result within k, or zero when none appears. MRR@k is its mean across evaluated questions, consistent with Elastic's ranking evaluation definitions. A first-position hit gets full reciprocal-rank credit even when a second essential passage is missing. Use recall and a required-evidence check for questions that need several facts.

Neither metric measures whether the generator used the evidence correctly. Increasing k may improve measured recall while adding distracting text, tokens, and latency. Measure the final context after reranking and truncation as well as the retriever's original candidates.

Run a Reproducible Retrieval Evaluation

Save the following as rag_eval.py and run python rag_eval.py with Python 3. It uses only the standard library. These four hand-written cases deliberately contain failures; their scores are fixture output, not a benchmark for any retrieval product.

from statistics import mean


def retrieval_metrics(relevant, retrieved, k):
    if k < 1:
        raise ValueError("k must be positive")
    top = retrieved[:k]
    if not relevant:
        return None  # No relevant evidence: evaluate abstention separately.
    recall = len(set(top) & relevant) / len(relevant)
    rr = next((1 / rank for rank, item in enumerate(top, 1)
               if item in relevant), 0.0)
    return recall, rr


# IDs represent versioned chunks in one frozen corpus.
cases = [
    {"id": "refund", "relevant": {"refund-v2", "exception-v2"},
     "retrieved": ["refund-v2", "refund-v2", "shipping-v1"],
     "forbidden": set()},
    {"id": "restore", "relevant": {"restore-v3"},
     "retrieved": ["restore-v1", "restore-v3", "shipping-v1"],
     "forbidden": set()},
    {"id": "sla", "relevant": {"sla-v2"},
     "retrieved": ["shipping-v1", "restore-v1"],
     "forbidden": set()},
    {"id": "private-payroll", "relevant": set(),
     "retrieved": ["payroll-private"],
     "forbidden": {"payroll-private"}},
]

k = 3
scores = []
duplicates = leaks = unanswerable = 0
for case in cases:
    top = case["retrieved"][:k]
    duplicates += len(top) - len(set(top))
    leaks += bool(set(top) & case["forbidden"])
    result = retrieval_metrics(case["relevant"], top, k)
    if result is None:
        unanswerable += 1
        continue
    scores.append(result)
    print(f'{case["id"]}: recall={result[0]:.3f} rr={result[1]:.3f}')

print(f"answerable cases: {len(scores)}")
if scores:
    print(f"mean recall@{k}: {mean(s[0] for s in scores):.3f}")
    print(f"MRR@{k}: {mean(s[1] for s in scores):.3f}")
print(f"duplicate slots: {duplicates}")
print(f"cases with forbidden evidence: {leaks}")
print(f"unanswerable cases requiring answer review: {unanswerable}")

# Sanity checks for the evaluator, not acceptance gates for a RAG system.
assert retrieval_metrics({"a"}, [], 3) == (0.0, 0.0)
assert retrieval_metrics({"a", "b"}, ["a", "a", "x"], 3) == (0.5, 1.0)
assert retrieval_metrics(set(), [], 3) is None
assert (duplicates, leaks, unanswerable) == (1, 1, 1)

The output is:

refund: recall=0.500 rr=1.000
restore: recall=1.000 rr=0.500
sla: recall=0.000 rr=0.000
answerable cases: 3
mean recall@3: 0.500
MRR@3: 0.500
duplicate slots: 1
cases with forbidden evidence: 1
unanswerable cases requiring answer review: 1

The refund case exposes why MRR alone is insufficient: the first hit is relevant, yet half the labeled evidence is missing. Repeated chunks occupy real result slots; the harness keeps those positions rather than removing duplicates before ranking. The payroll case is excluded from retrieval averages because its relevant set is empty, but its permission failure remains visible.

To use this with your pipeline, replace each fixed retrieved list with IDs returned under the recorded caller identity. Keep question text and labels in a versioned dataset. Capture intermediate candidates and the final model context separately. The small forbidden-ID sets here are test assertions; production authorization must check every returned document against the actual permission policy.

Test Unanswerable and Unauthorized Requests

A question can have no answer in the authorized corpus even when similar documents exist. Include nonexistent features, missing dates, misleading premises, and requests requiring private evidence. Require an appropriate refusal, clarification, or statement of insufficient evidence. Also measure unnecessary abstention on answerable questions; an assistant that refuses everything should not pass.

Run the same question as an allowed user, a different tenant, and a user whose access was revoked. Check retrieval, assembled context, citations, and cached responses. Microsoft's document access control guidance places permission enforcement at retrieval time. A generator declining to quote a forbidden passage does not repair the earlier access violation. Treat any unauthorized evidence entering the model context as a release blocker.

Review Answers with an Explicit Rubric

For each answer, check required facts, unsupported claims, contradictions, and whether each citation supports its associated claim. An existing URL is not enough. Preserve the actual context passed to the generator so reviewers can distinguish a missing passage from an ignored passage. Apply access controls and retention limits to these traces too.

A model judge can help triage answers, but calibrate it against independent human labels. Give reviewers the same rubric and resolve disagreements before treating their labels as references. Ask the judge for a verdict with evidence spans; compare false passes and false failures by scenario. Recheck calibration after changing its model or prompt.

For pairwise comparisons, hide system names and swap answer order to detect position sensitivity. Include concise correct answers and verbose incorrect answers in calibration. The MT-Bench study of LLM judges documents position and verbosity biases. Keep deterministic permission checks outside the judge, and have people review consequential failures.

Turn Failures into a Release Decision

Use the first failed boundary to choose the next investigation
Observed failureInspect next
Evidence absent from candidatesIngestion, labels, query matching, and authorization filters
Evidence retrieved but absent from contextReranking, duplicate removal, and token truncation
Evidence present but answer wrongInstructions, source conflicts, and generation behavior
Answer accurate but unauthorizedIdentity propagation, permission synchronization, and cache isolation

Compare one change at a time against the same baseline cases. Record corpus, embedding, retriever, prompt, generator, and judge versions. Inspect per-question differences and scenario groups before the average. Repeat nondeterministic generation runs and report variation; a small aggregate movement on a tiny dataset is not reliable evidence of improvement.

Track end-to-end and stage latency, including p95, alongside input/output tokens and measured cost per request. A reranker or larger context needs to earn its additional work through better answers. Choose release thresholds from your product's failure tolerance and baseline, not a universal quality score. Require zero observed authorization leaks in the suite, while recognizing that passing finite tests cannot prove the absence of every leak.

Keep newly discovered production failures as regression cases and refresh the representative sample as usage changes. Next, work through RAG evaluation and quality engineering to connect this harness to a broader quality program.

Share this article

Stuck on implementation?

Get private, 1-on-1 help with system design, performance, scaling, or any technical challenge.

Book a Session

Related Production Resources

Course

Free learning tracks

Turn this guide into a structured production engineering path.

Lab

Interactive engineering labs

Practice the same ideas through scenario-based simulators.

Reference

Production cheatsheets

Keep the operational commands and checks nearby.

Glossary

Key terms

Review the vocabulary behind the architecture.

Discussion

Questions, corrections, or production notes? Add them here so other learners can benefit.

Comments load on demand

To keep this article fast and private by default, the GitHub-powered discussion loads only when you reach this section.

Prefer GitHub? Open the project discussions directly .

Continue Reading

Related practical guides from the same production engineering path.

AI 16 min read

MCP Security in Production: How to Safely Run AI Agents with Tools, OAuth, and Gateways

Learn how to secure MCP-based AI agents with OAuth, token audience validation, gateway policy, tool permissions, SSRF protection, sandboxing, and audit logs.

MCP AI Agents
AI 13 min read

Vector Databases Explained: Embeddings, Similarity Search, and When You Need One

Vector databases power semantic search, recommendation engines, and RAG pipelines. Learn how embeddings work, the HNSW algorithm behind similarity search, chunking strategies, and when pgvector is enough vs when you need Pinecone.

Vector Database Embeddings
AI 12 min read

Fine-Tuning vs RAG vs Prompt Engineering: Which AI Strategy Do You Need?

Your AI project needs domain-specific knowledge. Should you fine-tune a model, build a RAG pipeline, or engineer better prompts? This decision matrix covers cost, accuracy, latency, maintenance, and when each approach wins.

AI RAG
AI 14 min read

Building AI Agents with Claude: From Chatbot to Autonomous Worker

Transform a simple chatbot into an autonomous agent that uses tools, maintains memory, recovers from errors, and orchestrates multi-step workflows. Practical Python guide using Claude API with production-ready patterns.

AI Claude