What Is RAG Pipeline Evaluation and Why Does It Matter?
7 October 2026
Your RAG system gives a wrong answer. Where did it go wrong?
If you only evaluate the final answer, you cannot say. The fault could sit in parsing, chunking, the index, the retriever, the reranker, the prompt or the model — and each needs a different fix.
RAG pipeline evaluation measures each stage separately so you can find out.
The RAG pipeline from end to end
A typical pipeline runs like this:
- Ingestion and parsing. Documents are pulled in and turned into text: PDFs, web pages, tables, scanned files.
- Chunking. Text is split into passages, carrying whatever metadata survives — titles, dates, access permissions.
- Embedding and indexing. Passages become vectors, often alongside a keyword index.
- Query processing. The user's question is rewritten, expanded or routed.
- Retrieval. Candidate passages come back, filtered by metadata and permissions.
- Reranking. Candidates are reordered by relevance.
- Context assembly. The top passages are packed into the prompt.
- Generation. The model writes the answer, ideally with citations.
- Post-processing. Guardrails, formatting, citation rendering.
Not every system has every stage. The evaluation should follow the pipeline you actually built.
What happens when one stage fails
Failures cascade, and they disguise themselves. A table the parser flattened never becomes a useful passage. A chunk that separates a rule from its exception retrieves cleanly and answers wrongly. A permission filter that is not applied at the vector layer returns a passage the user should never see — and the answer looks perfectly helpful.
By the time anyone reads the output, all of these look the same: a bad answer. Score only the final sentence and you will misdiagnose retrieval failures as hallucinations.
Why standard software testing falls short for RAG
Conventional tests check that a given input produces an expected output. RAG systems strain that model.
What makes RAG evaluation different
- Outputs vary. The same question can produce differently worded answers across runs, and several may be correct. Exact-match assertions fail good answers and pass bad ones that happen to contain the right words.
- Correctness depends on the corpus. The right answer is whatever your documents say, and your documents change.
- Failures are silent. Nothing throws an error when the model fills a gap with something plausible.
UAT does not close the gap either. Two testers running the same query get different model answers, so the signal is weak. Quantitative scoring against an agreed test set should run first, with human testing of usability, tone and edge-case judgement on top.
Key components of a sound RAG evaluation framework
A workable framework needs a test set of real and deliberately difficult questions, with the relevant passages identified for each; measurements at each stage; traces that keep the retrieved context; and scoring rules that stay fixed between runs.
| Stage | What to measure | Typical failure |
|---|---|---|
| Parsing and chunking | Whether key passages survive intact, with their metadata | Tables flattened; a rule split from its exception |
| Retrieval | Recall, precision, MRR, NDCG, with latency on the same calls | Paraphrased questions miss; stale versions rank first |
| Permission filtering | What each role receives, using a test account per role | A junior user's query returns a chunk only senior users should see |
| Context assembly | What actually reached the model, and in what order | The relevant passage truncated or buried among distractors |
| Generation | Claim-level support, citation accuracy, abstention | A confident answer from memory instead of the corpus |
| Adversarial | Injection and extraction success, by category | Instructions hidden in a document override the system prompt |
How to measure retrieval quality
Retrieval Quality asks whether the right context comes back. Recall: did the relevant passages appear at all? Precision: how much of what came back was relevant? MRR and NDCG: did the best passage rank near the top?
Measure latency on the same calls. Then slice the results by question type, document type and user role, because an average across everything hides the category that always fails.
Assessing groundedness and faithfulness
Once you know what was retrieved, ask whether the answer stayed inside it. Split the answer into claims and check each against the supplied context. Check that citations point to passages that actually contain the claim. Include questions the corpus cannot answer, and score whether the system abstains.
Keep this score apart from retrieval. A well-grounded answer built on the wrong passage is a retrieval failure; an ungrounded answer built on the right passage is a generation failure. AI grounding failures goes deeper on the second kind.
Common failure modes found during evaluation
- Paraphrase sensitivity. Asked in the user's words, the question retrieves nothing useful; asked in the document's words, it works.
- Version confusion. Superseded documents stay in the index and outrank current ones.
- Overreach. The answer stretches a passage beyond what it says.
- Missing abstention. The system answers questions the corpus does not cover.
- Permission leaks. Vector search bypasses row-level security.
Stanford's 2024 study "Hallucination-Free?" found that 17–33% of answers from leading retrieval-based legal research tools contained a hallucination — tools built specifically to answer from a curated corpus.
How adversarial inputs expose hidden RAG weaknesses
An attacker does not have to type into the chat box. A RAG system reads whatever is in its corpus, so a document containing instructions can reach the model through retrieval. That is prompt injection (OWASP LLM01) arriving by the side door.
Adversarial evaluation also probes the vector layer directly, with queries crafted to pull passages across permission boundaries. That maps to OWASP LLM02, sensitive information disclosure, with LLM08 covering vector and embedding weaknesses. Doing it properly needs your written authorisation and a test account for each role. The pass bar is zero leaks.
What a RAG evaluation report should include
A useful report covers:
- Scope. The system, the corpus and the object under test.
- The test set. How it was built and what it covers.
- Scores per area — retrieval, grounding, hallucination, adversarial — never one blended number.
- Results by slice. Question type, document type, user role.
- Findings with evidence. For each: the query, the retrieved passages, the answer, and why it failed.
- Latency. p50, p95 and p99, with gates on mean latency and timeout rate against your SLA.
- Regression readiness. What your setup needs in order to rerun this evaluation whenever the system changes.
- Coverage gaps. What was out of scope, written down rather than implied.
Turning findings into remediation actions
Findings only help if they are ordered. A remediation roadmap ranks fixes by risk and by where they sit in the pipeline. Fix upstream stages first, because a better retriever can shift every score downstream of it. Then rerun the same test set and compare against the baseline.
When to run a RAG pipeline evaluation
- Before launch, and before opening the system to new users or roles.
- After changing the model, the embedding model, the chunking strategy or the reranker.
- After a significant change to the corpus.
- When users report wrong answers and nobody can say which stage caused them.
- When someone outside the team — a customer, a board, a regulator — asks for evidence.
How independent evaluation adds credibility
An evaluation run by the team that built the pipeline is useful and limited: they know which questions it handles and tend to write the test set around them.
An independent evaluation uses a test set agreed in advance, the same scoring rules every time, and evidence you can hand to someone else. It is not certification. It is a record of what was tested, how, and what was found. The full methodology is in the breaklight whitepaper.
breaklight's core assessment evaluates the pipeline area by area — Retrieval Quality, Answer Grounding, Hallucination, Adversarial and Eval Ops, with long-context testing where your system relies on very large contexts. You get performance metrics, a findings report and a remediation roadmap in priority order, so your team knows which stage to fix first.
Talk to us