breaklight ← All blog posts Blog

RAG Accuracy and Grounding Assessments: A Complete Guide

7 October 2026

A retrieval-based system makes two promises. It will find the right material, and it will answer only from it.

Most testing checks neither directly. It reads the final answer and decides whether it sounds right.

This guide sets out what each promise means, how to measure it, and what the numbers tell you.

Understanding RAG: how retrieval-augmented generation works

Retrieval-augmented generation (RAG) puts a search step in front of a language model. The model doesn't answer from memory; it answers from documents fetched for that question.

A typical pipeline:

  1. Ingest. Documents are split into chunks and turned into embeddings, stored in an index.
  2. Retrieve. The user's question is matched against the index — by meaning, by keywords, or both.
  3. Rank and filter. Results are re-ranked and filtered, including by what the user is allowed to see.
  4. Generate. The best chunks are packed into a prompt, and the model writes an answer, ideally with citations.

Variants abound. Classic, hybrid and re-ranked retrieval are the common patterns. Knowledge-graph and GraphRAG systems walk relationships between entities. Agentic systems retrieve in a loop, deciding what to fetch next. Some systems skip retrieval and load the whole knowledge base into a very large context.

Each needs a slightly different test. The two promises stay the same.

What RAG accuracy is and why it matters

RAG accuracy is whether the system's answer is correct for the question, given your documents. It depends on both halves of the pipeline: retrieval has to bring back the right content, and generation has to use it properly.

It matters because users treat a retrieval system as a source of record. They were told it reads your documents. When it is wrong, they act on it.

Common retrieval failures that undermine accuracy

Each of these produces a wrong answer that looks like a model problem. None of them is.

What grounding in AI systems is

Grounding is whether each claim in the answer is supported by the context the model was given. A grounded answer says only what the retrieved material says, and says so when the material is silent.

A grounding assessment breaks the answer into individual claims and checks each against the context: supported, unsupported or contradicted. Citations are checked separately. A citation that points to a real document is not proof that the document supports the sentence it sits beside.

The difference between grounding and factual accuracy

They are related, but not the same, and the difference changes the fix.

An answer can be grounded but wrong. The model faithfully repeated your document, and the document is out of date. That is a corpus problem.

An answer can be correct but ungrounded. The model got it right from its general training, not from your documents. It happens to be true this time, but you can't trace it, and the same behaviour will invent something next time. For a system that claims to answer from your sources, that is still a failure.

You need both measures to know which problem you have.

How RAG accuracy and grounding are assessed

An assessment starts with an agreed test set: real questions, the documents that should answer them, and answers a domain expert would accept. It includes questions the corpus can't answer, and hostile ones.

Retrieval is scored first, on its own. Then the answers are scored for grounding and correctness. Keeping them separate is the whole point — score only the final sentence and you will misdiagnose retrieval failures as hallucinations.

Key metrics used in RAG evaluation

Metric What it tells you
Recall Did the relevant documents come back at all?
Precision How much of what came back was relevant?
MRR (mean reciprocal rank) How high did the first relevant result appear?
NDCG How good was the whole ranking, weighted towards the top?
Grounding (claim support) Is each claim backed by the retrieved context?
Hallucination What did the system invent, and how often?
Refusal behaviour Does it decline when the corpus has no answer?
Permission leaks Did any user get content their role doesn't cover? Our pass bar is zero.
Latency Measured on the same runs; reported as p50, p95 and p99.

Report these per area and per slice of questions. One blended score usually lies to you. Where a system relies on very large contexts, add long-context testing: where in the window accuracy degrades, and how badly.

The role of AI test tooling in grounding assessment

Tooling makes this repeatable. Metric libraries compute retrieval scores. Model judges check claims against context at a scale people can't match. In-framework evaluators are useful while you build.

None of it removes judgement. A model judge is only as good as its calibration against human labels in your domain. A metric library is not a release decision.

Why independent verification of RAG systems is essential

The people who built the pipeline also wrote its tests. They test the questions they designed for and score them with rules they chose. Their blind spots end up in both.

An independent assessment uses a test set built for your corpus by people who didn't design the retrieval, and scoring rules that don't move between runs. That produces a baseline you can trust and compare against after every change.

The stakes are visible in the public record. According to the Stanford HAI AI Index Report 2026, there were 362 AI incidents recorded in 2025, up from 233 the year before.

What happens after an assessment: remediation and next steps

The output of a good assessment is not a score. It is a list of findings, each traced to a stage, with evidence and a priority.

Retrieval fixes usually come first, because they change what every answer is built from: chunking, embeddings, ranking, index freshness, permission filters. Grounding fixes follow: instructions to answer only from context, citation behaviour, refusal when the context is silent.

Then re-run the same test set. If a fix can't be shown on the cases that exposed the problem, you don't know it worked. Keeping that test set running as the system changes is what Eval Ops is for.

How do you know if your RAG system needs an assessment?

breaklight's AI Enabled Retrieval Accuracy assessment covers retrieval quality, answer grounding, hallucination, adversarial testing and Eval Ops, with long-context testing in the core where your system needs it. It ends with performance metrics against an agreed test set, a findings report and a remediation roadmap in priority order — evidence for your release decision, not a certificate. The full method is in the breaklight whitepaper.

Talk to us
← All insights