Eval Ops Explained: Running AI Evaluation at Scale
7 October 2026
You tested the system before launch. Then someone swapped the model.
Nobody re-ran the tests, because the tests lived in a notebook only one engineer knew how to run.
Eval Ops is what stops that being normal.
What Eval Ops is
Eval Ops is the operating routine around AI evaluation: the test set, the scoring rules, the gates, and the wiring that runs them whenever the system changes.
Think of it as doing for evaluation what continuous integration did for unit tests. A test you ran once is evidence about one moment. A test that runs on every change is a control.
The question it answers is plain: after we change the model, the prompt or the index — what broke?
Why traditional QA falls short for AI systems
Traditional QA assumes the same input produces the same output, and that only code changes behaviour. Neither holds for an AI system.
| Traditional software QA | AI system evaluation | |
|---|---|---|
| Same input, same output? | Yes | Not reliably — answers vary between runs |
| Pass condition | Exact match against an expected value | Scored against criteria, often claim by claim |
| What changes behaviour | Code | Code, model version, prompt, corpus, index, settings |
| Typical failure | Crash or wrong value | Fluent, plausible and wrong |
| Who notices first | The test suite | Users, unless you measure |
This is also why UAT is weak on AI. Two testers running the same query get different answers, so a thumbs-up from a tester tells you little. Quantitative testing against an agreed test set has to come first. People then judge tone, usability and edge-case judgement on a system whose accuracy is already measured.
Core components of an Eval Ops framework
Continuous evaluation pipelines
The centre of Eval Ops is a versioned test set: real questions, deliberately nasty ones, and questions the system should refuse. It lives alongside the code and changes through review, like the code does.
Runs are triggered by change — a new model version, an edited prompt, a re-indexed corpus. Every run uses the same scoring rules, and results are stored with the version of everything that produced them. Without that, you can't compare today's score with the last one, and the trend line means nothing.
Metrics that actually measure AI quality
Score each stage separately. A retrieval-based system needs at least:
- Retrieval: recall, precision, MRR and NDCG — did the right content come back, and how high did it rank?
- Grounding: is each claim in the answer supported by the retrieved context?
- Hallucination: what does the system invent, and how often?
- Adversarial: does it hold up against prompt injection and access-control probes? For permission leaks, our pass bar is zero.
- Latency: measured on the same runs. We report p50, p95 and p99, and gate on mean latency and timeout rate against your SLA.
Then refuse to blend them. One blended score usually lies to you. A good gate can fail a critical slice — say, every question about refunds — even when the overall average looks fine.
Human-in-the-loop validation
Automation scales the checking. People decide what correct means.
Domain experts write and own the expected answers. Where you use a model as a judge, calibrate it against human labels on your own content and keep checking where they disagree. And keep a human on the consequential calls: anything that touches money, health, employment or legal position.
Tooling and infrastructure for Eval Ops
No single tool is Eval Ops. The categories do different jobs:
- Eval platforms manage your own test sets and runs.
- Observability tools monitor production traffic over time.
- Guardrails filter bad outputs at runtime.
- In-framework evaluators are handy during development.
- Scanners probe for known attack patterns.
Underneath, you need version control for the test set, traces detailed enough to reproduce a failure, and a path from "that answer was bad" back into a new regression case. If your monitoring can't reproduce a failure, it is analytics, not an evaluation loop.
Eval Ops in RAG systems: a special case
In a retrieval-based system, behaviour can change without anyone touching the code.
Someone uploads a new policy. An old one is deleted. The chunking settings change, or the embedding model is upgraded. Every one of those can change answers, and none will trip a code-based test.
So Eval Ops for RAG has extra demands:
- Runs trigger on corpus and index changes, not only on releases.
- Retrieval is scored separately from answers, or you will misdiagnose retrieval failures as hallucinations.
- Permission checks run per role, because new documents arrive with new access rules.
- Questions the corpus can't answer stay in the set, and get updated as the corpus grows.
If your system loads very large contexts instead of retrieving, track where in the window accuracy starts to degrade, and re-check it when the content changes.
How to build an Eval Ops practice inside your organisation
Start with a baseline assessment
You can't detect a regression without knowing where you started. Measure the system now, stage by stage, against an agreed test set. That baseline becomes the reference every later run is compared against.
Define evaluation criteria before you build
Decide what a correct answer is, which questions must be refused, which slices are critical and what your latency SLA is — before the feature exists. Criteria written afterwards tend to drift towards whatever the system already does.
Integrate evaluation into your CI/CD pipeline
Run a fast subset on every prompt or configuration change, and the full set on model, index or corpus changes. Block a release when a critical slice fails. Record the model, prompt, index and test-set versions with every result.
One caution: tools and their recommended workflows change. Don't copy last year's tutorial into this year's release gate without checking it still applies.
Signs your AI system needs stronger Eval Ops
- Customer complaints are how you find out what broke.
- Nobody can say why the scores moved between the last run and this one.
- The test set isn't versioned, or lives in one person's head.
- You report a single overall score.
- A model upgrade went out because the demo looked fine.
- You can't turn a bad production answer into a test case.
What to expect from an AI evaluation consultancy
An outside assessment should test against an agreed test set, score each area separately, and run the same way every time so results stay comparable. It should tell you where your evaluation routine falls short, not sell you a dashboard.
In breaklight's core assessment, Eval Ops means exactly that: we look at what a repeatable routine requires — regression gates, CI/CD wiring — and report where your setup falls short as findings. We score the gap; you keep your stack. We don't take over your pipeline.
If you want the checks wired into your release process and firing every time the system changes, that is a separate extension, Continuous eval ops, scoped upfront or added later by change request.
breaklight's assessment gives you the baseline Eval Ops needs: performance metrics per area against an agreed test set, findings on regression readiness, and a remediation roadmap in priority order. It produces evidence you can build a routine on, not a certificate. The method is set out in the breaklight whitepaper.
Talk to us