breaklight ← All blog posts Blog

Eval Ops Explained: Running AI Evaluation at Scale

7 October 2026

You tested the system before launch. Then someone swapped the model.

Nobody re-ran the tests, because the tests lived in a notebook only one engineer knew how to run.

Eval Ops is what stops that being normal.

What Eval Ops is

Eval Ops is the operating routine around AI evaluation: the test set, the scoring rules, the gates, and the wiring that runs them whenever the system changes.

Think of it as doing for evaluation what continuous integration did for unit tests. A test you ran once is evidence about one moment. A test that runs on every change is a control.

The question it answers is plain: after we change the model, the prompt or the index — what broke?

Why traditional QA falls short for AI systems

Traditional QA assumes the same input produces the same output, and that only code changes behaviour. Neither holds for an AI system.

Traditional software QA AI system evaluation
Same input, same output? Yes Not reliably — answers vary between runs
Pass condition Exact match against an expected value Scored against criteria, often claim by claim
What changes behaviour Code Code, model version, prompt, corpus, index, settings
Typical failure Crash or wrong value Fluent, plausible and wrong
Who notices first The test suite Users, unless you measure

This is also why UAT is weak on AI. Two testers running the same query get different answers, so a thumbs-up from a tester tells you little. Quantitative testing against an agreed test set has to come first. People then judge tone, usability and edge-case judgement on a system whose accuracy is already measured.

Core components of an Eval Ops framework

Continuous evaluation pipelines

The centre of Eval Ops is a versioned test set: real questions, deliberately nasty ones, and questions the system should refuse. It lives alongside the code and changes through review, like the code does.

Runs are triggered by change — a new model version, an edited prompt, a re-indexed corpus. Every run uses the same scoring rules, and results are stored with the version of everything that produced them. Without that, you can't compare today's score with the last one, and the trend line means nothing.

Metrics that actually measure AI quality

Score each stage separately. A retrieval-based system needs at least:

Then refuse to blend them. One blended score usually lies to you. A good gate can fail a critical slice — say, every question about refunds — even when the overall average looks fine.

Human-in-the-loop validation

Automation scales the checking. People decide what correct means.

Domain experts write and own the expected answers. Where you use a model as a judge, calibrate it against human labels on your own content and keep checking where they disagree. And keep a human on the consequential calls: anything that touches money, health, employment or legal position.

Tooling and infrastructure for Eval Ops

No single tool is Eval Ops. The categories do different jobs:

Underneath, you need version control for the test set, traces detailed enough to reproduce a failure, and a path from "that answer was bad" back into a new regression case. If your monitoring can't reproduce a failure, it is analytics, not an evaluation loop.

Eval Ops in RAG systems: a special case

In a retrieval-based system, behaviour can change without anyone touching the code.

Someone uploads a new policy. An old one is deleted. The chunking settings change, or the embedding model is upgraded. Every one of those can change answers, and none will trip a code-based test.

So Eval Ops for RAG has extra demands:

If your system loads very large contexts instead of retrieving, track where in the window accuracy starts to degrade, and re-check it when the content changes.

How to build an Eval Ops practice inside your organisation

Start with a baseline assessment

You can't detect a regression without knowing where you started. Measure the system now, stage by stage, against an agreed test set. That baseline becomes the reference every later run is compared against.

Define evaluation criteria before you build

Decide what a correct answer is, which questions must be refused, which slices are critical and what your latency SLA is — before the feature exists. Criteria written afterwards tend to drift towards whatever the system already does.

Integrate evaluation into your CI/CD pipeline

Run a fast subset on every prompt or configuration change, and the full set on model, index or corpus changes. Block a release when a critical slice fails. Record the model, prompt, index and test-set versions with every result.

One caution: tools and their recommended workflows change. Don't copy last year's tutorial into this year's release gate without checking it still applies.

Signs your AI system needs stronger Eval Ops

What to expect from an AI evaluation consultancy

An outside assessment should test against an agreed test set, score each area separately, and run the same way every time so results stay comparable. It should tell you where your evaluation routine falls short, not sell you a dashboard.

In breaklight's core assessment, Eval Ops means exactly that: we look at what a repeatable routine requires — regression gates, CI/CD wiring — and report where your setup falls short as findings. We score the gap; you keep your stack. We don't take over your pipeline.

If you want the checks wired into your release process and firing every time the system changes, that is a separate extension, Continuous eval ops, scoped upfront or added later by change request.

breaklight's assessment gives you the baseline Eval Ops needs: performance metrics per area against an agreed test set, findings on regression readiness, and a remediation roadmap in priority order. It produces evidence you can build a routine on, not a certificate. The method is set out in the breaklight whitepaper.

Talk to us
← All insights