breaklight ← All blog posts Blog

From Prototype to Production: AI Application Verification Explained

7 October 2026

The prototype worked in the demo. Everyone asked it questions they knew it could answer.

Production is different: more documents, more users, more roles, and people who ask things nobody predicted — some of them on purpose.

Verification is how you find out what changes between the two before your users do.

Why verification is a critical step

A prototype proves something is possible. A release decision needs evidence that it holds up under the conditions it will actually meet.

The gap is wider for AI than for conventional software, because the failures are quiet. A wrong answer looks exactly like a right one. A permission leak looks like a helpful reply. Nothing crashes.

Trust is not a given either. KPMG and the University of Melbourne found in 2025 that 46% of people worldwide are willing to trust AI systems, across 47 countries. A confident wrong answer spends trust you were already short of.

What AI application verification actually involves

Functional verification vs behavioural verification

Functional verification asks whether the application works as software. The endpoint responds, sign-in works, the interface renders, data flows through the integrations. Conventional testing handles this well — and the breaklight platform also runs API, browser, security and performance test types if you need them alongside the assessment.

Behavioural verification asks whether the AI does the right thing. Does it find the right content, answer only from it, avoid inventing, and hold up under attack? That is where conventional tests run out. A UI test can tell you a button works and still tell you nothing about whether the answer is true.

Functional verification Behavioural verification
Question Does the software work? Does the AI behave correctly?
Reference point Expected output for a given input An agreed test set, your corpus and fixed scoring rules
Typical failure Error, timeout, broken flow Plausible wrong answer, data leak, quiet overreach
Repeatability Same input, same output Answers vary between runs, so scoring needs repeated runs

What gets tested during an AI application assessment

The breaklight core assessment answers five questions, in order:

  1. Did it find the right content? Retrieval Quality: recall, precision, MRR and NDCG, with latency measured on the same calls.
  2. Did it answer honestly from it? Answer Grounding: is every claim anchored in your corpus?
  3. Did it avoid inventing anything? Hallucination: what the system makes up under pressure, and how often.
  4. Does it hold up under attack? Adversarial: prompt injection, access-control leaks and the OWASP LLM risks that reach production.
  5. Can you keep proving it? Eval Ops: how ready your setup is for regression testing, reported as findings.

Long-context testing is part of the core when your system relies on very large contexts. Agents that plan and act are covered by the Agentic assessment extension.

Common failure modes found during verification

How retrieval failures undermine application trust

Prototypes usually run on a small, clean corpus. Production corpora are bigger and messier: duplicates, superseded versions, scanned documents, tables. Retrieval that worked on a curated folder starts returning near-misses and outdated versions.

Permissions are the sharpest version of this. A prototype often runs under a single test account, sometimes one with full access. In production, vector search can bypass row-level security, so a junior user's question returns a chunk only senior users should see. We test this with your written authorisation and a test account for each role, issuing the same queries under each role. The pass bar is zero leaks. It maps to OWASP LLM02, with LLM08 for the vector layer.

Grounding issues that slip through internal testing

Internal testers know the corpus. They ask questions that have answers, phrased the way the documents phrase them.

That hides the two grounding failures that matter most in production: answering from general knowledge when your corpus says something different, and answering when your corpus says nothing at all.

The verification process, step by step

  1. Scope. Scoped to your system, your corpus and your risk before anything runs.
  2. Connect by your rules. We test from our own secured environment and reach yours by the route your security policy allows — or install the platform inside your network.
  3. Score against an agreed test set. The same areas, run the same way every time, so results stay comparable.
  4. Findings and a remediation roadmap. Performance metrics, a findings report and the fixes in priority order.
  5. Verified deletion. When the retention period in your MSA and its data-processing annex ends, or sooner on request, a secure delete — verified, recorded and attested in writing.

Defining scope and success criteria before testing begins

Write down what you are verifying and what passing means before anything runs: which system, which corpus, which user roles, which risks. How deep each area goes is set by that risk.

Then agree the thresholds. Some should not bend: permission leaks should be zero. Others depend on your use case and are yours to set. For latency, judge mean latency and timeout rate against your production SLA, and look at p50, p95 and p99 so the tail is visible. If you have no SLA, agree targets before testing starts.

Agreeing this upfront stops the bar drifting to wherever the results happen to land.

How evidence-based assessments drive actionable outcomes

A finding without evidence is an opinion. Each finding should come with the query, what was retrieved, what the system said, and why that counts as a failure. That lets your engineers reproduce it and fix the right stage.

The remediation roadmap puts the fixes in priority order. You fix, rerun the same test set, and see what moved.

Independent verification vs internal QA

Internal QA is essential, and it is not a substitute. Your team built the system, wrote the prompts and chose the test questions. Their blind spots are the system's blind spots. Independent verification of AI-built applications covers why outside eyes catch different things.

UAT is not a substitute either. Two testers running the same query get different model answers, so the signal is weak. The right order is quantitative scoring first, then UAT on top for usability, tone and edge-case judgement.

When should you commission an AI application verification?

How breaklight approaches AI application verification

breaklight tests deployed AI systems: customer-facing assistants, internal knowledge search, document Q&A, and agents that retrieve and act. We test AI, and only AI.

A light-touch smoke assessment gives directional results. A full assessment gives evidence you can base a release decision on. Every assessment delivers performance metrics, a findings report and a remediation roadmap. On Eval Ops, we score the gap; you keep your stack. Wiring the checks into your release process is the Continuous eval ops extension.

We do not audit training data, test reinforcement learning or reward hacking, test self-learning systems that change between runs, or assess hosting and network security — use a penetration-testing provider for that. Multimodal testing is currently deferred. We produce evidence, not formal certification; for certification, speak to your auditor. Anything out of scope is recorded in your report's Coverage Gaps section.


If you have a prototype heading for production, the useful question is what it does with your corpus, your users and your roles. A breaklight assessment answers that with measured scores, evidence and a prioritised list of fixes, so the release decision rests on more than a good demo. The methodology is in the breaklight whitepaper.

Talk to us
← All insights