What we test
Core
Does the right context come back? Recall, precision, MRR and NDCG — with latency measured on the same calls.
Is every claim anchored in your corpus — or in the model's imagination?
What does the system invent under pressure, and how often?
Prompt injection, access-control leaks, and the OWASP LLM risks that reach production.
What a repeatable evaluation routine requires — regression gates, CI/CD wiring, and where your setup falls short. We score the gap; you keep your stack.
Extensions · scoped upfront or added by change request
Where, and how badly, accuracy degrades across the window — on the multi-fact, distractor-heavy tasks that break current systems.
The destination and the safety of the route — every tool the agent can reach inventoried and tested, not trusted.
Fairness across groups of users — and whether it still holds after the model or corpus changes.
The same checks wired into your release process, firing every time the system changes.
Tools we test with and alongside
Python, promptfoo, Langfuse and Ragas.
The core is Domains 1–5 — Retrieval Quality, Answer Grounding, Hallucination, Adversarial, and Eval Ops. How deep each one runs is set by time and risk when we scope the engagement.
Four extensions sit alongside them, each priced separately — scope them upfront, or add them later by Change Request:
The same five questions run the same way every time, so results stay comparable between engagements and over time.
All of them. AI Enabled Retrieval Accuracy is pattern-agnostic: classic retrieval patterns, cache-augmented generation, knowledge-graph and GraphRAG retrieval, hybrid and re-ranked pipelines, and agentic retrieval loops are each one implementation pattern underneath the same five domains.
The methodology measures whether the right context comes back and whether the answer is anchored in it — not which framework you chose to get there.
Because every deployment is different, we scope a fixed-fee Offer 01 assessment to your system, corpus, and risk — not an open-ended day-rate burn.
No matter the length, every engagement delivers:
Skipping it is the most expensive mistake. Teams that ship without pre-launch testing find out which failure modes they have by reading customer complaints. The cost of fixing a problem in production is somewhere between 10x and 100x the cost of fixing it pre-launch. The cost of not fixing it is unbounded.
Neither. UAT for AI systems does not work the way UAT for traditional software works: two testers running the same query get different model answers, and UAT signal collapses. The right shape: breaklight runs first, producing quantitative scores against an agreed test set. UAT then runs on top — humans test usability, tone, and edge-case judgment on a system already proven factually correct.
Four separate vendor categories, often confused. AI testing (breaklight) is a methodology applied at a defined point. AI observability (Arize, WhyLabs, Helicone, Datadog) is runtime monitoring of production AI traffic. AI eval platforms (Braintrust, Galileo, LangSmith) are SaaS dashboards for managing test runs over time. AI guardrails (Lakera, NeMo, Patronus) are runtime filters that block bad outputs before they reach users.
breaklight does AI testing only. We do not sell observability dashboards, run runtime guardrails, or operate as a SaaS eval platform. For runtime monitoring, talk to Arize, WhyLabs, Helicone, Langfuse, or Datadog. For eval platforms: Braintrust, Galileo, LangSmith, Vellum, or Humanloop — we can work alongside any of them. For guardrails: Lakera, NeMo, Guardrails AI, Patronus, or Aporia.
breaklight minimises data handling by design. Test execution runs on your infrastructure by default, and client data never leaves your perimeter for normal testing. Where breaklight does hold client artefacts, they are encrypted at rest, access-limited to the engagement team, isolated from other work, and access-logged. Testing uses your existing LLM endpoints under your existing vendor agreements.
You receive the findings report, the performance metrics, and the remediation roadmap. breaklight deletes client knowledge base content, customer queries, test outputs containing client-identifiable information, attack-surface details, and credentials — deletion confirmed in writing within 14 days. Only aggregated anonymised metrics are retained.
Latency is measured inside our retrieval testing, on the same calls that produce Recall, Precision, MRR and NDCG. breaklight instruments the retriever to capture wall-clock time from query submission to ranked chunks returned, reporting p50, p95 and p99 across the full test set. Default thresholds: p95 < 500ms / p99 < 1.2s for conversational deployments; p95 < 2s for batch. Drift across runs is gated — >20% raises a warn, >50% a fail.
Yes — access-control verification sits inside our adversarial testing. The common failure: vector search bypasses row-level security, and a question from a junior user returns a chunk from a document only senior users should see. breaklight seeds the corpus with role-tagged documents, issues identical queries as different user personas, and confirms retrieval respects the documented boundaries. Pass bar is zero leaks. Maps directly to OWASP LLM02.
Yes, with one caveat. Open-source LlamaIndex is fully testable end-to-end because chunking, embeddings and retriever internals are inspectable. LlamaCloud is testable as a black box for outcome measurement, but root-cause diagnostics often need visibility into the chunking and embedding layer, which managed services abstract away. LlamaIndex's built-in evaluation module is a useful in-loop developer tool — not a substitute for an independent third-party assessment.
Tell us what you've built and what you need proven. A person reads every message and replies — no bot, no drip sequence.
Talk to us →