Independent AI testing

breaklight — the consultancy that tests AI

A consultancy that tests AI — only AI.

Join the waitlist →

02 · The problem

DO YOU WANT TO KNOW IF YOUR AI IS RIGHT

Every engagement produces evidence, not opinions.

01 · Offering

RAG Accuracy & Grounding Assessment

A structured assessment of any RAG system, run on the breaklight methodology. Findings report, remediation roadmap, and a repeatable eval harness you keep.

  • Structured assessment of your existing RAG technology stack, LLMs, and document pipelines.
  • Flexible manual or automated tooling to extract your corpus and build a Golden Dataset for testing.
  • Report and remediation roadmap with vital metrics for next steps and ongoing review.
Join the waitlist Early access

04 · The method

The breaklight methodology.

Every assessment runs the same methodology, so the verdict is comparable, repeatable, and defensible.

  1. 1

    Retrieval

    Does the right context come back? Recall, precision, MRR and NDCG — with latency measured on the same calls.

  2. 2

    Grounding

    Is every claim anchored in your corpus — or in the model's imagination?

  3. 3

    Hallucination

    What does the system invent under pressure, and how often?

  4. 4

    Adversarial

    Prompt injection, access-control leaks, and the OWASP LLM risks that reach production.

  5. 5

    Eval Ops

    A golden dataset, regression gates and CI/CD wiring your team keeps after handover.

Built by experienced testers with over 25 years of experience across a wide range of domains and industries.

05 · FAQ

The questions buyers ask most.

Because every AI deployment is unique, our engagement timelines are flexible. We scope each project dynamically based on your specific systems, LLM selections, dataset requirements, and testing goals.

No matter the timeline, every engagement delivers:

  • A thorough technical review of your current stack and corpus.
  • Customized dataset extraction and evaluation harness implementation.
  • Key performance metrics, a comprehensive findings report, and a strategic roadmap for your next steps.

Skipping it is the most expensive mistake. Teams that ship without pre-launch testing find out which failure modes they have by reading customer complaints. The cost of fixing a problem in production is somewhere between 10x and 100x the cost of fixing it pre-launch. The cost of not fixing it is unbounded.

Neither. UAT for AI systems does not work the way UAT for traditional software works: two testers running the same query get different model answers, and UAT signal collapses. The right shape: breaklight runs first, producing quantitative scores against a defined golden dataset. UAT then runs on top — humans test usability, tone, and edge-case judgment on a system already proven factually correct.

Four separate vendor categories, often confused. AI testing (breaklight) is a methodology applied at a defined point. AI observability (Arize, WhyLabs, Helicone, Datadog) is runtime monitoring of production AI traffic. AI eval platforms (Braintrust, Galileo, LangSmith) are SaaS dashboards for managing test runs over time. AI guardrails (Lakera, NeMo, Patronus) are runtime filters that block bad outputs before they reach users.

breaklight does AI testing only. We do not sell observability dashboards, run runtime guardrails, or operate as a SaaS eval platform. For runtime monitoring, talk to Arize, WhyLabs, Helicone, Langfuse, or Datadog. For eval platforms: Braintrust, Galileo, LangSmith, Vellum, or Humanloop — our harness integrates with any of them. For guardrails: Lakera, NeMo, Guardrails AI, Patronus, or Aporia.

breaklight minimises data handling by design. Test execution runs on client infrastructure by default — the harness is deployed into your environment, and client data never leaves the client perimeter for normal testing. Where breaklight does hold client artefacts, they are encrypted at rest, access-limited to the engagement team, isolated from other work, and access-logged. The harness uses your existing LLM endpoints under your existing vendor agreements.

At handover you take ownership of the golden dataset, the runnable harness, the findings report, and any associated artefacts. breaklight deletes client knowledge base content, customer queries, test outputs containing client-identifiable information, attack-surface details, and credentials — deletion confirmed in writing within 14 days. Only aggregated anonymised metrics are retained.

Latency is measured inside our retrieval testing, on the same calls that produce Recall, Precision, MRR and NDCG — at zero added cost. breaklight instruments the retriever to capture wall-clock time from query submission to ranked chunks returned, reporting p50, p95 and p99 across the full golden dataset. Default thresholds: p95 < 500ms / p99 < 1.2s for conversational deployments; p95 < 2s for batch. Drift across runs is gated — >20% raises a warn, >50% a fail.

Yes — access-control verification sits inside our adversarial testing. The common failure: vector search bypasses row-level security, and a question from a junior user returns a chunk from a document only senior users should see. breaklight seeds the corpus with role-tagged documents, issues identical queries as different user personas, and confirms retrieval respects the documented boundaries. Pass bar is zero leaks. Maps directly to OWASP LLM02.

Yes, with one caveat. Open-source LlamaIndex is fully testable end-to-end because chunking, embeddings and retriever internals are inspectable. LlamaCloud is testable as a black box for outcome measurement, but root-cause diagnostics often need visibility into the chunking and embedding layer, which managed services abstract away. LlamaIndex's built-in evaluation module is a useful in-loop developer tool — not a substitute for an independent third-party assessment.

06 · Insights

What we're learning testing AI.

Join the waitlist

Early access. Limited engagement slots each quarter.