Four likely failures escaped the existing checks.
The assistant gave a 60-day refund window after retrieving an obsolete policy. The active policy allows 30 days.
AI Quality Audit
Give Relay 10–20 interactions that passed your current evals. We find likely failures they missed, show the evidence, and turn confirmed cases into regression tests.
Send 10 de-identified examples that already passed your current evals. No integration required.
The assistant gave a 60-day refund window after retrieving an obsolete policy. The active policy allows 30 days.
The blind spot
Existing checks often ask whether one answer looks acceptable. Relay asks a different question: where does an independently produced answer disagree, and is that disagreement evidence of a material error, missing assumption, unsupported claim, or important omission?
How it works
Select 10–20 representative interactions that your existing eval or review process did not flag. Record that status before Relay sees the examples.
A challenger answers independently. Relay extracts meaningful differences, then investigates them against the supplied context and optional trace.
A person confirms material findings, rejects false positives, groups recurring patterns, and chooses which cases become regression tests.
What you receive
Every retained finding stays tied to the original interaction, challenger response, disagreement, evidence, confidence, and human review outcome.
A prioritized list of likely material failures, with the exact evidence behind each finding.
Recurring categories and system behaviors that are hard to see when examples are reviewed one at a time.
Cases where an independent answer exposes a better approach—not a generic model leaderboard.
Human-approved examples your team can add to the eval suite it already uses.
See the deliverable
The public report uses fictional support interactions and manually composed findings. It shows the exact reporting structure without implying a customer result or proven detection rate.
Read the sample reportData handling
The first pilot is designed for ordinary support, knowledge, sales, research, or internal assistants—not medical, financial, children’s, or other highly regulated data.
Replace names, emails, account IDs, customer names, and internal project names with random labels. Remove tokens, credentials, exact addresses, and unnecessary trace fields entirely.
Relay names the proposed AI providers before a live run. No customer data is sent to a provider until that boundary is approved.
No website upload, public post, training of a Relay model, or customer-data analytics. Storage and deletion timing are agreed before delivery.
Good first fit
First design partners
Send a short description of your AI workflow. We will reply with the de-identification checklist, proposed provider boundary, and a sample file—before you share any interaction data.
Run a free 10-example audit