AI Quality Audit

Find the AI failures your current evals miss.

Give Relay 10–20 interactions that passed your current evals. We find likely failures they missed, show the evidence, and turn confirmed cases into regression tests.

Send 10 de-identified examples that already passed your current evals. No integration required.

Quality AuditSynthetic example
Human reviewed
Support assistant · 10 fictional interactions

Four likely failures escaped the existing checks.

10interactions audited
4candidate failures
3confirmed material
3regression tests
01
Stale policy retrieval

The assistant gave a 60-day refund window after retrieving an obsolete policy. The active policy allows 30 days.

High
Customer de-identifies firstApproved providers onlyHuman-confirmed findingsMarkdown · JSON · CSV

A passing eval can still hide a real failure.

Existing checks often ask whether one answer looks acceptable. Relay asks a different question: where does an independently produced answer disagree, and is that disagreement evidence of a material error, missing assumption, unsupported claim, or important omission?

A manual first engagement, not another eval platform.

1

Freeze the cohort

Select 10–20 representative interactions that your existing eval or review process did not flag. Record that status before Relay sees the examples.

2

Investigate disagreements

A challenger answers independently. Relay extracts meaningful differences, then investigates them against the supplied context and optional trace.

3

Review and operationalize

A person confirms material findings, rejects false positives, groups recurring patterns, and chooses which cases become regression tests.

A decision-ready audit, not a leaderboard.

Every retained finding stays tied to the original interaction, challenger response, disagreement, evidence, confidence, and human review outcome.

01

Hidden failures

A prioritized list of likely material failures, with the exact evidence behind each finding.

02

Failure patterns

Recurring categories and system behaviors that are hard to see when examples are reviewed one at a time.

03

Challenger advantages

Cases where an independent answer exposes a better approach—not a generic model leaderboard.

04

Regression tests

Human-approved examples your team can add to the eval suite it already uses.

Inspect a complete synthetic sample.

The public report uses fictional support interactions and manually composed findings. It shows the exact reporting structure without implying a customer result or proven detection rate.

Read the sample report
Top recurring patternsIllustrative
  1. 01
    Stale source selected2 candidate cases · 2 confirmed
  2. 02
    Required escalation skipped1 candidate case · 1 confirmed
  3. 03
    Uncertainty overstated1 candidate case · rejected in review

Minimize first. Process only what the audit needs.

The first pilot is designed for ordinary support, knowledge, sales, research, or internal assistants—not medical, financial, children’s, or other highly regulated data.

You remove identifiers and secrets

Replace names, emails, account IDs, customer names, and internal project names with random labels. Remove tokens, credentials, exact addresses, and unnecessary trace fields entirely.

You approve the processing boundary

Relay names the proposed AI providers before a live run. No customer data is sent to a provider until that boundary is approved.

Your pilot data stays out of the product

No website upload, public post, training of a Relay model, or customer-data analytics. Storage and deletion timing are agreed before delivery.

Teams already running AI in a real workflow.

Customer support assistantsEnterprise knowledge and RAGSales and research agentsInternal workflow copilots

Have ten examples your current evals passed?

Send a short description of your AI workflow. We will reply with the de-identification checklist, proposed provider boundary, and a sample file—before you share any interaction data.

Run a free 10-example audit