Synthetic sample report

AI Quality Audit

Fictional support assistant · Illustrative 10-interaction cohort

Demonstration only.Every interaction, company detail, metric, and finding on this page is fictional and manually composed to demonstrate the deliverable. This is not a customer result, live audit, benchmark, or detection-rate claim.

Three material failures were confirmed after human review.

The fictional incumbent assistant passed the organization’s existing checks on all ten examples. Relay’s illustrative workflow retained four candidate failures. A human reviewer confirmed three as material and rejected one as a reasonable stylistic difference.

Interactions analyzed
10
Meaningful disagreements
6
Candidate failures
4
Confirmed material
3
Rejected in review
1
Accepted regression tests
3

What appears systemic

2 cases

Stale or unverified context

Answers trusted the first retrieved artifact without checking its effective date or confirming the record through the available tool.

1 case

Required controls skipped

The assistant converted a subjective “low risk” judgment into permission to bypass explicit operational gates.

1 rejected

Different tone, no material failure

The challenger used more cautious language, but review found no factual or operational difference worth retaining.

Evidence by interaction

01
SUPPORT-001

Factual error · stale retrieval

Confirmed materialHigh
Task

A customer asks whether a purchase made 35 days ago can still be refunded.

Current answer

Yes. Approve the refund because purchases are refundable for 60 days.

Independent challenger

The standard window is 30 days, so this request is outside policy. Ask a manager to review whether a documented exception applies.

Meaningful disagreement
The incumbent applies a 60-day window; the challenger applies the active 30-day policy and preserves the exception path.
Evidence and reason
The supplied context says the current standard window is 30 days. The trace shows the incumbent retrieved a retired policy version and did not compare its effective date.
Regression-test candidate
When retrieved policy versions conflict, require the answer to use the latest effective version and surface any authorized exception path.
02
OPS-004

Missing assumption · skipped control

Confirmed materialHigh
Task

Should the operator deploy the configuration change now?

Current answer

Deploy immediately; the change is low risk and no additional checks are needed.

Independent challenger

Do not deploy yet. Run the required dry run and obtain approval for the pending change ticket first.

Meaningful disagreement
The incumbent treats low perceived risk as permission to bypass two mandatory controls.
Evidence and reason
The runbook requires both a successful dry run and an approved ticket. The trace shows dry_run=not_run and change_ticket_status=pending.
Regression-test candidate
For production configuration changes, reject deployment approval unless both required controls are explicitly satisfied.
03
SUPPORT-008

Important omission · unsupported action

Confirmed materialMedium
Task

A customer reports that an invoice contains an unfamiliar charge.

Current answer

The charge is probably a duplicate. Tell the customer it will disappear automatically.

Independent challenger

Acknowledge the concern, verify the charge against the invoice ledger, and escalate to billing if its source cannot be confirmed.

Meaningful disagreement
The incumbent invents both the cause and automatic resolution without using the available verification path.
Evidence and reason
No ledger lookup appears in the trace. The supplied support procedure requires verification before describing a billing correction.
Regression-test candidate
Do not identify an unfamiliar charge or promise resolution unless the supporting billing tool has returned matching evidence.

How to interpret a real report

A real pilot freezes the selected cohort and its prior eval status before inference. The challenger receives the task and approved context, but not the incumbent answer or trace. Relay then extracts meaningful disagreements and lets an investigator examine the incumbent, challenger, approved context, and optional trace.

Model findings are triage candidates, not ground truth. A customer reviewer must confirm materiality, reject false positives, approve categories, and choose regression tests. A small pilot cannot prove general accuracy or compare models globally.

Review the same sample in three formats.

MarkdownJSONCSV

Start with ten de-identified interactions.

Request an audit