Stale or unverified context
Answers trusted the first retrieved artifact without checking its effective date or confirming the record through the available tool.
Synthetic sample report
Fictional support assistant · Illustrative 10-interaction cohort
Executive summary
The fictional incumbent assistant passed the organization’s existing checks on all ten examples. Relay’s illustrative workflow retained four candidate failures. A human reviewer confirmed three as material and rejected one as a reasonable stylistic difference.
Recurring patterns
Answers trusted the first retrieved artifact without checking its effective date or confirming the record through the available tool.
The assistant converted a subjective “low risk” judgment into permission to bypass explicit operational gates.
The challenger used more cautious language, but review found no factual or operational difference worth retaining.
Confirmed findings
A customer asks whether a purchase made 35 days ago can still be refunded.
Yes. Approve the refund because purchases are refundable for 60 days.
The standard window is 30 days, so this request is outside policy. Ask a manager to review whether a documented exception applies.
Should the operator deploy the configuration change now?
Deploy immediately; the change is low risk and no additional checks are needed.
Do not deploy yet. Run the required dry run and obtain approval for the pending change ticket first.
A customer reports that an invoice contains an unfamiliar charge.
The charge is probably a duplicate. Tell the customer it will disappear automatically.
Acknowledge the concern, verify the charge against the invoice ledger, and escalate to billing if its source cannot be confirmed.
Method and limits
A real pilot freezes the selected cohort and its prior eval status before inference. The challenger receives the task and approved context, but not the incumbent answer or trace. Relay then extracts meaningful disagreements and lets an investigator examine the incumbent, challenger, approved context, and optional trace.
Model findings are triage candidates, not ground truth. A customer reviewer must confirm materiality, reject false positives, approve categories, and choose regression tests. A small pilot cannot prove general accuracy or compare models globally.
Portable deliverables
Run a real pilot