Usable with a quick edit9 of 10 checks passed
The reasoning is sound, evidence is correctly used, and the recommendation is clear. A PM could add a monitoring threshold in a quick edit without redoing any analysis.
Output
Recommendation: no-go for Monday’s rollout to all 42 agents. Reassess once the failed safety gates have been fixed and retested.
The pilot demonstrates value: 64% of drafts were accepted as-is or with minor edits, exceeding the 60% target. Median first-response time fell 41%, exceeding the 30% target, although the before-and-after comparison does not establish causality. Seven of eight agents want to keep it.
But we failed two explicit launch criteria:
- Wrong or harmful drafts: 1.8% against a 1% ceiling. Nine failures is above the maximum five permitted on the 500-ticket set.
- No refund commitments: failed. Two drafts promised refunds. Agents catching them prevented those messages reaching customers; it does not satisfy a criterion about what drafts may say. Human review was already part of the agreed design.
Six errors used outdated billing prices. That gives us a concrete fix, but billing represents 38% of tickets, so this is a substantial exposure. Updating the reference data alone would not establish that the remaining errors or refund promises are resolved.
Required before reconsidering launch:
- Engineering and the billing policy owner: update and verify current prices and policies, with a named owner and process for keeping reference data current. Investigate the other three harmful drafts.
- Engineering: add and test a control that prevents refund commitments from appearing in agent-facing drafts, including paraphrases.
- PM and evaluation owner: rerun the full 500-ticket evaluation after the fixes, meeting the original thresholds, and add targeted pricing and refund tests. Report those additional results separately.
Excluding billing is a possible narrower release, not an automatic pass: we lack queue-level results and do not know whether refund failures are confined to billing. It would require explicit scope approval and validation of both routing and the remaining queues.
Welsh-language tickets should bypass drafting until language handling is fixed. Disable tables or use plain text pending the formatting fix.
Keep Friday’s training as preparation, with launch clearly marked as pending. The booking should not determine readiness. After the gates pass, expand in stages with monitoring and a named owner able to disable drafting.
Grades and run details
Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
- passUses the supplied evidence correctly69%
- passAddresses the actual decision94%
- passRespects explicit constraints44%
- passIdentifies material uncertainty97%
- passAvoids unsupported claims74%
- passProduces the required deliverable86%
- passChecks every agreed criterion49%
- passMakes a clear call93%
- partialLimits the downside of being wrong54%
- passGets the base of every number right83%
Run
- Run
- #1
- Time to output
- 34 s
- Submitted
- 25 Sept 2026