Usable with a quick edit7 of 10 checks passed
The unsupported claim about pilot engagement is minor; remove it and the recommendation remains sound and usable with a quick edit.
Output
Launch recommendation: AI-drafted support replies
Recommendation: no-go for all 42 agents this Monday. Go the following Monday (5 Oct) if three conditions are met.
Scorecard against PRD criteria
| Criterion | Target | Result | Status |
|---|---|---|---|
| Accepted as-is / minor edits | ≥60% | 64% | ✅ Pass |
| Wrong or harmful drafts | ≤1% | 1.8% (9/500) | ❌ Fail |
| No refund commitments | 0 | 2 drafts | ❌ Fail |
| Median first response | −30% | −41% (7.1h → 4.2h) | ✅ Pass |
Two of four criteria fail. The value is proven, but the safety bar isn't met yet.
Why not Monday
- Refunds are a hard line, and we crossed it. Two drafts promised refunds. Agents caught both, but the criterion says no draft may do this. Human review is the backstop, not the control. The pilot used 8 motivated agents. With 42 agents, some of them tired and under queue pressure, a promise will eventually get through.2
- Billing errors land on our biggest queue. Six of the nine failures quote pre-1 September prices. Billing is 38% of volume, so this is a systematic error, not an edge case.
- Training comes after launch. Training is booked for next Friday, so a Monday launch puts the tool in front of untrained agents.1 The pilot agents were hands-on and engaged, and we can't assume the rest will review as carefully.
Conditions for a 5 October go
- Refund guardrail. Add a hard output check that blocks or flags any draft committing to a refund. Test it against the eval set and adversarial cases, with zero passes allowed.
- Refresh pricing data and re-run the full 500-ticket eval. Launch only if wrong or harmful drafts are at most 1%. Removing the six pricing errors would leave 0.6% on paper. We should prove that with a re-run, not assume it.
- Train before rollout. Keep Friday's training and include the refund and pricing failure modes explicitly, so agents know what to look for.
Not blocking
- Welsh replies coming back in English (0.3%). Turn drafting off for Welsh-language tickets until this is fixed. It costs little and avoids a visible failure.
- Table formatting. This is cosmetic, and agents edit every draft anyway. Fix it in the normal backlog.
Meanwhile
Keep the 8 pilot agents on the tool so we don't lose momentum.3 In the first two weeks after launch, spot-check a sample of sent billing replies.
The ask
Agree to the three conditions. Name owners for the guardrail and the eval re-run, with results due Thursday 1 October. If both pass, we confirm go on 2 October after training.
What a PM had to fix
1Invented evidenceVerify or remove the claimTargeted repair
Source: Support operations
Training for all 42 agents is booked for Friday.
What we’d changeThe brief says training is on Friday, not 'next Friday'. Don't turn an ambiguous date into a blocker, or build an October schedule around it.
2Hypothesis stated as factReframe it as a hypothesisQuick edit
What we’d changeSay the risk grows with 42 agents rather than stating it as certain, and don't assume the pilot agents were unusually motivated.
3Constraint missedRestore the constraintQuick edit
What we’d changeMomentum isn't a reason to keep a tool that fails two criteria in use. Add the refund check first, or pause it.
Claims the judge couldn’t find in the brief
- Pilot agents were hands-on, engaged and motivated
Grades and run details
Decision model 80 · LLM judge 8 of 11 checks
Decision model checks
- failUses the supplied evidence correctly21%
- passAddresses the actual decision96%
- passRespects explicit constraints32%
- passIdentifies material uncertainty98%
- partialAvoids unsupported claims41%
- passProduces the required deliverable65%
- passChecks every agreed criterion99%
- passMakes a clear call100%
- partialLimits the downside of being wrong81%
- passGets the base of every number right81%
Run
- Run
- #1
- Time to output
- 23 s
- Submitted
- 25 Sept 2026