Usable with a quick edit9 of 10 checks passed
The recommendation is complete, evidence-based, and actionable; the only meaningful gap is a missing post-launch pause threshold, which is a quick edit.
Output
Recommendation: Go with conditions
Not a clean go. Two of the four agreed launch criteria are currently failed, and I don't think we should waive criteria we set ourselves.
| Criterion | Result | Status |
|---|---|---|
| ≥60% accepted | 64% | Met |
| ≤1% wrong/harmful | 1.8% (9/500) | Failed |
| No refund commitments | 2 drafts promised refunds | Failed |
| ≥30% faster first response | −41% (7.1h → 4.2h) | Met |
The value is real: faster responses, good acceptance, and 7 of 8 pilot agents want to keep it. Both failures also have identifiable causes.
- Wrong drafts: Six of the nine errors come from stale pricing, because the reference data stops in August. Fixing that should bring us to about 3/500 (0.6%), but that is a projection until we re-run the eval.
- Refunds: Agents caught both drafts, but the criterion says no draft, and with 42 agents we shouldn't rely on catching them by habit.
Conditions for Monday
- Refresh the reference data with post-1 September pricing and re-run the full 500-ticket eval. It must come in at 1% or below (5 or fewer wrong).
- Add a refund guardrail that blocks or flags commitment language such as "we'll refund." It must show zero refund commitments on the re-run plus a targeted set of refund-request tickets.
- Kill switch and flag button. Ops can disable the feature instantly, and agents can flag bad drafts in one click.
Fallback
- If condition 2 fails, no-go for Monday.
- If only condition 1 fails, launch on the three non-billing queues (62% of tickets) and keep billing off until it passes.
- Friday training goes ahead for all 42 agents either way. It should cover the refund rule and checking prices before sending.
Known issues, not blockers
- Welsh tickets (~0.3%) get English replies. Agents will write these manually, and we'll suppress drafts for them if detection is feasible.
- Table formatting is cosmetic, since agents edit every draft. We'll fix it after launch.
After launch
- Sample wrong-draft rates weekly for the first month.
- Track how often agents edit or discard drafts, as a check on whether review discipline slips.
- Ask the one pilot agent who didn't want to keep it why.
I'd like the eval re-run results by Thursday so we can make the final call before Friday's training.
Grades and run details
Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
- passUses the supplied evidence correctly44%
- passAddresses the actual decision98%
- passRespects explicit constraints37%
- passIdentifies material uncertainty100%
- partialAvoids unsupported claims25%
- passProduces the required deliverable84%
- passChecks every agreed criterion100%
- passMakes a clear call100%
- partialLimits the downside of being wrong85%
- passGets the base of every number right79%
Run
- Run
- #1
- API response time
- 17 s
- Submitted
- 29 Sept 2026