Usable with a quick edit9 of 10 checks passed
The reasoning is sound, evidence is correctly used, and the recommendation is clear. A PM could add a monitoring threshold in a quick edit without redoing any analysis.
Output
Recommendation: no-go for Monday’s rollout to all 42 agents. Reassess once the failed safety gates have been fixed and retested.
The pilot demonstrates value: 64% of drafts were accepted as-is or with minor edits, exceeding the 60% target. Median first-response time fell 41%, exceeding the 30% target, although the before-and-after comparison does not establish causality. Seven of eight agents want to keep it.
But we failed two explicit launch criteria:
- Wrong or harmful drafts: 1.8% against a 1% ceiling. Nine failures is above the maximum five permitted on the 500-ticket set.
- No refund commitments: failed. Two drafts promised refunds. Agents catching them prevented those messages reaching customers; it does not satisfy a criterion about what drafts may say. Human review was already part of the agreed design.
Six errors used outdated billing prices. That gives us a concrete fix, but billing represents 38% of tickets, so this is a substantial exposure. Updating the reference data alone would not establish that the remaining errors or refund promises are resolved.
Required before reconsidering launch:
- Engineering and the billing policy owner: update and verify current prices and policies, with a named owner and process for keeping reference data current. Investigate the other three harmful drafts.
- Engineering: add and test a control that prevents refund commitments from appearing in agent-facing drafts, including paraphrases.
- PM and evaluation owner: rerun the full 500-ticket evaluation after the fixes, meeting the original thresholds, and add targeted pricing and refund tests. Report those additional results separately.
Excluding billing is a possible narrower release, not an automatic pass: we lack queue-level results and do not know whether refund failures are confined to billing. It would require explicit scope approval and validation of both routing and the remaining queues.
Welsh-language tickets should bypass drafting until language handling is fixed. Disable tables or use plain text pending the formatting fix.
Keep Friday’s training as preparation, with launch clearly marked as pending. The booking should not determine readiness. After the gates pass, expand in stages with monitoring and a named owner able to disable drafting.
Check by check
Got wrong · 1
- Limits the downside of being wrongNo specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
Got right · 9
- Uses the supplied evidence correctlyAll factual statements about the current situation are directly from the supplied brief or context, or follow from arithmetic; no invented facts.
- Addresses the actual decisionThe output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).
- Respects explicit constraintsThe output is a launch recommendation, addresses the meeting, and stays within the 400-word limit (345 words).
- Identifies material uncertaintyIt identifies the lack of queue-level results and the unknown scope of refund failures, and explains how those would be resolved.
- Avoids unsupported claimsThe note about causality is framed as a methodological observation, not as a factual claim; no unsupported causes or forecasts are presented as fact.
- Produces the required deliverableThe recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
- Checks every agreed criterionAll four launch criteria are explicitly assessed: acceptance rate, error rate, refund commitment, and time reduction.
- Makes a clear callClear 'no-go' call up front, with a statement that reassessment follows fixes and retesting.
- Gets the base of every number rightAll derived figures (41% time reduction, 1.8% error rate, 9/500, 38% of tickets) are computed correctly and bases are clear.
Grades and run details
Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
- passUses the supplied evidence correctly76%
- passAddresses the actual decision89%
- passRespects explicit constraints44%
- passIdentifies material uncertainty96%
- passAvoids unsupported claims72%
- passProduces the required deliverable76%
- passChecks every agreed criterion44%
- passMakes a clear call92%
- partialLimits the downside of being wrong52%
- passGets the base of every number right84%
Run
- Run
- #1
- Time to output
- 34 s
- Submitted
- 25 Sept 2026