Usable with a quick edit9 of 10 checks passed
The call and evidence are sound; a PM would only need to add specific post-launch monitoring thresholds and a rollback trigger.
Output
Recommendation: Go with conditions; otherwise no-go for Monday.
The pilot met two launch criteria: 64% of drafts were accepted as-is or with minor edits, above the 60% target, and median time to first response fell 41%, exceeding the 30% target. Seven of eight agents also want to keep the feature.
But the launch does not currently meet the safety criteria. The eval found harmful drafts in 1.8% of cases, above the 1% ceiling. Six errors came from outdated billing prices—a significant exposure because billing handles 38% of tickets. And two drafts promised refunds, directly violating the “no draft may commit to a refund” criterion, even though agents caught them.
Before enabling drafts broadly, require: - Updated billing reference data and a rerun of the 500-ticket eval with no more than five wrong or harmful drafts. - A reliable safeguard against refund commitments, verified to produce zero such drafts. - Welsh-language tickets routed away from AI drafting until that issue is fixed.
If those gates pass before Monday, proceed with the launch and monitor closely. If they do not, delay the launch rather than rely on agents to catch known failure modes. The formatting issue should also be addressed or clearly covered in agent training, but it is not the main launch blocker.
Check by check
Got wrong · 1
- Limits the downside of being wrongIt says monitor closely but does not give a specific post-launch signal, threshold, pause/rollback action, or flag irreversible exposures.
Got right · 9
- Uses the supplied evidence correctlyAll factual claims about the current situation come from the supplied pilot, eval, bug, or support-ops context or follow by arithmetic.
- Addresses the actual decisionIt commits clearly to go with conditions for Monday and says the gates that would change the call to full go or delay.
- Respects explicit constraintsIt stays well under 400 words, addresses the launch meeting, and its proposed conditions would actually enforce the failed criteria.
- Identifies material uncertaintyIt names the unresolved gates—billing data rerun, refund safeguard, Welsh routing—and says before-Monday pass/fail changes the call.
- Avoids unsupported claimsIt does not present agent catch as acceptable or invent causes; its billing-exposure point is supported by the 38% billing share.
- Produces the required deliverableThe recommendation is complete, under the length limit, aimed at the meeting, and usable with only light edits.
- Checks every agreed criterionIt explicitly tests all four criteria: 64% acceptance, 1.8% error rate, refund language, and 41% time reduction.
- Makes a clear callIt makes one clear call—go with conditions, otherwise no-go—and states what would change it.
- Gets the base of every number rightIts derived percentages use correct denominators and steps: 9/500, 64% pilot acceptance, and 40.8% median-time reduction.
Grades and run details
Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
- passUses the supplied evidence correctly90%
- passAddresses the actual decision97%
- passRespects explicit constraints71%
- passIdentifies material uncertainty99%
- passAvoids unsupported claims76%
- passProduces the required deliverable95%
- passChecks every agreed criterion77%
- passMakes a clear call98%
- partialLimits the downside of being wrong47%
- passGets the base of every number right88%
Run
- Run
- #1
- API response time
- 11 s
- Submitted
- 29 Sept 2026