Needs more than a quick edit3 of 10 checks passed
A PM would need to add a plan to re-run the eval set after fixes and explicit post-launch monitoring triggers; the unsupported 'training cutoff' claim also needs correction.
Output
Launch Recommendation: Go with Conditions
As the Product Manager, my recommendation for Monday’s rollout to all 42 agents is Go with Conditions.
While the pilot successfully proved the core value proposition, the evaluation set revealed critical safety gaps that we must mitigate before scaling.3
---
Why We Are Ready (The "Go" Factors)
- Efficiency Gains Exceeded Target: Median time to first response dropped by 41% (from 7.1 to 4.2 hours), surpassing our 30% goal.
- High Team Adoption: 7 out of 8 pilot agents want to keep using the feature.
- Strong Overall Quality: 64% of drafts were accepted as-is or with minor edits, meeting our 60% threshold.
---
Critical Conditions for Monday's Launch
To protect our customers and brand, the following four conditions must be met prior to Monday:
1. Fix the Billing Knowledge Gap (Mandatory): * The Issue: 6 out of the 9 harmful drafts on the eval set quoted old prices because the model's training data cutoff is August4, and prices changed September 1st. * Action: Update the model's reference document with current pricing immediately. 2. Implement Guardrails for Refunds: * The Issue: Two drafts promised refunds, violating our zero-tolerance policy (though agents successfully caught both). * Action: Deploy a strict system prompt constraint blocking drafts from mentioning or committing to refunds.2 3. Address Known Minor Bugs: * The Issue: Welsh-language tickets generate English replies, and data tables are poorly formatted. * Action: Temporarily route Welsh tickets around the AI tool until fixed, and issue a known-issues note regarding table formatting during Friday's agent training. 4. Targeted Billing Training: * Because billing accounts for 38% of our ticket volume, dedicate a specific section of Friday's all-hands training to remind agents to double-check pricing figures in billing drafts.
---
By locking in these quick fixes before Friday's training session, we capture a 41% speed gain while keeping accuracy well within safe operational limits.1
What a PM had to fix
1Hypothesis stated as factReframe it as a hypothesisSubstantial rework
Source: Eval set (500 tickets)
Wrong or harmful drafts: 9 of 500 (1.8%).
What we’d changeRequire a re-run of the 500-ticket eval after the fixes, and launch only if it passes. Nothing yet shows accuracy is within the limit.
2Constraint missedRestore the constraintTargeted repair
Source: Launch criteria agreed in the PRD
No draft may commit to a refund.
What we’d changeA prompt instruction isn't a guarantee. Test the control, and add a check that stops refund commitments reaching agents' drafts.
3Constraint missedRestore the constraintTargeted repair
Source: Launch criteria agreed in the PRD
Wrong or harmful drafts (wrong account data, wrong policy, invented features) at most 1% on the 500-ticket eval set. 3. No draft may commit to a refund.
What we’d changeCheck each agreed criterion and say plainly that two fail: wrong or harmful drafts at 1.8% against 1%, and two refund commitments against none.
4Invented evidenceVerify or remove the claimQuick edit
Source: Eval set (500 tickets)
the model's reference data stops in August
What we’d changeThe brief says the model's reference data stops in August, not its training data. That's what makes a pricing refresh a quick fix.
Check by check
Got wrong · 5
- Addresses the actual decisionThe output does not state what result or condition would change the 'Go with Conditions' call (e.g., failing to meet the conditions would make it a no-go).
- Identifies material uncertaintyIt does not name specific unknowns that could change the decision or say how they would be resolved, such as whether the pricing fix alone will bring error rate below 1%.
- Checks every agreed criterionIt does not explicitly compare the 1.8% harmful rate against the 1% bar, leaving the second launch criterion unchecked.
- Makes a clear callThe call is clear but does not state what would change it, e.g., that failing to meet the conditions would lead to a no-go.
- Limits the downside of being wrongNo post-launch signal, threshold to pause or roll back, or irreversible actions are identified.
Mixed · 2
- Uses the supplied evidence correctlyThe claim about a 'training data cutoff' is not supported; the context only mentions reference data stopping in August, so an unsupported fact about the current system was introduced.The two graders disagreed on this one.
- Avoids unsupported claimsThe output presents the 'training data cutoff' as a fact when the source only says reference data stops in August, turning an interpretation into an unsupported claim.The two graders disagreed on this one.
Got right · 3
- Respects explicit constraintsThe output stays under 400 words and respects the requested form of a launch recommendation for the meeting.
- Produces the required deliverableThe recommendation is under 400 words, addressed to the meeting, and usable with light edits, though it lacks a retest plan.
- Gets the base of every number rightThe 41% time reduction uses the correct base (7.1 hours to 4.2 hours), and other percentages match the supplied data.
Claims the judge couldn’t find in the brief
- The model's training data cutoff is August
Grades and run details
Decision model 60 · LLM judge 3 of 11 checks
Decision model checks
- passUses the supplied evidence correctly56%
- partialAddresses the actual decision51%
- partialRespects explicit constraints20%
- partialIdentifies material uncertainty71%
- passAvoids unsupported claims20%
- partialProduces the required deliverable41%
- partialChecks every agreed criterion72%
- partialMakes a clear call75%
- failLimits the downside of being wrong96%
- passGets the base of every number right83%
Run
- Run
- #1
- Time to output
- 6 s
- Submitted
- 25 Sept 2026