Needs more than a quick edit4 of 10 checks passed
A PM would need to re-check the causal claims and redo the uncertainty handling because the output treats inferences as established and does not specify what would change the call.
Output
Hi Lena,
Good catch—I’ve been digging into the numbers. While the new navigation is a factor, it’s not the whole story.1 Here’s what happened:
- The SSO Outage: On Thursday, 17 Sept, a two-hour single sign-on failure locked out ~30% of our teams2, directly impacting mid-week engagement.
- Navigation Friction: Support tickets about navigation jumped from 40 to 212, mostly with “where did X go?” questions. However, teams that were in the earlier beta only dropped 1.1%, suggesting the learning curve is real but manageable once users adjust.3
- Seasonality: Weekly active teams dropped 4% the exact same week last year due to the start of the school year.4
Next Steps: We’re shipping contextual tooltips for the moved features today and monitoring daily active recovery. I'll share a full breakdown in our standup.
What a PM had to fix
1Hypothesis stated as factReframe it as a hypothesisStart again
Source: Metrics
Mobile only: −5.2%.
What we’d changeLead with the strongest comparison: mobile fell 5.2% with no navigation change, so the navigation isn't established as a cause. Say what you'll check and when.
2Invented evidenceVerify or remove the claimSubstantial rework
Source: Incident
About 30% of teams use single sign-on.
What we’d change30% of teams use SSO. That doesn't mean all of them were trying to get in during those two hours.
3Hypothesis stated as factReframe it as a hypothesisQuick edit
What we’d changeBeta teams may differ from everyone else. Their small drop doesn't show a manageable learning curve.
4Hypothesis stated as factReframe it as a hypothesisQuick edit
Source: Metrics
which we put down to the start of the school year.
What we’d changeLast year's cause was an attribution, not a finding. Say 'which we put down to'.
Check by check
Got wrong · 6
- Uses the supplied evidence correctlyIt introduces causal statements such as 'directly impacting mid-week engagement' and 'new navigation is a factor' that are not in the supplied context or derivable by arithmetic.
- Addresses the actual decisionIt does not say what result or condition would change the answer; it only says it will monitor and share a breakdown.
- Identifies material uncertaintyIt does not name the specific unknowns that could change the decision or tie the monitoring to what would settle the call.
- Avoids unsupported claimsIt presents causal interpretations such as the SSO outage directly impacting engagement and navigation being a factor as established fact, though the evidence only suggests them.
- Separates effect from explanationIt lists SSO outage, navigation friction, and seasonality as 'what happened' rather than clearly labelling them as hypotheses with supporting or contradicting evidence.
- Owns the problemIt reports the SSO outage but does not say what the team has already done about it or recommend a way forward for that specific problem.
Got right · 4
- Respects explicit constraintsIt is a channel reply addressed to Lena and is under 200 words.
- Produces the required deliverableThe reply is in the right channel format, addressed to Lena, under 200 words, and usable as a status update.
- Answers the question askedThe first lines answer the question by saying the new navigation is a factor but not the whole story.
- Ready for the reader as deliveredIt is written for Lena, uses no placeholders, and all dates and owners are taken from the material or clearly proposed.
Claims the judge couldn’t find in the brief
- The new navigation is a factor, though not the whole story.
- The SSO outage directly impacted mid-week engagement.
- This suggests the learning curve is real but manageable once users adjust.
Grades and run details
Decision model 55 · LLM judge 4 of 11 checks
Decision model checks
- failUses the supplied evidence correctly38%
- partialAddresses the actual decision73%
- passRespects explicit constraints55%
- partialIdentifies material uncertainty69%
- failAvoids unsupported claims23%
- passProduces the required deliverable57%
- partialAnswers the question asked32%
- partialSeparates effect from explanation68%
- partialOwns the problem94%
- passReady for the reader as delivered84%
Run
- Run
- #1
- Time to output
- 5 s
- Submitted
- 27 Sept 2026