Needs more than a quick edit6 of 9 checks passed
The unsupported claim about ticket splitting and the missing trust check require verification before a PM could safely act.
Output
Shorter checklist test — what it tells us and what's next
Hi all,
The shorter onboarding checklist came back "not significant." That doesn't make it a failure, and I don't think we should drop it.
What the test actually says
- B activated at 32.9% vs 31.2% for A, a +1.7pp lift.
- The 95% interval runs from −1.4pp to +4.8pp. Most of that range is positive, and the worst plausible case is a small loss.
- We sized the test to detect 3pp. A real 1–2pp gain was always likely to come back "not significant." The result means the effect is smaller than 3pp or zero.1 It doesn't mean the change didn't work.
Supporting signal
Support tickets tagged "onboarding confusion" fell from 44 to 29 during the test. We didn't split tickets by arm, so this is directional only. It does point the same way as the activation result.
Recommendation: ship it
- It's already built, and rolling it out costs about a day.
- The likely effect is neutral to positive, and five steps are simpler to maintain than nine.
- Running four more weeks would only detect effects of about 2pp. Our point estimate is 1.7pp, so we would probably still get an ambiguous answer, a month later.
Guardrails
- Watch 7-day activation for the four weeks after rollout. If it drops more than about 1.5pp below the current baseline, roll back.
- Split onboarding-confusion tickets by checklist version from now on, so next time we can tell whether a ticket drop is real.
For future tests
Before launch, agree what we'll do if the result comes back inconclusive. For a cheap, low-risk change like this, "not significant but probably not harmful" should default to ship.
Happy to discuss.
What a PM had to fix
1Numbers wrongRedo the arithmeticTargeted repair
Source: Readout
Activation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).
What we’d changeCorrect the reading: the interval runs from −1.4 to +4.8pp, so an effect above 3pp isn't ruled out. The test just wasn't sized to reliably detect smaller ones.
Check by check
Got wrong · 3
- Uses the supplied evidence correctlyThe claim 'We didn't split tickets by arm' is not stated in the supplied context and has no support.
- Avoids unsupported claimsPresents 'We didn't split tickets by arm' as an established fact without evidence or qualification.
- Trusts the data before reading itNo check of sample ratio, logging issues, or any trust signal before interpreting the result.
Got right · 6
- Addresses the actual decisionCommits unambiguously to 'ship it', specifies monitoring and rollback conditions.
- Respects explicit constraintsDelivers a short note under 300 words to the onboarding team addressing conclusions and next steps.
- Identifies material uncertaintyNames the uncertainty around effect size, bounds it with the confidence interval, and defines a rollback trigger.
- Produces the required deliverableThe output is a complete, usable note in the requested form and within length, ready with light edits.
- Interprets power correctlyExplains the test was powered to detect 3pp and that a 1–2pp effect would likely be non-significant.
- Gets the base of every number rightAll percentages and differences are derived correctly from the supplied activation rates and counts.
Claims the judge couldn’t find in the brief
- We didn't split tickets by arm.
Grades and run details
Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
- failUses the supplied evidence correctly23%
- passAddresses the actual decision92%
- passRespects explicit constraints49%
- passIdentifies material uncertainty97%
- partialAvoids unsupported claims23%
- passProduces the required deliverable89%
- passInterprets power correctly98%
- failTrusts the data before reading it83%
- passGets the base of every number right80%
Run
- Run
- #1
- Time to output
- 13 s
- Submitted
- 25 Sept 2026