Usable with a quick edit9 of 12 checks passed
The core reasoning and evidence are correct; a quick edit to add a sample ratio check, guardrail thresholds and a one-line explanation for no peeking would make it fully ready.
Output
Basket delivery-fee experiment
For: Ravi (build), Ines (analysis), Chloe (approval) Decision requested: Approve a test of showing the existing delivery fee in the basket.
Hypothesis and scope
Showing the delivery cost before checkout will reduce surprise and fee-related support contacts, without materially reducing orders.
- Control: Current experience; delivery fee first appears on the final checkout step.
- Variant: Show the applicable fee in the basket: “Delivery £3.99” below the free-delivery threshold, or “Free delivery” when the basket qualifies. Keep the existing “free on orders over £60” rule and wording.
- Change nothing else. Do not add the free-delivery banner: that would test a second intervention and make the result hard to interpret.
Build and assignment — Ravi
Randomise 50/50 by customer, not session, and keep each customer in the same arm across visits and devices where identity is available. Stratify assignment by new versus returning basket viewers. Exclude staff and test accounts. Log assignment, basket value/fee shown, basket views, checkout starts, orders, order value, and relevant support contacts.
Measurement — Ines
Primary metric: Basket-viewer conversion: percentage of assigned customers who place at least one order within 7 days of their first eligible basket view. Count each customer once, in their assigned arm (intention-to-treat). This captures any deterrence before checkout; checkout completion alone would miss it.
Guardrails / secondary measures: Checkout starts per basket viewer; checkout completion (orders ÷ checkout starts); average order value and revenue per assigned customer; delivery-fee-related support contacts per assigned customer. Support contacts may be too sparse for a definitive test, so treat them as directional, not a substitute for the primary outcome.
Sample, duration, and analysis
Baseline basket-viewer conversion is 8%. Power the test to detect a 0.5 percentage-point absolute difference (Chloe’s smallest worthwhile order effect), with 80% power and a two-sided 5% significance level. This requires approximately 95,000 unique customers total (about 47,500 per arm), assuming independent customer-level outcomes.
There are about 40,000 basket viewers weekly, but many return across weeks. Run for at least four full weeks and until the sample target is reached; then allow 7 days for outcomes to mature before reading results. If the target is not reached, extend in full weeks. Ines should confirm the final sample-size calculation and report actual enrolment.
Do not stop early for significance or check daily results to make a stopping decision. Monitor only for implementation or customer-harm issues. Analyse the primary metric at the customer level, with a 95% confidence interval and the pre-specified two-sided test.
Decision
Recommend rollout if the variant improves basket-viewer conversion by at least 0.5 percentage points, the result is statistically significant, and guardrails show no material harm. Do not roll out if there is significant harm. If results are inconclusive, report that clearly; do not interpret “not significant” as proof of no effect.
Check by check
Got wrong · 2
- An unambiguous primary metricIt names one primary metric with a rationale but omits a planned trust check such as a sample ratio check.
- Guardrails with thresholdsGuardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
Mixed · 1
- Fixes the duration up frontIt sets a fixed duration but does not explain why daily peeking inflates false positives, only instructs not to do it.The two graders disagreed on this one.
Got right · 9
- Uses the supplied evidence correctlyEvery statement about the current situation is taken directly from the supplied context or follows by arithmetic.
- Addresses the actual decisionThe output clearly requests approval of the test and provides a complete spec for Chloe to decide, with a post-test decision rule.
- Respects explicit constraintsThe output fixes all three problems in the draft, stays under 700 words, and is addressed to the three named readers.
- Identifies material uncertaintyIt notes support contacts may be too sparse for a definitive test and that inconclusive results are possible, with a plan for each.
- Avoids unsupported claimsThe hypothesis is clearly labelled, and no interpretations are presented as established fact.
- Produces the required deliverableThe spec is complete, under 700 words, and directly usable by Ravi, Ines and Chloe.
- Tests one change at a timeIt explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
- Decision rule written before the testIt maps rollout, no rollout and inconclusive results to actions, with thresholds of 0.5pp improvement and statistical significance.
- Sized from the real trafficSample size is calculated from the 8% baseline and 0.5pp effect, and the four-week duration accounts for returning visitors and whole-week cycles.
Grades and run details
Decision model 79 · LLM judge 10 of 13 checks
Decision model checks
- passUses the supplied evidence correctly12%
- passAddresses the actual decision62%
- partialRespects explicit constraints28%
- passIdentifies material uncertainty35%
- passAvoids unsupported claims66%
- passProduces the required deliverable62%
- passTests one change at a time95%
- passFixes the duration up front56%
- partialAn unambiguous primary metric83%
- partialDecision rule written before the test28%
- partialSized from the real traffic24%
- partialGuardrails with thresholds98%
Run
- Run
- #1
- API response time
- 22 s
- Submitted
- 30 Sept 2026