Usable with a quick edit9 of 12 checks passed
The core reasoning and evidence are correct; a quick edit to add a sample ratio check, guardrail thresholds and a one-line explanation for no peeking would make it fully ready.
Output
Basket delivery-fee experiment
For: Ravi (build), Ines (analysis), Chloe (approval) Decision requested: Approve a test of showing the existing delivery fee in the basket.
Hypothesis and scope
Showing the delivery cost before checkout will reduce surprise and fee-related support contacts, without materially reducing orders.
- Control: Current experience; delivery fee first appears on the final checkout step.
- Variant: Show the applicable fee in the basket: “Delivery £3.99” below the free-delivery threshold, or “Free delivery” when the basket qualifies. Keep the existing “free on orders over £60” rule and wording.
- Change nothing else. Do not add the free-delivery banner: that would test a second intervention and make the result hard to interpret.
Build and assignment — Ravi
Randomise 50/50 by customer, not session, and keep each customer in the same arm across visits and devices where identity is available. Stratify assignment by new versus returning basket viewers. Exclude staff and test accounts. Log assignment, basket value/fee shown, basket views, checkout starts, orders, order value, and relevant support contacts.
Measurement — Ines
Primary metric: Basket-viewer conversion: percentage of assigned customers who place at least one order within 7 days of their first eligible basket view. Count each customer once, in their assigned arm (intention-to-treat). This captures any deterrence before checkout; checkout completion alone would miss it.
Guardrails / secondary measures: Checkout starts per basket viewer; checkout completion (orders ÷ checkout starts); average order value and revenue per assigned customer; delivery-fee-related support contacts per assigned customer. Support contacts may be too sparse for a definitive test, so treat them as directional, not a substitute for the primary outcome.
Sample, duration, and analysis
Baseline basket-viewer conversion is 8%. Power the test to detect a 0.5 percentage-point absolute difference (Chloe’s smallest worthwhile order effect), with 80% power and a two-sided 5% significance level. This requires approximately 95,000 unique customers total (about 47,500 per arm), assuming independent customer-level outcomes.
There are about 40,000 basket viewers weekly, but many return across weeks. Run for at least four full weeks and until the sample target is reached; then allow 7 days for outcomes to mature before reading results. If the target is not reached, extend in full weeks. Ines should confirm the final sample-size calculation and report actual enrolment.
Do not stop early for significance or check daily results to make a stopping decision. Monitor only for implementation or customer-harm issues. Analyse the primary metric at the customer level, with a 95% confidence interval and the pre-specified two-sided test.
Decision
Recommend rollout if the variant improves basket-viewer conversion by at least 0.5 percentage points, the result is statistically significant, and guardrails show no material harm. Do not roll out if there is significant harm. If results are inconclusive, report that clearly; do not interpret “not significant” as proof of no effect.
Grades and run details
Decision model 88 · LLM judge 10 of 13 checks
Decision model checks
- passUses the supplied evidence correctly14%
- passAddresses the actual decision54%
- partialRespects explicit constraints17%
- passIdentifies material uncertainty35%
- passAvoids unsupported claims64%
- passProduces the required deliverable50%
- passTests one change at a time94%
- passFixes the duration up front62%
- passAn unambiguous primary metric40%
- partialDecision rule written before the test27%
- passSized from the real traffic39%
- partialGuardrails with thresholds98%
Run
- Run
- #1
- API response time
- 22 s
- Submitted
- 30 Sept 2026