Needs more than a quick edit10 of 12 checks passed
Guardrail thresholds are vague ('materially worse', 'drops meaningfully'), so a PM would need to define them before the test can run.
Output
Experiment spec: delivery fee in basket
Owner: PM, Checkout | Readers: Ravi (build), Ines (analysis), Chloe (approval)
Hypothesis
Showing the delivery fee in the basket reduces surprise at checkout (12% of last quarter's support contacts) without reducing orders. Chloe's concern is that an early fee deters people before they start checkout, so the test measures that directly.
Changes from the team draft
| Draft | Problem | Fix |
|---|---|---|
| Primary metric: checkout completion | It only counts people who already started checkout. If the early fee puts people off, they never enter the denominator, so the harm would be hidden. | Primary metric is orders per basket viewer. |
| Split by session | Returning customers (about half of weekly viewers) would see both versions. | Randomise by customer. |
| Check daily, stop when significant | Repeated peeking inflates false positives. | Fixed sample size and duration, one readout. |
| Add the free-delivery banner to the variant | We couldn't tell whether the fee or the banner caused any change. | Banner excluded. Test it separately afterwards. |
Design
- Control: fee first shown at final checkout step (today).
- Variant: basket shows the delivery line: £3.99, or "Free" at £60+. Nothing else changes. The basket and checkout amounts must always match.
- Unit: customer ID, with a 50/50 split. Assignment is sticky for the whole test. Logged-out users are bucketed by a persistent device ID (Ravi to confirm how many there are).
- Enrolment: at first basket view during the test. Log an exposure event at that moment.
Metrics
Primary: share of enrolled customers who place at least one order within 7 days of first basket view. Baseline is about 8%.
Guardrails: - Checkout start rate (baseline about 12.9%, i.e. 8% ÷ 62%). - Revenue per enrolled customer (baseline about £3.76, i.e. 8% × £47). - Delivery-fee support contacts per 1,000 customers. This is the benefit metric, but volumes are low, so read it directionally.
Diagnostics (not decision metrics): - Checkout completion (orders ÷ checkout starts). Read it only alongside the primary metric, since it's conditional on starts. - New vs returning customers, and baskets under vs over £60. Both are exploratory.
Sample size and duration
- Smallest effect worth detecting: 0.5 percentage points on orders (8.0% → 8.5%), per Chloe.
- Requirement: 80% power, 5% two-sided significance. This needs about 46,000 customers per arm, roughly 92,000 in total.
- Time to reach it: Week 1 brings about 40,000 unique customers. After that, only about 20,000 new customers enrol per week, because returning customers are already assigned. That reaches about 100,000 by the end of week 4.
- Duration: enrol for 4 full weeks, plus 7 days of follow-up for the last cohort. Readout is at about 5 weeks, with no early stopping. Ines can monitor data quality but not results.
- Limit: the test can't reliably detect effects smaller than 0.5pp. A "no significant difference" result means "no effect larger than about 0.5pp", not "no effect".
Data quality checks (before readout)
- Sample ratio check: arms should be within normal variation of 50/50.
- Exposure is logged for every enrolled customer.
- Fee shown in basket equals fee charged, in all cases.
Decision rule (Chloe to confirm)
- Ship if the 95% confidence interval for the change in orders rules out a drop of 0.5pp or more, no guardrail is materially worse, and delivery-fee contacts do not rise.
- Don't ship if orders fall by 0.5pp or more, or revenue per customer drops meaningfully.
- Otherwise, treat as inconclusive and discuss. Do not extend the test to "find" significance.
Next
If the variant ships or is neutral, run a follow-up test of the free-delivery-over-£60 banner, where the fee is already visible. Its effect is likely concentrated in baskets just under £60.
Open items
- Ravi: logged-out identity handling; a flag to switch the variant off quickly.
- Chloe: confirm the decision rule and that the 0.5pp threshold means absolute percentage points on orders per basket viewer.
- Ines: confirm the 7-day order attribution and the support-contact tagging for fee-related contacts.
Grades and run details
Decision model 79 · LLM judge 11 of 13 checks
Decision model checks
- failUses the supplied evidence correctly4%
- passAddresses the actual decision61%
- partialRespects explicit constraints23%
- passIdentifies material uncertainty80%
- passAvoids unsupported claims35%
- passProduces the required deliverable67%
- passTests one change at a time99%
- passFixes the duration up front100%
- passAn unambiguous primary metric96%
- passDecision rule written before the test95%
- partialSized from the real traffic31%
- partialGuardrails with thresholds94%
Run
- Run
- #1
- API response time
- 29 s
- Submitted
- 30 Sept 2026