Needs more than a quick edit11 of 12 checks passed
Guardrail thresholds are vague ('materially worse', 'drops meaningfully'), so a PM would need to define them before the test can run.
Output
Experiment spec: delivery fee in basket
Owner: PM, Checkout | Readers: Ravi (build), Ines (analysis), Chloe (approval)
Hypothesis
Showing the delivery fee in the basket reduces surprise at checkout (12% of last quarter's support contacts) without reducing orders. Chloe's concern is that an early fee deters people before they start checkout, so the test measures that directly.
Changes from the team draft
| Draft | Problem | Fix |
|---|---|---|
| Primary metric: checkout completion | It only counts people who already started checkout. If the early fee puts people off, they never enter the denominator, so the harm would be hidden. | Primary metric is orders per basket viewer. |
| Split by session | Returning customers (about half of weekly viewers) would see both versions. | Randomise by customer. |
| Check daily, stop when significant | Repeated peeking inflates false positives. | Fixed sample size and duration, one readout. |
| Add the free-delivery banner to the variant | We couldn't tell whether the fee or the banner caused any change. | Banner excluded. Test it separately afterwards. |
Design
- Control: fee first shown at final checkout step (today).
- Variant: basket shows the delivery line: £3.99, or "Free" at £60+. Nothing else changes. The basket and checkout amounts must always match.
- Unit: customer ID, with a 50/50 split. Assignment is sticky for the whole test. Logged-out users are bucketed by a persistent device ID (Ravi to confirm how many there are).
- Enrolment: at first basket view during the test. Log an exposure event at that moment.
Metrics
Primary: share of enrolled customers who place at least one order within 7 days of first basket view. Baseline is about 8%.
Guardrails: - Checkout start rate (baseline about 12.9%, i.e. 8% ÷ 62%). - Revenue per enrolled customer (baseline about £3.76, i.e. 8% × £47). - Delivery-fee support contacts per 1,000 customers. This is the benefit metric, but volumes are low, so read it directionally.
Diagnostics (not decision metrics): - Checkout completion (orders ÷ checkout starts). Read it only alongside the primary metric, since it's conditional on starts. - New vs returning customers, and baskets under vs over £60. Both are exploratory.
Sample size and duration
- Smallest effect worth detecting: 0.5 percentage points on orders (8.0% → 8.5%), per Chloe.
- Requirement: 80% power, 5% two-sided significance. This needs about 46,000 customers per arm, roughly 92,000 in total.
- Time to reach it: Week 1 brings about 40,000 unique customers. After that, only about 20,000 new customers enrol per week, because returning customers are already assigned. That reaches about 100,000 by the end of week 4.
- Duration: enrol for 4 full weeks, plus 7 days of follow-up for the last cohort. Readout is at about 5 weeks, with no early stopping. Ines can monitor data quality but not results.
- Limit: the test can't reliably detect effects smaller than 0.5pp. A "no significant difference" result means "no effect larger than about 0.5pp", not "no effect".
Data quality checks (before readout)
- Sample ratio check: arms should be within normal variation of 50/50.
- Exposure is logged for every enrolled customer.
- Fee shown in basket equals fee charged, in all cases.
Decision rule (Chloe to confirm)
- Ship if the 95% confidence interval for the change in orders rules out a drop of 0.5pp or more, no guardrail is materially worse, and delivery-fee contacts do not rise.
- Don't ship if orders fall by 0.5pp or more, or revenue per customer drops meaningfully.
- Otherwise, treat as inconclusive and discuss. Do not extend the test to "find" significance.
Next
If the variant ships or is neutral, run a follow-up test of the free-delivery-over-£60 banner, where the fee is already visible. Its effect is likely concentrated in baskets just under £60.
Open items
- Ravi: logged-out identity handling; a flag to switch the variant off quickly.
- Chloe: confirm the decision rule and that the 0.5pp threshold means absolute percentage points on orders per basket viewer.
- Ines: confirm the 7-day order attribution and the support-contact tagging for fee-related contacts.
Check by check
Got wrong · 1
- Guardrails with thresholdsGuardrails are named but lack specific numeric thresholds; 'materially worse' and 'drops meaningfully' are too vague to block rollout unambiguously.
Got right · 11
- Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the brief or derived by correct arithmetic.
- Addresses the actual decisionThe spec commits to a clear decision rule with ship/don't ship/inconclusive actions, framed for Chloe to confirm.
- Respects explicit constraintsThe output is under 700 words, addresses the named readers, and is a complete experiment spec.
- Identifies material uncertaintyIt names unknowns like logged-out identity, decision rule confirmation, and low support-contact volumes, and says how to resolve them.
- Avoids unsupported claimsHypotheses and forecasts are clearly labelled, and no interpretations are presented as established fact.
- Produces the required deliverableThe spec is complete, in the right form, and usable by Ravi, Ines, and Chloe with light edits.
- Tests one change at a timeThe free-delivery banner is excluded from the variant and a separate follow-up test is proposed, with the attribution problem explained.
- Fixes the duration up frontDuration is fixed at 4 weeks enrollment plus 7 days follow-up, with no early stopping, and the false-positive risk of peeking is explained.
- An unambiguous primary metricThe primary metric is orders per enrolled customer within 7 days, with a clear rationale and a planned sample ratio check.
- Decision rule written before the testEvery outcome (ship, don't ship, inconclusive) is mapped to an action with thresholds separating them.
- Sized from the real trafficSample size is correctly calculated from the 8% baseline and 0.5pp effect, and duration follows from the traffic pattern over whole weeks.
Grades and run details
Decision model 92 · LLM judge 11 of 13 checks
Decision model checks
- passUses the supplied evidence correctly5%
- passAddresses the actual decision62%
- passRespects explicit constraints31%
- passIdentifies material uncertainty81%
- passAvoids unsupported claims34%
- passProduces the required deliverable72%
- passTests one change at a time98%
- passFixes the duration up front100%
- passAn unambiguous primary metric82%
- passDecision rule written before the test91%
- partialSized from the real traffic48%
- partialGuardrails with thresholds95%
Run
- Run
- #1
- API response time
- 29 s
- Submitted
- 30 Sept 2026