Needs more than a quick edit8 of 12 checks passed
The spec exceeds the 700-word limit and lacks guardrail thresholds, so a PM would need to trim it and add those thresholds before use.
Output
Experiment Spec: Delivery Fee Shown in Basket
Owner: Checkout PM · Build: Ravi · Analysis: Ines · Approver: Chloe
Why we're testing this
Today the £3.99 delivery fee (free over £60) first appears on the final checkout step. Last quarter, 12% of support contacts were about unexpected delivery fees. Design expects that showing the fee earlier will build trust. Chloe's concern is that it will put people off before they start checkout. This test tells us which effect is bigger.
Hypothesis
Showing the delivery fee in the basket will change the share of basket viewers who place an order. We don't assume a direction. The test is two-sided because either outcome is plausible.
Variants
- Control: Current experience. The fee appears only at the final checkout step.
- Variant: The basket shows a "Delivery: £3.99" line, or "Free" when the basket is over £60. The same fee still appears at checkout.
Change from draft: The free-delivery-over-£60 banner is removed from this test. If both changes ship together, we can't tell which one caused any result. The banner could also raise orders while the fee line lowers them, so the two effects would cancel and we'd learn nothing. We'll test the banner as a follow-up once this result is in.
Primary metric
Order conversion: the share of customers who view their basket during the test and go on to place at least one order.
Change from draft: The draft used checkout completion (orders ÷ checkout starts). That metric would be misleading here. If the variant puts off fee-sensitive customers before checkout, the people who remain in checkout are the ones already willing to pay. Completion would then rise even if total orders fell. Order conversion counts everyone who saw the basket, so it catches losses at either stage.
Secondary and guardrail metrics
- Checkout start rate and checkout completion rate. These are diagnostic only and show where any change happens.
- Average order value. Seeing the fee may push some customers over £60.
- Revenue per basket viewer. This includes delivery fee income.
- Fee-related support contacts per 1,000 orders. This is the problem we set out to fix.
Randomisation
- Unit: customer (logged-in account ID). The draft split by session, which needs to change. About half of each week's basket viewers are returning customers. With a session split, they would see the fee in the basket on one visit and not on the next. That would contaminate both groups and confuse customers.
- Split: 50/50, sticky for the whole test.
- Ravi: Please log the assignment when a customer first views their basket, and confirm that logged-out traffic is either excluded or bucketed consistently.
Sample size and duration
- Assumptions: baseline 8% order conversion, two-sided test at α = 0.05, 80% power.
- Minimum effect we want to detect: 0.5 percentage points (8.0% → 8.5%). We're reading Chloe's threshold as absolute points on order conversion. That is roughly a 6% relative change.
- Customers needed: about 47,500 per arm, or about 95,000 total.
How long that takes: About 40,000 customers view their basket in week 1, then about 20,000 new customers each week after. That gives roughly 40k, 60k, 80k and 100k cumulative customers. The test therefore runs for 4 full weeks.
Caveat for Ines: The 8% baseline is per basket viewer per week. Over a 4-week window, each customer has more chances to order, so the per-customer baseline will be higher. At 12%, for example, we'd need about 135,000 customers, which means about 6 weeks. Before launch, Ines will calculate the 4-week per-customer baseline from historical data. The final duration will be fixed from that number and recorded here.
Stopping rule
Change from draft: We won't stop the test as soon as the dashboard shows significance. If we check every day and stop on the first significant result, the false-positive rate climbs well above 5%. The dashboard will often show early "wins" that aren't real.
- The test runs for the pre-set duration, in full weeks only.
- The dashboard can be watched for bugs and sample-ratio mismatch. Nobody makes ship decisions from it mid-test.
- Only reason to stop early: a guardrail failure. That means a broken experience, a sample-ratio mismatch (p < 0.001), or order conversion down by more than 1.5 points.
Decision rules
The confidence interval (CI) below is the interval around the difference in order conversion between the variant and control.
| Result on order conversion | Decision |
|---|---|
| Up by a significant amount | Ship the variant |
| Down by a significant amount | Don't ship. Look into other ways to reduce the support contacts |
| Not significant, and the CI excludes −0.5 points | Ship. We lose no meaningful number of orders and gain trust and fewer support contacts |
| Not significant, and the CI includes −0.5 points | Treat as inconclusive. Chloe decides whether to extend the test or drop the change |
Sign-off
- ☐ Ravi: build plus assignment logging
- ☐ Ines: baseline check and final duration
- ☐ Chloe: approval
Grades and run details
Decision model 83 · LLM judge 9 of 13 checks
Decision model checks
- failUses the supplied evidence correctly7%
- passAddresses the actual decision73%
- partialRespects explicit constraints16%
- passIdentifies material uncertainty90%
- passAvoids unsupported claims45%
- passProduces the required deliverable59%
- passTests one change at a time100%
- passFixes the duration up front98%
- passAn unambiguous primary metric39%
- passDecision rule written before the test94%
- passSized from the real traffic55%
- partialGuardrails with thresholds84%
Run
- Run
- #1
- Time to output
- 31 s
- Submitted
- 30 Sept 2026