Needs more than a quick edit8 of 12 checks passed
The spec exceeds the 700-word limit and lacks guardrail thresholds, so a PM would need to trim it and add those thresholds before use.
Output
Experiment Spec: Delivery Fee Shown in Basket
Owner: Checkout PM · Build: Ravi · Analysis: Ines · Approver: Chloe
Why we're testing this
Today the £3.99 delivery fee (free over £60) first appears on the final checkout step. Last quarter, 12% of support contacts were about unexpected delivery fees. Design expects that showing the fee earlier will build trust. Chloe's concern is that it will put people off before they start checkout. This test tells us which effect is bigger.
Hypothesis
Showing the delivery fee in the basket will change the share of basket viewers who place an order. We don't assume a direction. The test is two-sided because either outcome is plausible.
Variants
- Control: Current experience. The fee appears only at the final checkout step.
- Variant: The basket shows a "Delivery: £3.99" line, or "Free" when the basket is over £60. The same fee still appears at checkout.
Change from draft: The free-delivery-over-£60 banner is removed from this test. If both changes ship together, we can't tell which one caused any result. The banner could also raise orders while the fee line lowers them, so the two effects would cancel and we'd learn nothing. We'll test the banner as a follow-up once this result is in.
Primary metric
Order conversion: the share of customers who view their basket during the test and go on to place at least one order.
Change from draft: The draft used checkout completion (orders ÷ checkout starts). That metric would be misleading here. If the variant puts off fee-sensitive customers before checkout, the people who remain in checkout are the ones already willing to pay. Completion would then rise even if total orders fell. Order conversion counts everyone who saw the basket, so it catches losses at either stage.
Secondary and guardrail metrics
- Checkout start rate and checkout completion rate. These are diagnostic only and show where any change happens.
- Average order value. Seeing the fee may push some customers over £60.
- Revenue per basket viewer. This includes delivery fee income.
- Fee-related support contacts per 1,000 orders. This is the problem we set out to fix.
Randomisation
- Unit: customer (logged-in account ID). The draft split by session, which needs to change. About half of each week's basket viewers are returning customers. With a session split, they would see the fee in the basket on one visit and not on the next. That would contaminate both groups and confuse customers.
- Split: 50/50, sticky for the whole test.
- Ravi: Please log the assignment when a customer first views their basket, and confirm that logged-out traffic is either excluded or bucketed consistently.
Sample size and duration
- Assumptions: baseline 8% order conversion, two-sided test at α = 0.05, 80% power.
- Minimum effect we want to detect: 0.5 percentage points (8.0% → 8.5%). We're reading Chloe's threshold as absolute points on order conversion. That is roughly a 6% relative change.
- Customers needed: about 47,500 per arm, or about 95,000 total.
How long that takes: About 40,000 customers view their basket in week 1, then about 20,000 new customers each week after. That gives roughly 40k, 60k, 80k and 100k cumulative customers. The test therefore runs for 4 full weeks.
Caveat for Ines: The 8% baseline is per basket viewer per week. Over a 4-week window, each customer has more chances to order, so the per-customer baseline will be higher. At 12%, for example, we'd need about 135,000 customers, which means about 6 weeks. Before launch, Ines will calculate the 4-week per-customer baseline from historical data. The final duration will be fixed from that number and recorded here.
Stopping rule
Change from draft: We won't stop the test as soon as the dashboard shows significance. If we check every day and stop on the first significant result, the false-positive rate climbs well above 5%. The dashboard will often show early "wins" that aren't real.
- The test runs for the pre-set duration, in full weeks only.
- The dashboard can be watched for bugs and sample-ratio mismatch. Nobody makes ship decisions from it mid-test.
- Only reason to stop early: a guardrail failure. That means a broken experience, a sample-ratio mismatch (p < 0.001), or order conversion down by more than 1.5 points.
Decision rules
The confidence interval (CI) below is the interval around the difference in order conversion between the variant and control.
| Result on order conversion | Decision |
|---|---|
| Up by a significant amount | Ship the variant |
| Down by a significant amount | Don't ship. Look into other ways to reduce the support contacts |
| Not significant, and the CI excludes −0.5 points | Ship. We lose no meaningful number of orders and gain trust and fewer support contacts |
| Not significant, and the CI includes −0.5 points | Treat as inconclusive. Chloe decides whether to extend the test or drop the change |
Sign-off
- ☐ Ravi: build plus assignment logging
- ☐ Ines: baseline check and final duration
- ☐ Chloe: approval
Check by check
Got wrong · 1
- Guardrails with thresholdsGuardrail metrics are listed but no thresholds are given for average order value, revenue, or support contacts to block rollout; only order conversion has a stop threshold.
Mixed · 3
- Uses the supplied evidence correctlyAll claims about the current situation are directly from the brief or derived by correct arithmetic.The two graders disagreed on this one.
- Respects explicit constraintsThe output is approximately 780 words, exceeding the 'under 700 words' limit.The two graders disagreed on this one.
- Produces the required deliverableThe spec is over the 700-word limit, so it does not fully meet the requested form.The two graders disagreed on this one.
Got right · 8
- Addresses the actual decisionThe spec commits to a clear decision framework with rules for shipping, not shipping, and inconclusive results, framed for Chloe.
- Identifies material uncertaintyIt identifies the uncertain per-customer baseline over 4 weeks and the inconclusive outcome, with a plan to resolve the baseline.
- Avoids unsupported claimsHypotheses and forecasts are clearly labelled, and no confident claims go beyond the supplied evidence.
- Tests one change at a timeThe free-delivery banner is removed from the variant and the spec explains why bundling would make results uninterpretable.
- Fixes the duration up frontIt sets a fixed 4-week duration and explains that daily peeking inflates false positives.
- An unambiguous primary metricOrder conversion is the single primary metric with a clear rationale, and a sample-ratio mismatch check is included.
- Decision rule written before the testAll outcomes (significant up, significant down, non-significant with CI excluding -0.5pp, non-significant with CI including -0.5pp) map to stated actions.
- Sized from the real trafficSample size is calculated from the 8% baseline and 0.5pp effect, duration follows from traffic accumulation, and it runs full weeks.
Grades and run details
Decision model 79 · LLM judge 9 of 13 checks
Decision model checks
- failUses the supplied evidence correctly3%
- passAddresses the actual decision71%
- passRespects explicit constraints30%
- passIdentifies material uncertainty92%
- passAvoids unsupported claims45%
- passProduces the required deliverable59%
- passTests one change at a time100%
- passFixes the duration up front98%
- partialAn unambiguous primary metric41%
- passDecision rule written before the test94%
- partialSized from the real traffic34%
- partialGuardrails with thresholds88%
Run
- Run
- #1
- Time to output
- 31 s
- Submitted
- 30 Sept 2026