Usable with a quick edit9 of 12 checks passed
The spec is well-structured and fixes the draft's main issues; a PM could quickly add guardrail thresholds and lock the duration to full weeks.
Output
Experiment spec: delivery-fee visibility in basket
Owner: Checkout PM Build: Ravi · Analysis: Ines · Approval: Chloe
Objective and hypothesis
Test whether showing delivery costs before checkout increases orders by reducing surprises—or decreases orders by discouraging customers earlier.
Last quarter, 12% of support contacts concerned unexpected delivery fees. This motivates the test but is not the percentage of customers affected.
Experience
Control: Current basket; delivery fee first appears at the final checkout step.
Variant: Show a clearly labelled delivery-fee line and fee-inclusive total in the basket. Update both when the basket changes. Use the production pricing logic: £3.99 delivery, free for orders over £60. Ravi must verify boundary behaviour at exactly £60 and match existing checkout pricing.
Keep the final checkout fee display unchanged. Everything else stays identical.
Do not add the free-delivery banner. It could independently encourage larger baskets, making it impossible to attribute results to earlier fee visibility. Test it separately.
Eligibility and assignment
Include customers viewing an orderable basket in supported markets. Apply existing employee, test-account and bot exclusions identically, before assignment.
Randomise 50/50 by customer, not session, at their first eligible basket view. Persist assignment across visits and devices where identity is available. Use a stable anonymous identifier for guests and a documented account-linking rule; Ravi and Ines must validate identity handling before launch.
Each customer enters the experiment once. Analyse everyone assigned, regardless of whether they start checkout or successfully see the treatment.
Metrics
Primary: Seven-day order conversion: customers placing at least one order within seven days of assignment ÷ all assigned eligible basket viewers.
This captures Chloe’s concern about discouraging checkout entry. Orders ÷ checkout starts is unsuitable as primary: the treatment can change who starts checkout.
Secondary/diagnostic: - Checkout-start rate per assigned customer. - Checkout completion among starters, interpreted cautiously because starter populations may differ. - Revenue per assigned customer, including non-buyers. - AOV and proportion of orders qualifying for free delivery. - Unexpected-fee support contacts per assigned customer within seven days, using consistent contact classification.
Operational guardrails: Basket/checkout errors, incorrect fee calculations and page performance. These can trigger a safety rollback, not an early declaration of success.
Sample size and duration
Interpret Chloe’s minimum worthwhile change as 0.5 percentage points absolute: 8.0% to 8.5%, not a 0.5% relative lift.
With a matching 8% baseline, two-sided 5% significance and 80% power, plan for approximately 95,000 unique customers total, evenly split. Ines must verify the historical baseline uses the same seven-day definition and finalise the calculation before launch.
Do not count 40,000 weekly viewers as entirely new experiment entrants. Estimated cumulative reach is 40,000 in week one, then roughly 20,000 additional customers weekly: about four weeks to enrol the target, plus seven days for outcomes to mature.
Lock the enrolment target before launch; stop enrolment when reached, then wait seven days. Extend only if unique-customer accrual is slower—not based on results.
Analysis and decision
Ines reports the intention-to-treat conversion difference, 95% confidence interval and absolute rates. Check allocation imbalance, logging, identity consistency and treatment exposure before interpreting results.
No daily significance-based stopping. Daily dashboards are for operational safety only.
Recommend rollout if the primary improvement is statistically significant, the point estimate is at least +0.5 percentage points, and there are no unresolved operational or economic concerns.
Otherwise, retain control pending review. Distinguish evidence of harm, evidence ruling out a worthwhile gain, and an inconclusive result; “not significant” does not mean “no effect.”
Launch approval
Ravi signs off implementation and QA; Ines signs off measurement and the locked analysis plan; Chloe approves the isolated treatment, decision rule and any economic guardrail thresholds before launch.
Grades and run details
Decision model 83 · LLM judge 10 of 13 checks
Decision model checks
- passUses the supplied evidence correctly7%
- passAddresses the actual decision65%
- partialRespects explicit constraints21%
- passIdentifies material uncertainty45%
- passAvoids unsupported claims66%
- passProduces the required deliverable59%
- passTests one change at a time100%
- passFixes the duration up front77%
- passAn unambiguous primary metric84%
- partialDecision rule written before the test33%
- partialSized from the real traffic37%
- partialGuardrails with thresholds98%
Run
- Run
- #1
- API response time
- 36 s
- Submitted
- 30 Sept 2026