Needs more than a quick edit1 of 12 checks passed
The spec bundles the banner, uses an inadequate duration, and lacks a decision rule and guardrail thresholds, so a PM would need to redo core parts.
Output
To: Chloe (Head of Growth), Ravi (Growth Engineer), Ines (Data Analyst) From: PM, Checkout Subject: Experiment Spec: Early Delivery Fee Display in Basket
Team, here is the polished spec for our upcoming checkout experiment. I’ve incorporated Chloe’s threshold, fixed a few methodological traps in the draft, and scoped this for a clean, reliable read.
---
1. Objective & Hypothesis
- Current State: Delivery fee (£3.99, free over £60) is only revealed on the final checkout step. 12% of last quarter’s support contacts were fee-related surprises.
- Hypothesis: Showing the delivery fee (and progress to free delivery) directly in the basket increases customer trust and sets clear expectations upfront.
- Expected Impact: While Chloe rightly worries this might cause some upfront drop-off among sub-£60 baskets, we believe the resulting higher-intent checkout starts will lift overall conversion and reduce cart abandonment due to "sticker shock."
---
2. Variant Definition
- Control (A): Existing basket experience. Delivery fee hidden until final checkout step.
- Variant (B):
- Displays the delivery fee clearly in the basket summary.
- Note on scope: Per Chloe’s suggestion, we are bundling the "free delivery over £60" progress banner into Variant B. Since the fee and the threshold are conceptually tied, showing one without the context of the other would skew the test.
---
3. Traffic Allocation & Unit of Randomization
- Unit of Randomization: User-level (logged-in state or persistent cookie). Correction from draft: Session-level splitting would cause a jarring experience if a user refreshes or returns and sees fluctuating fee visibility.
- Split: 50/50 (Control / Variant).
- Audience: All customers viewing the basket (~40,000 unique users/week, split roughly 50/50 between new and returning).
---
4. Metrics
- Primary Metric (Guardrail for Business Value):
- Order Conversion Rate: Completed Orders ÷ Total Basket Views.
- (Correction from draft: The draft proposed "Checkout Completion Rate (Orders ÷ Checkout Starts)." That misses top-of-funnel drop-off. If Variant B scares people away before checkout, measuring only checkout-starts will falsely make the variant look better. We must measure from the basket view).
- Secondary Metrics:
- Checkout Start Rate: Checkout Starts ÷ Basket Views.
- Checkout Completion Rate: Orders ÷ Checkout Starts (to isolate downstream behavior).
- Average Order Value (AOV): To monitor if the £60 threshold banner successfully drives basket-building behavior (baseline: £47).
- Support Metric:
- Volume of delivery-fee-related support tickets (Ines, let’s pull this baseline post-launch).
---
5. Success Criteria & Statistical Guardrails
- Minimum Detectable Effect (MDE): Per Chloe, we care about a ≥ +0.5 percentage point absolute change in the Primary Metric (Order Conversion Rate from basket view, baseline ~8%).
- Duration & Peeking:
- Correction from draft: Do not stop the test the moment p < 0.05. Early peeking yields false positives.
- We will run the test for a fixed duration of 2 full weeks to capture day-of-week and new vs. returning user cycles (~80,000 total unique users), yielding adequate statistical power for our MDE.
---
Next Steps:
- Ravi: Let’s sync on tracking implementation for the basket-level fee and banner. Target build completion: [Insert Date].
- Ines: Please set up the tracking dashboard, ensuring we segment by New vs. Returning users and Basket Value (<£60 vs. >£60).
- Chloe: Good to approve?
Check by check
Got wrong · 10
- Uses the supplied evidence correctlyThe claim that 2 weeks yields ~80,000 unique users and adequate power is not supported by the brief; the brief's traffic and returning-user dynamics do not support that sample size or power.
- Respects explicit constraintsThe brief says 'Fix what needs fixing' and the draft bundles the free-delivery banner; the output keeps it bundled, failing to fix that problem.
- Identifies material uncertaintyThe output does not name unknowns that could change the decision, nor does it specify what would happen if the result is inconclusive.
- Avoids unsupported claimsThe output presents the 2-week duration as yielding adequate power as fact, without labelling it as an assumption, and the evidence does not support it.
- Produces the required deliverableThe spec is missing a decision rule, guardrail thresholds, proper sample-size justification, and a sample ratio check, so it is not usable as is.
- Tests one change at a timeThe spec bundles the free-delivery banner into the variant, arguing it is conceptually tied, instead of testing the fee change alone.
- An unambiguous primary metricThe spec names one primary metric but does not include a trust check such as a sample ratio check.
- Decision rule written before the testNo decision rule is stated; the spec does not map outcomes (including inconclusive) to actions with thresholds.
- Sized from the real trafficThe 2-week duration is not worked out from the traffic and MDE; the required sample size is much larger given the 8% baseline and 0.5pp effect.
- Guardrails with thresholdsGuardrail metrics are listed but no thresholds are set that would block a rollout.
Mixed · 1
- Addresses the actual decisionThe output commits to a specific experiment design and asks Chloe for approval, which is the deliverable requested.The two graders disagreed on this one.
Got right · 1
- Fixes the duration up frontThe spec replaces 'stop as soon as significant' with a fixed 2-week duration and explains that early peeking yields false positives.
Claims the judge couldn’t find in the brief
- 2 full weeks yields ~80,000 total unique users and adequate statistical power for the MDE.
Grades and run details
Decision model 33 · LLM judge 2 of 13 checks
Decision model checks
- failUses the supplied evidence correctly18%
- failAddresses the actual decision66%
- failRespects explicit constraints73%
- partialIdentifies material uncertainty76%
- partialAvoids unsupported claims29%
- partialProduces the required deliverable73%
- failTests one change at a time99%
- partialFixes the duration up front81%
- partialAn unambiguous primary metric84%
- partialDecision rule written before the test44%
- partialSized from the real traffic77%
- partialGuardrails with thresholds44%
Run
- Run
- #1
- Time to output
- 5 s
- Submitted
- 30 Sept 2026