Needs more than a quick edit8 of 12 checks passed
The spec is strong but needs substantive edits to add guardrail thresholds, a complete decision rule including inconclusive and no-worse-than-0.5pp outcomes, and to remove unsupported current-system assumptions.
Output
Experiment: show delivery fees in the basket
Owner/build: Ravi · Analysis: Ines · Approval: Chloe Decision: Should Basketful show delivery charges before checkout?
Hypothesis and scope
Showing the fee in the basket will reduce fee surprises, but could discourage ordering. Unexpected delivery fees accounted for 12% of support contacts last quarter; this is motivation, not a baseline customer complaint rate.
Test fee visibility only. Do not add the proposed free-delivery promotional banner: it introduces a second change and would prevent us attributing results to earlier disclosure.
Experience
- Control: Current experience; delivery fee first appears at the final checkout step.
- Treatment: Show a clearly labelled delivery line beside the basket total: £3.99 delivery, or Free delivery when eligible. Update the fee and total as the basket changes.
- Use the existing pricing service and eligibility rules. Confirm behaviour at exactly £60 before launch; “over £60” must not silently become “£60 or more.”
- Leave checkout, pricing and other messaging unchanged.
Eligibility and assignment
Include customers viewing a non-empty, orderable basket. Exclude staff, bots and test accounts using rules fixed before launch.
Randomise 50/50 by customer, not session, at their first eligible basket view. Persist assignment across visits and devices where identity is available. For signed-out customers, use a stable first-party identifier; Ravi must document identity merging and cross-device limitations before launch.
Analyse everyone assigned, whether or not they start checkout or successfully see the fee: intention to treat.
Metrics
Primary: Customer order conversion — percentage of assigned customers placing at least one order within seven days of assignment. Report treatment minus control in percentage points, with a two-sided 95% confidence interval.
Checkout completion is diagnostic only: treatment may change who starts checkout, so orders ÷ checkout starts could improve while overall ordering falls.
Secondary/diagnostic: - Checkout-start rate and checkout completion. - Delivery-fee-related support contacts per assigned customer within seven days, using consistent tagging. - Orders and revenue per assigned customer; average order value and free-delivery eligibility as diagnostics.
Guardrails: Contribution margin per assigned customer, if reliably available; otherwise net revenue per assigned customer as an explicitly limited proxy. Also monitor basket/checkout errors. Chloe and Ines must approve numerical economic and reliability tolerances before launch.
Sample size and schedule
Chloe’s smallest worthwhile order-conversion change is 0.5 percentage points: 8.0% to 8.5%, not a 0.5% relative lift.
At 5% two-sided significance and 80% power, detecting that increase requires approximately 48,000 customers per arm, or 96,000 total.
Ines must first verify that the 8% baseline matches the eligibility rules and seven-day outcome window; recalculate and freeze the sample target if necessary.
Do not count returning customers again. Expected recruitment is approximately:
- Week 1: 40,000 unique customers.
- Each subsequent week: 20,000 additional customers.
- Four weeks: approximately 100,000 unique customers.
Enroll for at least four complete weeks and until the frozen sample target is reached, ending at a full-week boundary. Then wait seven days for the final cohort’s outcomes: roughly five weeks to readout.
Monitoring and decision
Ravi validates assignment persistence, fee accuracy, exposure/order linkage and sample-ratio balance before ramp-up.
Daily checks are for instrumentation and predefined safety breaches—not efficacy stopping. No “stop when significant,” sample extensions based on results, or repeated significance decisions.
Ines delivers one final primary analysis. Chloe’s default rollout rule is: - Estimated conversion gain at least +0.5 points; - 95% confidence interval excludes zero; - Guardrails pass.
Otherwise, do not roll out on this result alone. A non-significant result does not prove no harm. Secondary trust improvements inform follow-up, not a post-hoc replacement for the primary success rule.
Claims the judge couldn’t find in the brief
- The current experience includes a pricing service and eligibility rules that can be used to show the fee.
- Signed-out customers can be assigned using a stable first-party identifier, with identity merging and cross-device limitations to document.
- Delivery-fee-related support contacts can be measured per assigned customer using consistent tagging.
- Contribution margin or net revenue per assigned customer is available as a guardrail metric.
- Staff, bots, test accounts, non-empty orderable baskets, and basket/checkout errors can be identified and excluded or monitored.
Grades and run details
Decision model 88 · LLM judge 8 of 13 checks
Decision model checks
- passUses the supplied evidence correctly7%
- passAddresses the actual decision52%
- partialRespects explicit constraints24%
- passIdentifies material uncertainty73%
- passAvoids unsupported claims69%
- passProduces the required deliverable47%
- passTests one change at a time97%
- passFixes the duration up front82%
- passAn unambiguous primary metric76%
- partialDecision rule written before the test34%
- passSized from the real traffic59%
- partialGuardrails with thresholds99%
Run
- Run
- #1
- Time to output
- 47 s
- Submitted
- 1 Oct 2026