Needs more than a quick edit9 of 12 checks passed
The spec lacks a carry-over buffer at switch boundaries and a trust check for the primary metric, and it defers sample-size calculation to a simulation instead of sizing from the given traffic. A PM would need to add these before use.
Output
Brisk order-batching experiment spec
Decision required: Whether to enable batching in all 14 cities on 5 November 2026, ahead of the 6 November code freeze. Batching pairs two eligible orders from nearby restaurants for one courier.
Decision principles and sign-off
Launch only if the experiment shows a worthwhile courier-cost reduction, customer lateness stays within an agreed limit, and courier earnings protections are met. An inconclusive result is not a pass.
Before launch, the COO and Head of Operations must approve the customer-delay limit below; the courier relations lead and Legal must confirm the Porto protocol in writing. If those approvals or the schedule cannot be met, do not claim evidence for an all-city rollout before the freeze.
Design
Randomise city × two-hour operating windows, not individual orders. Orders interact through shared courier supply, and individual randomisation would contaminate the comparison. In each treatment window, the batching policy is enabled; in each control window, batching is disabled. Randomise and balance windows within each city, weekday and daypart, aiming for an even treatment/control split. Keep the assignment schedule fixed before the test begins.
An order is analysed according to its city-window at placement (intention to treat). Batch only orders placed in the same treatment window; do not pair across windows or with control orders. This limits cross-arm contamination at switchovers. Record actual batching eligibility and use, but do not reclassify orders based on whether they were batched.
Timeline and Porto condition
- 30 September: Send Porto’s couriers’ association written notice. The notice requirement applies to the experiment as well as any rollout. Start Porto treatment no earlier than 15 October, allowing 14 full days.
- By 7 October: Complete the per-city switch, monitoring, and rollback controls; test the assignment and cost calculations.
- 15 October–4 November: Run the experiment in all cities, subject to Porto approval.
- 5 November: Review results and, only if all launch gates pass, enable batching in all 14 cities in one day. Otherwise keep it off. Do not shorten the experiment to meet the freeze.
Porto must not enter treatment until Legal and the courier relations lead confirm the notice and trial protocol comply with the agreement. Agree a written earnings-protection mechanism with the association before starting; any top-ups count in courier cost. If this cannot be agreed, exclude Porto from treatment and do not describe the result as support for a 14-city rollout.
Measures and launch gates
Estimate treatment effects against control, with confidence intervals and pre-specified analysis. Baselines: courier cost €7.40/order, orders over 45 minutes 7.5%, courier earnings €13.20/active hour.
| Owner | Measure and proposed launch gate |
|---|---|
| CFO / COO — primary | All-in courier cost per completed order, including pay, distance payments, incentives and earnings-protection top-ups. Launch only if estimated savings are at least 5% and the 95% confidence interval supports a real saving (excludes zero savings). Report the interval; the point estimate alone is not proof. |
| Head of Operations — customer guardrail | Share delivered more than 45 minutes after ordering. Propose a maximum acceptable increase of 0.5 percentage points versus control; launch only if the 95% confidence interval rules out a larger increase. COO and Operations must approve this “noticeably longer” limit before the test. Also report delivery-time distributions, cancellations and complaints. |
| Courier relations — earnings guardrail | Earnings per active hour, including all payments and top-ups, with active time defined consistently across arms. No rollout if the overall result shows a material decline against a margin agreed in advance. In Porto, meet the association-approved earnings protection; do not treat a non-significant decline as proof of compliance. Report by city as well as overall. |
| Growth — secondary | 30-day reorder rate for orders exposed to each arm, analysed by assigned window. This outcome will mature after the November decision (roughly mid-December for the final cohort); report it then. It is not a pre-freeze launch gate. |
The proposed 0.5pp lateness margin and the earnings non-inferiority margin are policy choices, not facts in the brief; the accountable leads must sign them off before randomisation. Do not change margins after seeing results.
Analysis and power
The schedule contains 1,764 city-windows (14 cities × 6 windows/day × 21 days), before any exclusions. The analyst’s estimate of 740 assumes independent windows; successive windows may be correlated, and couriers and demand carry across time. Before the test, use historical city-window data and the planned randomisation to simulate power, accounting for serial correlation and city/daypart differences. Use blocked randomisation-based inference or an appropriate model with uncertainty at the randomised-window level. Do not claim the test is powered merely because 1,764 exceeds 740. If adjusted power is inadequate, or results are too imprecise to meet a gate by 5 November, the decision is inconclusive: no rollout before freeze.
Operations and rollback
Monitor treatment uptake, cost, lateness and earnings daily, but do not stop early for a favourable result. Pause a city for a safety, service or earnings breach under pre-agreed operational thresholds; document the incident and retain it in intention-to-treat analysis. Keep the city switch available for immediate rollback. Sign-off: COO, CFO, Head of Operations, courier relations lead; Growth acknowledges the delayed reorder readout.
Check by check
Got wrong · 2
- An unambiguous primary metricIt names courier cost per order as the primary metric but does not include a planned trust check such as a sample ratio check or verification that arms received the scheduled share of windows.
- Sized from the real trafficIt does not derive sample size from the 1.9 million monthly orders and the 5% effect; it only references the analyst's 740 windows and defers to a simulation without calculating required duration from real traffic.
Mixed · 1
- Avoids contamination between the armsIt explains the shared-courier problem and uses a switchback design, but does not handle carry-over at switches with a buffer or an assignment rule that excludes orders near the switch boundary.The two graders disagreed on this one.
Got right · 9
- Uses the supplied evidence correctlyAll claims about the current situation are taken directly from the brief or supplied context, with no invented facts.
- Addresses the actual decisionThe output commits to a clear decision rule (launch if gates pass, otherwise don't, inconclusive means no rollout) and states what would change it.
- Respects explicit constraintsThe spec respects the word limit, addresses each exec's needs, handles the Porto notice and earnings requirement, and fits the timeline before the freeze.
- Identifies material uncertaintyIt identifies correlation across windows, the need for power simulation, and that margins are policy choices, and says how each would be resolved.
- Avoids unsupported claimsCausal claims like contamination are presented as design rationale, not as established facts, and hypotheses are labelled with 'may'.
- Produces the required deliverableThe output is a complete experiment spec under 1,200 words, addressed to the execs, and usable as a decision document.
- A realistic plan that beats the freezeIt gives Porto notice on 30 Sep, starts the experiment on 15 Oct, runs for three weeks, and delivers a decision on 5 Nov before the 6 Nov freeze, with realistic acknowledgment of correlation.
- Decision rule written before the testEvery outcome (gates pass, guardrail breached, inconclusive) maps to a stated action with thresholds, all written before the test.
- Guardrails with thresholdsIt names late deliveries (0.5pp max increase) and courier earnings (no material decline) as guardrails with thresholds that would block rollout.
Grades and run details
Decision model 83 · LLM judge 9 of 13 checks
Decision model checks
- passUses the supplied evidence correctly17%
- passAddresses the actual decision56%
- partialRespects explicit constraints28%
- passIdentifies material uncertainty94%
- passAvoids unsupported claims61%
- partialProduces the required deliverable38%
- passAvoids contamination between the arms90%
- passA realistic plan that beats the freeze96%
- partialAn unambiguous primary metric57%
- passDecision rule written before the test91%
- partialSized from the real traffic65%
- passGuardrails with thresholds47%
Run
- Run
- #1
- API response time
- 44 s
- Submitted
- 30 Sept 2026