Needs more than a quick edit9 of 12 checks passed
The spec lacks a carry-over buffer at switch boundaries and a trust check for the primary metric, and it defers sample-size calculation to a simulation instead of sizing from the given traffic. A PM would need to add these before use.
Output
Brisk order-batching experiment spec
Decision required: Whether to enable batching in all 14 cities on 5 November 2026, ahead of the 6 November code freeze. Batching pairs two eligible orders from nearby restaurants for one courier.
Decision principles and sign-off
Launch only if the experiment shows a worthwhile courier-cost reduction, customer lateness stays within an agreed limit, and courier earnings protections are met. An inconclusive result is not a pass.
Before launch, the COO and Head of Operations must approve the customer-delay limit below; the courier relations lead and Legal must confirm the Porto protocol in writing. If those approvals or the schedule cannot be met, do not claim evidence for an all-city rollout before the freeze.
Design
Randomise city × two-hour operating windows, not individual orders. Orders interact through shared courier supply, and individual randomisation would contaminate the comparison. In each treatment window, the batching policy is enabled; in each control window, batching is disabled. Randomise and balance windows within each city, weekday and daypart, aiming for an even treatment/control split. Keep the assignment schedule fixed before the test begins.
An order is analysed according to its city-window at placement (intention to treat). Batch only orders placed in the same treatment window; do not pair across windows or with control orders. This limits cross-arm contamination at switchovers. Record actual batching eligibility and use, but do not reclassify orders based on whether they were batched.
Timeline and Porto condition
- 30 September: Send Porto’s couriers’ association written notice. The notice requirement applies to the experiment as well as any rollout. Start Porto treatment no earlier than 15 October, allowing 14 full days.
- By 7 October: Complete the per-city switch, monitoring, and rollback controls; test the assignment and cost calculations.
- 15 October–4 November: Run the experiment in all cities, subject to Porto approval.
- 5 November: Review results and, only if all launch gates pass, enable batching in all 14 cities in one day. Otherwise keep it off. Do not shorten the experiment to meet the freeze.
Porto must not enter treatment until Legal and the courier relations lead confirm the notice and trial protocol comply with the agreement. Agree a written earnings-protection mechanism with the association before starting; any top-ups count in courier cost. If this cannot be agreed, exclude Porto from treatment and do not describe the result as support for a 14-city rollout.
Measures and launch gates
Estimate treatment effects against control, with confidence intervals and pre-specified analysis. Baselines: courier cost €7.40/order, orders over 45 minutes 7.5%, courier earnings €13.20/active hour.
| Owner | Measure and proposed launch gate |
|---|---|
| CFO / COO — primary | All-in courier cost per completed order, including pay, distance payments, incentives and earnings-protection top-ups. Launch only if estimated savings are at least 5% and the 95% confidence interval supports a real saving (excludes zero savings). Report the interval; the point estimate alone is not proof. |
| Head of Operations — customer guardrail | Share delivered more than 45 minutes after ordering. Propose a maximum acceptable increase of 0.5 percentage points versus control; launch only if the 95% confidence interval rules out a larger increase. COO and Operations must approve this “noticeably longer” limit before the test. Also report delivery-time distributions, cancellations and complaints. |
| Courier relations — earnings guardrail | Earnings per active hour, including all payments and top-ups, with active time defined consistently across arms. No rollout if the overall result shows a material decline against a margin agreed in advance. In Porto, meet the association-approved earnings protection; do not treat a non-significant decline as proof of compliance. Report by city as well as overall. |
| Growth — secondary | 30-day reorder rate for orders exposed to each arm, analysed by assigned window. This outcome will mature after the November decision (roughly mid-December for the final cohort); report it then. It is not a pre-freeze launch gate. |
The proposed 0.5pp lateness margin and the earnings non-inferiority margin are policy choices, not facts in the brief; the accountable leads must sign them off before randomisation. Do not change margins after seeing results.
Analysis and power
The schedule contains 1,764 city-windows (14 cities × 6 windows/day × 21 days), before any exclusions. The analyst’s estimate of 740 assumes independent windows; successive windows may be correlated, and couriers and demand carry across time. Before the test, use historical city-window data and the planned randomisation to simulate power, accounting for serial correlation and city/daypart differences. Use blocked randomisation-based inference or an appropriate model with uncertainty at the randomised-window level. Do not claim the test is powered merely because 1,764 exceeds 740. If adjusted power is inadequate, or results are too imprecise to meet a gate by 5 November, the decision is inconclusive: no rollout before freeze.
Operations and rollback
Monitor treatment uptake, cost, lateness and earnings daily, but do not stop early for a favourable result. Pause a city for a safety, service or earnings breach under pre-agreed operational thresholds; document the incident and retain it in intention-to-treat analysis. Keep the city switch available for immediate rollback. Sign-off: COO, CFO, Head of Operations, courier relations lead; Growth acknowledges the delayed reorder readout.
Grades and run details
Decision model 83 · LLM judge 9 of 13 checks
Decision model checks
- passUses the supplied evidence correctly13%
- passAddresses the actual decision55%
- partialRespects explicit constraints19%
- passIdentifies material uncertainty93%
- passAvoids unsupported claims60%
- partialProduces the required deliverable34%
- passAvoids contamination between the arms87%
- passA realistic plan that beats the freeze93%
- partialAn unambiguous primary metric36%
- passDecision rule written before the test92%
- partialSized from the real traffic18%
- passGuardrails with thresholds49%
Run
- Run
- #1
- API response time
- 44 s
- Submitted
- 30 Sept 2026