Needs more than a quick edit9 of 12 checks passed
The spec is strong but would need material edits to the sizing and timeline, including converting the 740-window note, justifying a full three-week run, and adding explicit trust checks. A PM could not use it without reworking the experiment duration and power rationale.
Output
Brisk order-batching experiment: pre-freeze decision
Decision owner: Staff PM Required approvers: COO, CFO, Head of Operations, courier relations lead Decision deadline: 4 November 2026; rollout on 5 November, before the 6 November freeze.
1. Decision and hypothesis
Ship batching to all 14 cities only if it demonstrates at least 5% lower courier cost per order, with acceptable customer experience and protected courier earnings.
At today’s €7.40 baseline, 5% is approximately €0.37 per order. Simulated savings are not sufficient evidence.
The thresholds below are proposed sign-off terms, not established definitions of “noticeably longer.” Approvers must agree them before randomisation. An inconclusive result means no all-city rollout before the freeze, not permission to lower the bar.
2. Replace order randomisation with city switchbacks
Order-level randomisation is unsuitable: treatment and control orders share couriers, restaurant queues and dispatch capacity. Batching changes availability for both groups, contaminating the comparison.
Instead, switch each city between:
- Treatment: batching enabled under frozen restaurant-proximity, route-detour and eligibility rules; maximum two orders per courier.
- Control: current single-order dispatch.
Randomise six local-time blocks daily: 11–13, 13–15, …, 21–23. Balance treatment across cities, weekdays and clock times, with approximately equal exposure per city. Generate and lock the schedule in advance; do not let operators choose favourable treatment periods.
This estimates the effect of offering batching across a city, not merely its effect on successfully batched orders. Analyse all orders in the measurement cohort, including those never batched.
Carryover: Complete existing assignments safely at switches; never split an accepted batch. Set a fixed washout period using historical and shadow-mode assignment-to-completion times, before observing treatment results. Exclude order arrivals during that period from the primary cohort in every block, regardless of whether the state changed. Follow included orders through completion; do not truncate outcomes at block boundaries. If carryover cannot be contained within two-hour blocks, use longer blocks and re-power.
3. Metrics and launch gates
Use concurrent controls and pre-specified, one-sided 95% confidence bounds. Every gate must pass.
| Stakeholder | Metric | Launch requirement |
|---|---|---|
| CFO / COO | Total courier compensation attributable to the order cohort ÷ completed orders | Upper confidence bound on treatment/control ratio ≤0.95 |
| Operations | Share delivered more than 45 minutes after ordering; baseline 7.5% | Upper bound on increase ≤0.5 percentage points |
| COO / Operations | Mean order-to-door time | Upper bound on increase ≤2 minutes |
| Courier relations | Earnings per active hour; network baseline €13.20 | Lower bound on relative change ≥−2% outside Porto |
| Porto | Contract-defined average hourly earnings | No-reduction obligation met, as described below |
Cost includes distance pay, incentives, failed-delivery compensation and any batching-related earnings protection. Finance freezes allocation rules before launch. Track cancellation/completion rates so apparent savings cannot come from abandoning orders.
Define active hours as available-for-work plus engaged time, excluding genuine offline breaks—not just delivery time. Use consistent definitions across arms. Report both total earnings/total hours and courier-level distributions.
Report results by city, peak/off-peak, first/second batch position and courier earnings decile. Batch-position comparisons are descriptive, not randomised causal estimates. Examine late-delivery tails and cancellations; averages must not conceal material local harm.
Growth: Thirty-day reorder is exploratory, not a freeze gate. The first main-trial cohort matures on 14 November. Customer crossover also limits interpretation; do not present this switchback as a clean test of persistent customer retention.
4. Porto is a contractual prerequisite
Send written notice on 30 September, covering experimental on/off assignment changes and potential rollout. Legal and courier relations must confirm valid receipt, scope and the earliest permitted activation. Plan Porto activation for 15 October, providing a buffer beyond 14 days; late receipt moves that date.
Before activation, agree the contract’s hourly-earnings definition and comparison method with the association. The network’s €13.20 is not a Porto-specific contractual baseline.
Implement a funded make-whole mechanism accepted as satisfying the no-reduction requirement. Include those payments in experimental cost. A statistical non-inferiority margin cannot waive the agreement. Without compliant notice and earnings protection, Porto stays off—and the all-14-city decision is no-go.
5. Power, analysis and validity
The analyst’s 740 windows is not a sufficient sample-size justification:
- Adjacent windows share demand, weather, couriers and backlog.
- Washout reduces usable exposure.
- Detecting a 5% saving versus zero differs from proving a saving of at least 5%.
- Customer and earnings gates may require more data than cost.
Before launch, replay historical city-time data under the proposed schedule. Incorporate autocorrelation, city heterogeneity, washout and realistic take-up. Estimate power for every gate and the probability of passing them jointly, using simulation scenarios spanning 6–11% savings. Target at least 90% joint power under an explicitly agreed planning scenario.
Estimate intention-to-treat effects using city and calendar/clock-time adjustment, with fleet weights fixed from pre-experiment volume. Use schedule-respecting randomisation inference and dependence-aware uncertainty—not independent-window standard errors.
Publish city-level uncertainty. Pre-agree local safety thresholds and multiplicity handling; a pooled pass does not automatically authorise a city showing credible harm.
Feasibility gate: If adequate evidence cannot fit before the deadline, recommend no all-city pre-freeze launch.
6. Calendar and operating controls
- 30 September: Approve protocol and proposed margins; send Porto notice.
- 1–7 October: Build city switches, assignment logging, earnings protection and tested emergency rollback.
- 8–14 October: Shadow-mode validation; finalise washout, power, schedule and registered analysis.
- 15 October–2 November: Main experiment: 19 days, at most 1,596 nominal city-blocks before exclusions—not 1,596 independent observations.
- 3–4 November: Reconcile outcomes and compensation; analyse; obtain four signatures.
- 5 November: If approved, stage rollout across cities within the day, with monitored checkpoints.
Operations owns live monitoring. Pre-register severe-delay, cancellation, earnings and system-failure stop rules and a sequential-monitoring procedure. Emergency pauses remain permissible; log overrides and retain affected observations in assignment-based analysis. No opportunistic early declaration of success.
7. Sign-off package
Deliver effect estimates and confidence bounds, city diagnostics, notice/compliance evidence, incident logs, and Finance’s net savings forecast including protection payments and operating costs.
Go: Every statistical, contractual, local-safety and operational gate passes. No-go: Any gate fails, remains inconclusive, or cannot be evaluated reliably by 4 November. A subset rollout requires a separate, explicitly approved decision; it is not success against the all-city objective.
Grades and run details
Decision model 88 · LLM judge 9 of 13 checks
Decision model checks
- passUses the supplied evidence correctly10%
- passAddresses the actual decision82%
- partialRespects explicit constraints25%
- passIdentifies material uncertainty99%
- passAvoids unsupported claims64%
- passProduces the required deliverable44%
- passAvoids contamination between the arms100%
- passA realistic plan that beats the freeze99%
- passAn unambiguous primary metric27%
- passDecision rule written before the test98%
- failSized from the real traffic24%
- passGuardrails with thresholds89%
Run
- Run
- #1
- Time to output
- 81 s
- Submitted
- 1 Oct 2026