Usable with a quick edit10 of 12 checks passed
The spec is comprehensive and only needs a quick edit to add a trust check and rephrase one unsupported claim; no re-checking of evidence or redoing of reasoning is required.
Output
Experiment spec: Order batching before the peak-season freeze
Decision owner: COO · Sign-off: COO, CFO, Head of Operations, Courier Relations Lead · Copied: Head of Growth Status: Draft for sign-off, 30 September 2026
1. Decision
Should we switch batching on in all 14 cities before the code freeze on 6 November 2026? The switch works per city, so the result can also be "on in some cities."
2. What changed from the draft plan, and why
| Draft | Problem | Fix |
|---|---|---|
| Randomise orders 50/50 | Treatment and control orders share the same couriers. Batching frees courier time, and that spare capacity speeds up control orders, which shrinks the measured difference. Earnings per courier hour can't be split by order at all. | Switchback test: randomise each city in 2-hour windows. Everything in a window is either on or off. |
| Runs in every city from day one | Porto needs 14 days' written notice. The earliest a notice sent today can take effect is 14 October. | Send notice today. Porto joins on 14 Oct. |
| "Late deliveries" | Undefined | Share of orders delivered >45 min after ordering (baseline 7.5%) |
| 30-day reorder rate as a success metric | It can't be measured before the freeze. In a switchback, each customer also sees both arms. | Monitor it after rollout. It is not a decision metric. |
| No thresholds or stop rules | We can't make a pass/fail call | Thresholds in §5, stop rules in §7 |
| Two weeks, assuming 740 independent windows | Windows are not independent. Batches straddle switch times, couriers reposition, and windows on the same city-day are correlated. | Exclude a washout period from each window, cluster the analysis, and run 21 days (~1,700 windows). |
3. Timeline
| Date | Milestone |
|---|---|
| Wed 30 Sep | Written notice to the Porto couriers' association. It covers both the test and a possible rollout. Legal to confirm wording today. |
| 30 Sep – 7 Oct | Engineering builds the per-city switch. Analytics builds the randomisation schedule and dashboard. |
| Thu 8 Oct | Test starts in 13 cities |
| Wed 14 Oct | Porto joins (notice period complete) |
| Wed 28 Oct | Test ends |
| 29 Oct – 1 Nov | Analysis |
| Mon 2 Nov | Decision meeting |
| Tue 3 Nov | Rollout (one day) |
| 4–5 Nov | Buffer for rollback or fixes. Freeze on 6 Nov. |
The dates are fixed by two constraints. The Porto notice cannot slip past today without losing Porto days. The test cannot run past 28 October without eating into the rollout buffer.
4. Design
- Unit: one city in one 2-hour window. Each day has six windows (11–13, 13–15, 15–17, 17–19, 19–21, 21–23).
- Assignment: randomised 50/50 within each city and time slot. Each city gets a balanced mix of on and off across lunch, dinner, weekdays and weekends.
- Switch behaviour: new batches form only while the switch is on. Batches already in progress when it turns off complete normally.
- Washout: the first 30 minutes of each window are excluded from analysis. This leaves 90 analysed minutes per window.
- Sample: 13 cities × 6 windows × 21 days plus Porto × 6 × 15 days gives 1,728 windows. That is about 2.3× the analyst's 740. The extra covers the washout and the correlation between windows. Analytics to confirm the power calculation using the correlation seen in the first week's data. If power is short, we extend to 30 October and give up one buffer day.
- Analysis: compare window-level metrics with adjustments for city, day and time slot. Treat all windows in the same city on the same day as one group when calculating uncertainty. Report the network result and each city's result.
5. Metrics and decision rule
| Metric | Owner | Baseline | Rule |
|---|---|---|---|
| Courier cost per order | CFO | €7.40 | Primary. Pass if the estimated reduction is ≥5% (≥€0.37) and the 95% confidence interval excludes zero. |
| Share of orders delivered >45 min | Head of Ops | 7.5% | Guardrail. Fail if the increase is more than 1.0 percentage point (to 8.5%). |
| Courier earnings per active hour | Courier Relations | €13.20 | Guardrail. Fail if the drop is more than €0.40 (3%). In Porto, any drop fails. |
| Median delivery time, second order in a batch | Head of Ops | n/a | Diagnostic. Simulation expects +3–6 min. |
| Complaints and refunds per 1,000 orders | Head of Ops | n/a | Diagnostic. This is our short-term stand-in for customer harm. |
| Share of orders batched | All | 0% | Diagnostic. Shows whether batching is actually happening. |
| 30-day reorder rate | Growth | n/a | Tracked after rollout. Not a decision metric. |
Rollout rule: a city rolls out on 3 Nov if the network result passes the primary metric and both guardrails, and that city's own result shows no clear guardrail failure.
The 1.0pp and €0.40 margins are proposals. Signers must confirm or change them before 8 October. They cannot be changed after the test starts.
6. Tensions signers should know about
What "at least 5%" can mean. We can reliably detect whether batching cuts cost by 5% versus not at all. We cannot prove the true effect is above 5% if it is only 6%, which is the low end of the simulation. The decision rule therefore uses the point estimate. If the estimate lands between 4% and 6%, the COO decides with that uncertainty in view.
Cost savings and courier earnings pull in opposite directions. Couriers are paid per order plus distance, so a lower cost per order is lower pay per order. Earnings per hour only hold up if couriers complete more orders per hour.
- During the test, the same couriers are logged in whether batching is on or off. That makes the test a fair early read on hourly earnings.
- After rollout, the saving only stays a saving if we need fewer courier hours. Fewer hours affects couriers' income, even if the hourly rate is protected.
- The CFO and Courier Relations Lead should agree on this before sign-off, not after results.
Porto. "Must not reduce average hourly earnings" is a contractual condition, not a statistical margin. Porto gives us only ~90 windows, too few to prove "no reduction."
- If network earnings per hour fall at all, Porto stays off unless it comes with an earnings protection mechanism. One example is a per-batch supplement, which legal and Courier Relations should scope now.
- If Porto rolls out, we measure its hourly earnings against the pre-batching baseline every week.
"Noticeably longer." We are defining this as the >45-minute late share plus the complaint rate. We are not using average delivery time, because a 3–6 minute delay on second orders is expected by design.
7. Stop rules (checked daily by the on-call analyst)
- City level: the >45-minute late share in "on" windows exceeds 10% for two consecutive days. Batching goes off in that city pending review.
- Network level:
- Any courier safety incident linked to batching, or a formal complaint from a couriers' association: pause everywhere.
- Earnings per hour in "on" windows below €12.50 for three consecutive days: pause everywhere.
- Porto only: any clear drop in hourly earnings leads to an immediate pause and notification to Courier Relations.
8. What each signer is agreeing to
- COO: the dates in §3, the rollout rule in §5, and making the call if the result falls between 4% and 6%.
- CFO: cost per order as the primary metric, with the ≥5% rule. Also acknowledges that post-rollout savings depend on how many courier hours we buy (§6).
- Head of Operations: the 1.0pp late-share margin and the stop rules.
- Courier Relations Lead:
- the €0.40 margin;
- the separate rule for Porto;
- sending the Porto notice today.
9. Open items
1. Legal, today: Porto notice wording, and whether one notice covers both the test and the rollout. 2. Analytics, by 7 Oct: - power calculation that accounts for correlated windows; - randomisation schedule; - dashboard. 3. Engineering, by 7 Oct: per-city switch with switch times logged, tested in one city after hours. 4. Courier Relations and Legal, by 2 Nov: Porto earnings protection option. 5. Signers, by 8 Oct: confirm the guardrail margins in §5.
Grades and run details
Decision model 88 · LLM judge 11 of 13 checks
Decision model checks
- partialUses the supplied evidence correctly8%
- passAddresses the actual decision43%
- partialRespects explicit constraints18%
- passIdentifies material uncertainty100%
- passAvoids unsupported claims22%
- passProduces the required deliverable56%
- passAvoids contamination between the arms100%
- passA realistic plan that beats the freeze100%
- passAn unambiguous primary metric33%
- partialDecision rule written before the test27%
- passSized from the real traffic41%
- passGuardrails with thresholds99%
Run
- Run
- #1
- Time to output
- 63 s
- Submitted
- 30 Sept 2026