Usable with a quick edit11 of 12 checks passed
The spec is complete and actionable; a quick edit to add a sample ratio check or similar trust check would make it fully robust.
Output
Order Batching Experiment Spec
Decision: Roll out batching to all 14 cities before the 6 Nov 2026 code freeze, or not. Owner: Staff PM. Sign-off: COO, CFO, Head of Operations, Courier Relations Lead. Head of Growth is copied in.
1. What changed from the intern's draft
| Draft | Problem | Fix |
|---|---|---|
| Randomise orders 50/50 within each city | Treated and control orders share the same couriers and the same pool of nearby orders. Halving the batchable pool understates batching, and control orders are affected by treatment (couriers are tied up on batches). Both effects bias the result. | Switch batching on and off by city and two-hour window (switchback). |
| Porto included from day one | The agreement requires 14 days' written notice of any change to order assignment, and a test is such a change. | Send notice today. Porto joins on 15 Oct. |
| Four metrics, no thresholds | Nobody can say what "good" means. | Pre-agreed decision rules (section 4). |
| 30-day reorder rate as a success metric | A test starting in October cannot produce a 30-day read before 6 Nov. | Monitor it, but it does not gate the decision. |
| Two weeks, no justification | See section 3. | Three full weeks. |
2. Design
- Unit: city × two-hour window. The windows are 11–13, 13–15, 15–17, 17–19, 19–21 and 21–23, so six per city per day.
- Assignment: Within each city-day, the windows form three adjacent pairs. In each pair, one window is randomly batching-on and the other is off. This balances lunch, afternoon and dinner across arms and controls for city and weekday.
- Attribution: An order belongs to the window in which it was placed. The pre-specified sensitivity analysis drops orders placed in the first 30 minutes of each window, because in-flight batches and courier positioning carry over from the previous window.
- Exposure: Batching pairs only nearby restaurants, using the same rules the simulation assumed. The batching configuration is frozen before the test starts.
3. Duration and power
The analyst's figure of about 740 windows assumes windows are independent. They are not: adjacent windows share couriers and demand, and cities differ. I therefore treat 740 as a floor and plan for up to double that.
- Engineering needs one week, so the switch is ready on 7 Oct and the test runs Thu 8 Oct – Wed 28 Oct (21 days, three of every weekday).
- 13 cities × 6 windows × 21 days = 1,638 windows. Porto adds 6 × 14 days = 84. The total is about 1,720 windows, roughly 2.3× the floor.
- We do not stop early on a good-looking interim result. The end date is fixed.
- Analysis: A window-level regression on total cost ÷ total orders, with city and weekday × hour fixed effects and standard errors clustered by city-day. We will confirm with randomisation inference.
4. Metrics and decision rules
All four sign-offs are needed on these thresholds by Tue 6 Oct, before any data exist.
| Owner | Metric | Role | Rule to ship |
|---|---|---|---|
| COO / CFO | Courier cost per order (today €7.40) | Primary | Estimated reduction is ≥5% and the 95% CI excludes zero |
| Head of Ops | Share of orders delivered >45 min after ordering (today 7.5%) | Guardrail | Increase of no more than +1.0 pp (upper 95% CI bound ≤ 8.5%) |
| Courier Relations | Courier earnings per active hour (today €13.20) | Guardrail | Lower 95% CI bound of the difference ≥ −€0.13 (about −1%) |
| Head of Growth | Reorder rate | Monitor only | Reported, does not gate |
The COO's "without making customers wait noticeably longer" is operationalised as the +1.0 pp cap. The simulations add 3–6 minutes to the second order in each batch, so this cap is the threshold most likely to be contested. The Head of Ops should confirm or change it before the start.
Tension to surface: Cost per order is forecast to fall 6–11%, and couriers are paid per order plus distance. Earnings per hour therefore stay flat only if batching raises orders per courier-hour by at least as much as pay per order falls. The test settles this, and the CFO's and Courier Relations' metrics are not independent.
Secondary metrics (descriptive only): - Mean delivery time, and delivery time for the second order in a batch. - Share of orders batched. - Orders per courier-hour. - Customer support contacts. - Cancellations.
Reorder rate: We report the 7-day rate for test customers at the decision. The 30-day rate is tracked after rollout for customers by batched-order exposure, and it cannot be treated as causal. Growth should know this now.
5. Porto
- Written notice goes to the couriers' association today (30 Sept). Legal should word it to cover both the test and permanent rollout. The earliest compliant start is 14 Oct, so Porto joins the test on 15 Oct.
- The contract says changes must not reduce average hourly earnings. Porto therefore has a stricter rule than other cities: the estimated earnings difference must be ≥ €0 for Porto to ship.
- Porto has only about 84 windows, so its own estimate will be noisy. We pool it for cost and lateness. For earnings we report Porto separately, daily, with the stop rule below.
- If Porto fails or is ambiguous, we ship in the other 13 cities and keep Porto off. Legal should confirm that running the test in Porto does not itself breach the agreement. If it might, exclude Porto from the test and roll out there only on the pooled evidence plus the earnings rule.
6. Stop rules
The per-city switch lets us turn batching off immediately. Any of these triggers a pause in that city and a review within 24 hours:
- The >45-minute share in batching-on windows exceeds off windows by more than 3 pp over two consecutive days.
- Earnings per active hour in batching-on windows are more than 5% below off windows over three consecutive days. In Porto, the threshold is any sustained shortfall that the Courier Relations Lead judges real.
- Any safety or serious courier-welfare incident tied to batching.
7. Timeline
| Date | Milestone |
|---|---|
| Wed 30 Sept | Spec circulated. Porto notice sent. Engineering starts the per-city switch. |
| Tue 6 Oct | Thresholds signed off. Randomisation schedule locked. |
| Wed 7 Oct | Switch ready, QA complete |
| Thu 8 Oct | Test starts in 13 cities |
| Thu 15 Oct | Porto joins |
| Wed 28 Oct | Test ends |
| Thu 29 Oct | Analysis and go/no-go with all four sign-offs |
| Fri 30 Oct | Rollout to approved cities (one day) |
| 30 Oct – 5 Nov | Post-rollout monitoring with the switch live. The freeze is on 6 Nov. |
The 29 Oct decision leaves a week of buffer. If the test slips past about 2 Nov, there is no safe rollout before the freeze.
8. Risks and limits
- Seasonality: October demand is below peak. Higher order density should raise batch rates, but it also strains courier supply. We track batch rate by hour and expect the peak effect to differ from the test estimate.
- Carryover: Some carryover between windows will remain even with the 30-minute washout. The sensitivity analysis shows how much it matters.
- Short-run effects: Couriers may change their behaviour as they learn batching, and three weeks may not capture that. We monitor weekly trends within the test.
- Possible outcomes: The result can be a full rollout, a partial rollout (for example excluding Porto or any city that fails a guardrail), or no rollout. Partial rollout by city is allowed by the switch, but only for cities that individually pass the stop rules. City-level results will be noisy, so we do not make ship decisions city by city beyond this.
9. Sign-off
| Role | Agrees to |
|---|---|
| COO | Decision rule, timeline |
| CFO | Cost metric definition (fully loaded courier pay ÷ delivered orders), ≥5% threshold |
| Head of Ops | +1.0 pp lateness cap, stop rules |
| Courier Relations | Earnings guardrail, Porto notice and rule |
Check by check
Got wrong · 1
- An unambiguous primary metricThe spec names courier cost per order as the primary metric with a rationale, but it does not include a planned trust check such as a sample ratio check or verification that the arms receive the scheduled share of windows and similar order volumes.
Got right · 11
- Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the brief or follow by arithmetic.
- Addresses the actual decisionThe spec commits to a conditional rollout decision with clear thresholds and states what would change the call.
- Respects explicit constraintsThe spec is under 1,200 words, respects the Porto notice and earnings requirement, and fits the timeline before the freeze.
- Identifies material uncertaintyIt names seasonality, carryover, courier adaptation, and Porto legal risk, and says how each would be resolved or monitored.
- Avoids unsupported claimsInterpretations and forecasts are clearly labelled as such; no factual claim is presented as established without support.
- Produces the required deliverableThe output is a complete experiment spec in the requested form, under the word limit, and usable by the sign-off group.
- Avoids contamination between the armsIt explains the shared-courier contamination problem and uses a switchback design with a 30-minute buffer to handle carry-over.
- A realistic plan that beats the freezeIt correctly treats 740 windows as a floor, accounts for dependence and weekly cycles with 21 days, and fits the build, Porto notice, and decision before 6 November.
- Decision rule written before the testThe decision rule is stated before the test: ship if cost reduction ≥5% with CI excluding zero and guardrails hold; otherwise no rollout, including the inconclusive case.
- Sized from the real trafficSample size is derived from the analyst's 740-window floor, increased for dependence, and the duration runs three full weeks using the actual traffic and operating hours.
- Guardrails with thresholdsGuardrails (late deliveries, courier earnings) are named with explicit thresholds that would block rollout.
Grades and run details
Decision model 83 · LLM judge 12 of 13 checks
Decision model checks
- passUses the supplied evidence correctly13%
- passAddresses the actual decision14%
- partialRespects explicit constraints27%
- passIdentifies material uncertainty98%
- passAvoids unsupported claims34%
- passProduces the required deliverable63%
- passAvoids contamination between the arms100%
- passA realistic plan that beats the freeze100%
- partialAn unambiguous primary metric48%
- partialDecision rule written before the test77%
- partialSized from the real traffic40%
- passGuardrails with thresholds97%
Run
- Run
- #1
- API response time
- 67 s
- Submitted
- 30 Sept 2026