Usable with a quick edit12 of 12 checks passed
The spec is complete, evidence-based, and directly actionable; a PM could use it with only minor wording tweaks and no need to re-verify facts or redesign the experiment.
Output
Brisk order-batching experiment spec
Decision owner: Staff PM Sign-off: COO, CFO, Head of Operations, Courier Relations Decision deadline: 4 November 2026; rollout on 5 November, before the 6 November freeze.
1. Decision and hypothesis
Test whether batching can reduce total courier cost per delivered order by at least 5%, without materially worsening delivery times or reducing courier earnings.
At today’s €7.40 baseline, 5% means approximately €0.37 saved per order. Simulations are directional evidence, not sufficient grounds for rollout.
Decision: launch across all 14 cities only if the economic, customer, courier and operational gates below pass. Otherwise, leave batching off before the freeze. A promising but inconclusive result is not a pass.
2. Treatment and eligibility
Treatment enables one courier to carry at most two orders from nearby restaurants. Control retains current assignment.
Before testing, Operations and Engineering will lock:
- Restaurant-proximity, pickup-readiness and maximum-detour rules.
- Maximum predicted delivery times for both orders.
- Courier payment rules, including distance calculation.
- Exception handling, cancellation treatment and customer communications.
Do not change these rules during the confirmatory experiment. A material change requires a new test.
Measure the policy’s effect across all orders, not merely successfully batched orders. Report batching rate and first-/second-order outcomes as diagnostics; comparing batched orders with unbatched orders is selection-biased.
3. Experimental design
Replace order-level randomization with city-wide switchbacks. Treatment orders would otherwise change courier availability and dispatch conditions for control orders, contaminating the comparison.
Each city has six two-hour blocks daily between 11:00 and 23:00. Randomize batching on/off by city-block, with balanced assignment across cities, dates and time of day. Use constrained random schedules—not deterministic alternation.
Before the confirmatory test:
- Set a washout period using historical order/trip completion data and the pilot.
- Exclude the same initial portion of every block from measurement, whether or not its state changes.
- Validate that residual trips and courier repositioning do not materially contaminate measurement. If they do, lengthen blocks or washout and recalculate feasibility.
Attribute order outcomes to order-placement time; follow included orders through completion even after their block ends. Define payout and active-hour attribution consistently before launch.
Analysis is intention-to-treat by assigned city-block. Weight results to the expected rollout order mix, rather than weighting small and large cities equally. Report city-level results and heterogeneity alongside the aggregate.
4. Metrics and acceptance gates
The following customer thresholds are proposed definitions of “not noticeably longer” and require sign-off before testing.
| Stakeholder | Metric | Gate |
|---|---|---|
| CFO / COO | Total courier cost per delivered order, including distance pay, incentives and any earnings protection | One-sided 95% confidence bound supports at least 5% savings |
| Operations | Share delivered more than 45 minutes after ordering; baseline 7.5% | Upper confidence bound on increase ≤ 0.5 percentage points |
| COO / Operations | Mean and p90 order-to-delivery time | Upper confidence bounds on increases ≤ 1 minute mean, 2 minutes p90 |
| Courier Relations | Courier earnings per active hour; baseline €13.20 | Lower confidence bound on treatment effect ≥ €0/hour |
Define active hours identically in both arms, including available waiting time—not just time carrying orders. Report earnings distribution and effects by city and courier segment, not only the pooled mean.
Use baseline-adjusted treatment/control comparisons, not historical averages as the counterfactual. Log cancellations, failed deliveries, assignment failures and complaints as safety diagnostics; cost savings driven by worse completion are unacceptable.
Porto is a separate launch gate: pooled earnings gains cannot override a Porto earnings reduction. Legal and Courier Relations must approve the compliance interpretation and protection mechanism. Any guarantee or top-up must be part of the tested policy and counted in courier cost.
Growth: 30-day reorder is exploratory, not a pre-freeze gate. Even customers first exposed on 14 October do not mature until 13 November. Report available short-term indicators separately and agree a post-freeze retention study; do not present immature reorder data as a 30-day result.
5. Power and analysis
The intern’s “two weeks” is not justified. The analyst’s 740 windows assumes independence and addresses detecting a 5% effect—not necessarily proving savings exceed 5%.
From 14 October through 2 November, the maximum is:
14 cities × 6 blocks × 20 days = 1,680 blocks, before data losses.
That is only 2.27 times the nominal requirement. Serial correlation, washout, uneven volume and separate guardrail requirements can consume this headroom.
By 13 October, Data Science must replay the proposed schedule on historical/pilot data and establish:
- Power to pass the savings threshold under plausible true savings, including 6%, 8% and 11%.
- Power for delivery and earnings non-inferiority, especially Porto.
- Sensitivity to carryover, missing data and city differences.
Use inference respecting the randomized schedule, with serial-correlation and small-number-of-cities sensitivity checks. Pre-register estimators, confidence bounds, exclusions, city safety rules and any multiplicity adjustment for city-specific claims.
Feasibility gate: if adequate power cannot fit before the deadline, do not promise an all-city rollout. Extend testing beyond the freeze or return with a narrower decision.
6. Timeline and operating controls
- 30 September: obtain metric-threshold agreement; send Porto’s written assignment-change notice. Confirm legally valid delivery and when its 14-day period expires.
- By 7 October: deliver city switches, instrumentation, audit logs and rollback.
- 8–13 October: technical/safety pilot outside Porto; validate washout, payments and power. Pilot data excluded from confirmatory results.
- 14 October–2 November: confirmatory test, with Porto starting only after the full notice period has elapsed.
- 3–4 November: finalize outcomes, review compliance and make the signed decision.
- 5 November: one-day rollout only if all gates pass.
Operations owns live monitoring; Engineering owns rollback; Courier Relations owns courier communications and payment disputes. Stop affected cities immediately for payment errors, serious safety incidents or predefined severe service degradation. Such stops are recorded, not silently excluded.
No repeated efficacy peeking. If the test ends early for safety or loses its required sample, it does not automatically qualify for launch. Keep the kill switch available throughout peak season.
Check by check
Got right · 12
- Uses the supplied evidence correctlyAll factual claims about the current situation are directly from the brief or derived by arithmetic, with no invented numbers or facts.
- Addresses the actual decisionThe output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
- Respects explicit constraintsThe spec respects the word limit, addresses all named stakeholders, handles the Porto notice and earnings requirement, and fits the timeline before the code freeze.
- Identifies material uncertaintyIt identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
- Avoids unsupported claimsInterpretations and forecasts are clearly labelled as such (e.g., simulations as directional, serial correlation as a risk), and no confident claims go beyond the supplied evidence.
- Produces the required deliverableThe output is a complete experiment spec under 1,200 words, structured for the sign-off group, and contains all sections needed to act on it.
- Avoids contamination between the armsIt explicitly rejects order-level randomization due to shared couriers, adopts a city-block switchback design, and includes a washout period to prevent carry-over contamination.
- A realistic plan that beats the freezeIt converts 740 windows to about 9 days implicitly, explains that correlation and weekly cycles require a longer test (20 days, ~3 weeks), and provides dates that fit before 6 November with Porto notice given on 30 September.
- An unambiguous primary metricCourier cost per order is the unambiguous primary metric with a clear rationale (COO's 5% savings target), and trust checks like washout validation, diagnostics, and pre-registration are planned.
- Decision rule written before the testThe rule is stated upfront: launch if all gates pass, do not launch otherwise, and an inconclusive result is treated as a no-go.
- Sized from the real trafficThe duration uses the supplied traffic (14 cities, 6 blocks/day) and runs whole weeks (14 Oct–2 Nov), with a power analysis step to confirm adequacy against the 5% effect and correlation.
- Guardrails with thresholdsGuardrails are named (late deliveries, delivery time, courier earnings) with specific thresholds (0.5 pp, 1 min/2 min, €0/hour) that would block rollout.
Grades and run details
Decision model 92 · LLM judge 13 of 13 checks
Decision model checks
- passUses the supplied evidence correctly10%
- passAddresses the actual decision80%
- passRespects explicit constraints20%
- passIdentifies material uncertainty98%
- passAvoids unsupported claims62%
- passProduces the required deliverable60%
- passAvoids contamination between the arms100%
- passA realistic plan that beats the freeze76%
- partialAn unambiguous primary metric40%
- passDecision rule written before the test85%
- partialSized from the real traffic81%
- passGuardrails with thresholds89%
Run
- Run
- #1
- API response time
- 64 s
- Submitted
- 30 Sept 2026