Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Brisk. We want to know whether to roll out order batching (one courier carrying two orders from nearby restaurants) to all 14 of our cities before the peak-season code freeze on 6 November 2026. Write the experiment spec. Our COO, CFO, Head of Operations and courier relations lead will all sign it off, and each wants something different from it. Keep it under 1,200 words. A draft plan from our data science intern is below. Fix what needs fixing.

What the model was given7 items: About Brisk, What batching should do, What each exec wants to see, Courier agreement in Porto, Timeline, The intern's draft plan, Analyst's note
About BriskFood delivery in 14 European cities, about 1.9 million orders a month. Couriers are paid per order plus distance. Courier cost per order averages €7.40.
What batching should doCOO: “Roll batching out before the freeze if it cuts courier cost per order by at least 5% without making customers wait noticeably longer.” Simulations suggest it cuts courier cost per order by 6–11% and adds 3–6 minutes to the second order in each batch.
What each exec wants to seeCFO: courier cost per order. Head of Operations: the share of orders delivered more than 45 minutes after ordering (today 7.5%). Courier relations lead: courier earnings per active hour (today €13.20). Head of Growth, copied in: 30-day reorder rate.
Courier agreement in PortoOur agreement with the Porto couriers' association requires 14 days' written notice of any change to how orders are assigned, and says changes must not reduce couriers' average hourly earnings.
TimelineToday is 30 September. Engineering needs one week to put batching behind a switch that can be turned on and off per city at any time. Rollout to all cities takes a day. Operating hours are 11:00 to 23:00 in every city.
The intern's draft planRandomise orders 50/50 in every city: orders in the treatment group can be batched, control orders never are. Success metrics: courier cost per order, late deliveries, courier earnings per hour and 30-day reorder rate. Run for two weeks.
Analyst's noteIf we switch batching on and off by city in two-hour windows, detecting a 5% change in courier cost per order needs about 740 windows in total, assuming each window is independent of the others.
What a strong answer doesThe answer key the graders mark against

A spec that sees the draft can't work: batched and unbatched orders in the same city share one pool of couriers, so batching in the treatment group frees couriers for the control group and contaminates the comparison. It uses a switchback design instead (batching on or off by city in randomised two-hour windows), with a buffer at each switch (for example, orders placed shortly before a switch are excluded or assigned by dispatch time) so carry-over doesn't blur the arms. It picks courier cost per order as the one primary metric, because cutting it is batching's purpose and a fall is unambiguously good, and makes the others guardrails with thresholds: late deliveries (for example, no more than a 1pp rise from 7.5%), courier earnings per active hour (no fall, which the Porto agreement requires), with 30-day reorder rate named as a follow-up because it can't be read before the freeze. It sizes the test honestly: 14 cities × 6 windows a day is 84 windows, so 740 windows is about 9 days, but windows in the same city on the same day aren't independent, so it needs more, and a full weekly cycle plus time for couriers to adapt argue for about three weeks. It fits the timeline: build ready around 7 October, Porto given notice now and joining once its 14 days are up (or excluded), about three weeks of data, a read-out around 1 November and rollout before 6 November. It plans trust checks (the arms get the scheduled share of windows and similar order volumes; batching actually happens in 'on' windows), and states the decision rule up front: roll out if cost per order falls at least 5% and guardrails hold; don't if a guardrail is breached; and what happens if the result is inconclusive by the freeze.

Critical failures (cap the score at 40)
  • Randomises by order or by customer within a city, where both arms share the same couriers
  • Changes order assignment in Porto without the 14 days' notice
Case

Staff-level · v1.1 · synthetic · marketplace, logistics, switchback, network effects, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit5 of 12 checks passed

The timeline and sizing need rework to account for clustered windows and adaptation, a trust check is missing, and the decision rule does not cover inconclusive outcomes. A PM would have to redo the power analysis and add missing elements.

Output

Experiment Spec: Order Batching Rollout Evaluation

Author: Staff Product Manager, Core Logistics Target Decision Date: 4 November 2026 (Rollout: 5 November; Peak Code Freeze: 6 November) Stakeholders / Sign-offs: COO, CFO, Head of Operations, Courier Relations Lead Informed: Head of Growth

---

1. Executive Summary & Decision Framework

We are evaluating courier order batching (one courier delivering two proximate orders) across our 14 European markets (~1.9M monthly orders). Our goal is to determine whether to ship batching globally prior to the 6 November code freeze.

Go / No-Go Decision Criteria (Sign-off Contract)

Batching will roll out to all 14 cities on 5 November only if all four conditions are met: 1. Cost Efficiency (CFO & COO): Courier cost per order decreases by $\ge$ 5.0% (baseline: €7.40; target: $\le$ €7.03) at $p < 0.05$. 2. Customer Experience Guardrail (COO & Head of Ops): Share of orders delivered in >45 minutes increases by no more than 0.8 percentage points (baseline: 7.5%; ceiling: $\le$ 8.3%) with no statistically significant drop in order ratings. 3. Courier Welfare & Legal Compliance (Courier Relations & Porto Agreement): Courier earnings per active hour do not decline in any market (baseline: €13.20; target: $\ge$ €13.20, non-inferiority margin of 0%). 4. Leading Retention Indicator (Head of Growth): 7-day reorder rate exhibits no statistically significant negative degradation ($\alpha = 0.05$).

---

2. Flaws in the Intern’s Draft & Required Fixes

Draft ProposalCore FlawResolution
:---:---:---
Order-level 50/50 splitSevere Market Interference (SUTVA violation): Batched treatment orders consume courier capacity unevenly, skewing dispatch speed, courier transit times, and ETAs for control orders sharing the same courier pool.City-Time Switchback Design: Randomize batching policy across all couriers/orders within discrete 2-hour time blocks per city.
30-day reorder rate metricTimeline Impossibility: Testing in October for a 6 November release leaves insufficient time to observe a 30-day reorder window.Leading Indicator Proxies: Measure 7-day reorder rate, delivery ratings (CSAT), and support ticket contact rates during the test; monitor 30-day reorder post-launch.
Porto immediate testingLegal Breach: Porto couriers' agreement requires 14 days' advance written notice and a guarantee of non-negative impact on average hourly earnings.Targeted Notice & Phased Inclusion: Issue formal notice on 1 October. Exclude Porto from Week 1 testing; enter Porto into switchback on 15 October with an earnings-floor guarantee.
Unstratified durationFails to control for intra-day/intra-week demand cyclicality (lunch vs. dinner, weekday vs. weekend).Stratified Block Randomization: Balance treatment and control windows across peak (12:00–14:00, 19:00–21:00) and off-peak windows.

---

3. Experiment Design & Methodology

Design: Stratified City-Time Switchback

  • Unit of Randomization: City $\times$ 2-hour window (11:00–13:00, 13:00–15:00, ..., 21:00–23:00 = 6 windows/day per city).
  • Intervention:
  • Treatment Window: Batching algorithm active (up to 2 orders per courier from nearby pickup points).
  • Control Window: Pure single-order dispatch (status quo).
  • Buffer/Washout: 15-minute dispatch transition buffer between blocks. Orders placed in the final 15 minutes of a window cannot be batched if the subsequent window is Control, eliminating spillover across windows.

Power & Sample Size

  • Required Sample: 740 two-hour windows to detect a 5% shift in cost per order ($\alpha = 0.05, \beta = 0.80$).
  • Capacity:
  • 13 cities $\times$ 6 windows/day = 78 windows/day.
  • Porto joins on Day 8 (15 October): 14 cities $\times$ 6 windows/day = 84 windows/day.
  • Duration: 14 days (8 October – 21 October).
  • 7 days $\times$ 78 windows + 7 days $\times$ 84 windows = 1,134 windows.
  • This provides >150% of required sample, allowing high statistical power to detect sub-1% changes in courier hourly earnings and Ops tail-delays.

---

4. Stakeholder Metrics & Success Criteria

``` +---------------------------------------------------------------------------------------+ | PRIMARY METRIC | | Courier Cost per Order (CFO/COO) | Target: >= 5% reduction (<= €7.03 vs. €7.40) | +---------------------------------------------------------------------------------------+ | +---------------------------------+---------------------------------+ | | | +------------------+ +-------------------+ +--------------------+ | OPS GUARDRAIL | | COURIER GUARDRAIL | | GROWTH GUARDRAILS | | (Head of Ops) | | (Courier Lead) | | (Head of Growth) | | >45m Late Share | | Courier Earnings | | • 7-day Reorder | | Baseline: 7.5% | | per Active Hour | | • Order Rating | | Ceiling: <= 8.3% | | Baseline: €13.20 | | • Missing/Cold Food| | Second-order ETA | | Floor: >= €13.20 | | Refund Rate | | impact: <= +5m | | (Non-inferior) | | | +------------------+ +-------------------+ +--------------------+ ```

Metric Definitions & Instruments

1. Courier Cost per Order (CFO): Total courier payout (base + distance fees + batch bonuses) divided by completed orders within the window. 2. Late Delivery Share >45m (Head of Ops): P90 delivery time and proportion of orders delivered >45 minutes from placement. Simulation projects a 3–6 min increase for the 2nd order; total late share must not breach 8.3%. 3. Courier Earnings per Active Hour (Courier Relations): Gross earnings divided by active hours (online and on-trip/assigned). Risk: Because couriers are paid per order + distance, batching two orders saves distance, potentially reducing payout unless batch base rates or drop-off multipliers offset it. 4. Customer Experience & Retention Proxy (Growth): * 7-day reorder rate for cohorts exposed to $\ge 2$ batched deliveries. * Delivery-related customer support contacts per 1,000 orders.

---

5. Porto Compliance & Courier Safeguards

Our Porto agreement mandates a 14-day notice period for dispatch adjustments and guarantees that courier average hourly earnings will not fall.

  • Timeline Action: Formal written notice delivered to the Porto Couriers’ Association on Wednesday, 1 October 2026.
  • Phased Rollout: Porto remains 100% Control through 14 October. On 15 October, Porto enters the switchback rotation.
  • Earnings Floor Mechanism: If average courier earnings per active hour in treatment windows drop below €13.20 in Porto (or any city), an automatic “Batch Adjustment Top-Up” will credit couriers the difference at end-of-day. This eliminates legal exposure while allowing us to evaluate the raw algorithm’s unit economics.

---

6. Execution Timeline

DateMilestoneOwner
:---:---:---
30 SepSpec review and executive alignment.Staff PM
1 OctSend 14-day formal notice to Porto couriers' association.Courier Relations
1–7 OctEngineering builds city-level switchback toggle + 15m buffer logic.Engineering
8 OctExperiment Launch: 13 cities enter switchback (Porto held in control).Ops / Analytics
15 OctPorto enters switchback upon notice expiry.Ops
21 Oct (23:00)Experiment Concludes (1,134 windows logged).Analytics
22–27 OctCollect 7-day reorder data for final test cohorts.Data Science
28 Oct–2 NovDeep-dive analysis: cost, earnings, ETA tail, customer sentiment.Analytics / PM
3 NovFinal Sign-off Meeting with COO, CFO, Ops, Courier Lead.Staff PM
5 NovFull rollout across all 14 cities (1 day).Engineering
6 NovPeak-Season Code Freeze takes effect.Platform Eng

---

7. Circuit Breakers & Rollback Criteria

The experiment toggle will automatically shut down batching within 15 minutes if any of the following triggers trip during live operations: * Severe Operations Failure: City-level late delivery rate (>45 min) exceeds 12.0% across two consecutive peak windows. * Courier Backlash/Disruption: Courier unassigned order rate spikes by $>25\%$ or courier offline churn spikes by $>15\%$ compared to historical control. * Severe Food Quality Impact: Customer support contact rate for "cold food" or "late arrival" doubles ($>2.0\times$).

---

8. Stakeholder Sign-Off

By signing below, stakeholders agree that if the primary and guardrail criteria in Section 1 are satisfied by 3 November, order batching will be deployed to all 14 cities on 5 November 2026.

  • Chief Operating Officer: _________
  • Chief Financial Officer: _________
  • Head of Operations: _________
  • Courier Relations Lead: _________

Grades and run details

Decision model 54 · LLM judge 7 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly35%
  • passAddresses the actual decision38%
  • failRespects explicit constraints37%
  • partialIdentifies material uncertainty34%
  • partialAvoids unsupported claims44%
  • partialProduces the required deliverable30%
  • passAvoids contamination between the arms100%
  • partialA realistic plan that beats the freeze87%
  • partialAn unambiguous primary metric40%
  • partialDecision rule written before the test98%
  • partialSized from the real traffic58%
  • passGuardrails with thresholds83%
Run
Run
#1
API response time
37 s
Submitted
1 Oct 2026

Usable with a quick edit10 of 12 checks passed

The spec is complete and actionable; a quick edit to add a sample ratio check or similar trust check would make it fully robust.

Output

Order Batching Experiment Spec

Decision: Roll out batching to all 14 cities before the 6 Nov 2026 code freeze, or not. Owner: Staff PM. Sign-off: COO, CFO, Head of Operations, Courier Relations Lead. Head of Growth is copied in.

1. What changed from the intern's draft

DraftProblemFix
Randomise orders 50/50 within each cityTreated and control orders share the same couriers and the same pool of nearby orders. Halving the batchable pool understates batching, and control orders are affected by treatment (couriers are tied up on batches). Both effects bias the result.Switch batching on and off by city and two-hour window (switchback).
Porto included from day oneThe agreement requires 14 days' written notice of any change to order assignment, and a test is such a change.Send notice today. Porto joins on 15 Oct.
Four metrics, no thresholdsNobody can say what "good" means.Pre-agreed decision rules (section 4).
30-day reorder rate as a success metricA test starting in October cannot produce a 30-day read before 6 Nov.Monitor it, but it does not gate the decision.
Two weeks, no justificationSee section 3.Three full weeks.

2. Design

  • Unit: city × two-hour window. The windows are 11–13, 13–15, 15–17, 17–19, 19–21 and 21–23, so six per city per day.
  • Assignment: Within each city-day, the windows form three adjacent pairs. In each pair, one window is randomly batching-on and the other is off. This balances lunch, afternoon and dinner across arms and controls for city and weekday.
  • Attribution: An order belongs to the window in which it was placed. The pre-specified sensitivity analysis drops orders placed in the first 30 minutes of each window, because in-flight batches and courier positioning carry over from the previous window.
  • Exposure: Batching pairs only nearby restaurants, using the same rules the simulation assumed. The batching configuration is frozen before the test starts.

3. Duration and power

The analyst's figure of about 740 windows assumes windows are independent. They are not: adjacent windows share couriers and demand, and cities differ. I therefore treat 740 as a floor and plan for up to double that.

  • Engineering needs one week, so the switch is ready on 7 Oct and the test runs Thu 8 Oct – Wed 28 Oct (21 days, three of every weekday).
  • 13 cities × 6 windows × 21 days = 1,638 windows. Porto adds 6 × 14 days = 84. The total is about 1,720 windows, roughly 2.3× the floor.
  • We do not stop early on a good-looking interim result. The end date is fixed.
  • Analysis: A window-level regression on total cost ÷ total orders, with city and weekday × hour fixed effects and standard errors clustered by city-day. We will confirm with randomisation inference.

4. Metrics and decision rules

All four sign-offs are needed on these thresholds by Tue 6 Oct, before any data exist.

OwnerMetricRoleRule to ship
COO / CFOCourier cost per order (today €7.40)PrimaryEstimated reduction is ≥5% and the 95% CI excludes zero
Head of OpsShare of orders delivered >45 min after ordering (today 7.5%)GuardrailIncrease of no more than +1.0 pp (upper 95% CI bound ≤ 8.5%)
Courier RelationsCourier earnings per active hour (today €13.20)GuardrailLower 95% CI bound of the difference ≥ −€0.13 (about −1%)
Head of GrowthReorder rateMonitor onlyReported, does not gate

The COO's "without making customers wait noticeably longer" is operationalised as the +1.0 pp cap. The simulations add 3–6 minutes to the second order in each batch, so this cap is the threshold most likely to be contested. The Head of Ops should confirm or change it before the start.

Tension to surface: Cost per order is forecast to fall 6–11%, and couriers are paid per order plus distance. Earnings per hour therefore stay flat only if batching raises orders per courier-hour by at least as much as pay per order falls. The test settles this, and the CFO's and Courier Relations' metrics are not independent.

Secondary metrics (descriptive only): - Mean delivery time, and delivery time for the second order in a batch. - Share of orders batched. - Orders per courier-hour. - Customer support contacts. - Cancellations.

Reorder rate: We report the 7-day rate for test customers at the decision. The 30-day rate is tracked after rollout for customers by batched-order exposure, and it cannot be treated as causal. Growth should know this now.

5. Porto

  • Written notice goes to the couriers' association today (30 Sept). Legal should word it to cover both the test and permanent rollout. The earliest compliant start is 14 Oct, so Porto joins the test on 15 Oct.
  • The contract says changes must not reduce average hourly earnings. Porto therefore has a stricter rule than other cities: the estimated earnings difference must be ≥ €0 for Porto to ship.
  • Porto has only about 84 windows, so its own estimate will be noisy. We pool it for cost and lateness. For earnings we report Porto separately, daily, with the stop rule below.
  • If Porto fails or is ambiguous, we ship in the other 13 cities and keep Porto off. Legal should confirm that running the test in Porto does not itself breach the agreement. If it might, exclude Porto from the test and roll out there only on the pooled evidence plus the earnings rule.

6. Stop rules

The per-city switch lets us turn batching off immediately. Any of these triggers a pause in that city and a review within 24 hours:

  • The >45-minute share in batching-on windows exceeds off windows by more than 3 pp over two consecutive days.
  • Earnings per active hour in batching-on windows are more than 5% below off windows over three consecutive days. In Porto, the threshold is any sustained shortfall that the Courier Relations Lead judges real.
  • Any safety or serious courier-welfare incident tied to batching.

7. Timeline

DateMilestone
Wed 30 SeptSpec circulated. Porto notice sent. Engineering starts the per-city switch.
Tue 6 OctThresholds signed off. Randomisation schedule locked.
Wed 7 OctSwitch ready, QA complete
Thu 8 OctTest starts in 13 cities
Thu 15 OctPorto joins
Wed 28 OctTest ends
Thu 29 OctAnalysis and go/no-go with all four sign-offs
Fri 30 OctRollout to approved cities (one day)
30 Oct – 5 NovPost-rollout monitoring with the switch live. The freeze is on 6 Nov.

The 29 Oct decision leaves a week of buffer. If the test slips past about 2 Nov, there is no safe rollout before the freeze.

8. Risks and limits

  • Seasonality: October demand is below peak. Higher order density should raise batch rates, but it also strains courier supply. We track batch rate by hour and expect the peak effect to differ from the test estimate.
  • Carryover: Some carryover between windows will remain even with the 30-minute washout. The sensitivity analysis shows how much it matters.
  • Short-run effects: Couriers may change their behaviour as they learn batching, and three weeks may not capture that. We monitor weekly trends within the test.
  • Possible outcomes: The result can be a full rollout, a partial rollout (for example excluding Porto or any city that fails a guardrail), or no rollout. Partial rollout by city is allowed by the switch, but only for cities that individually pass the stop rules. City-level results will be noisy, so we do not make ship decisions city by city beyond this.

9. Sign-off

RoleAgrees to
COODecision rule, timeline
CFOCost metric definition (fully loaded courier pay ÷ delivered orders), ≥5% threshold
Head of Ops+1.0 pp lateness cap, stop rules
Courier RelationsEarnings guardrail, Porto notice and rule

Grades and run details

Decision model 83 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly16%
  • failAddresses the actual decision11%
  • partialRespects explicit constraints27%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims32%
  • passProduces the required deliverable61%
  • passAvoids contamination between the arms100%
  • passA realistic plan that beats the freeze100%
  • passAn unambiguous primary metric43%
  • partialDecision rule written before the test76%
  • passSized from the real traffic25%
  • passGuardrails with thresholds97%
Run
Run
#1
API response time
67 s
Submitted
30 Sept 2026

Usable with a quick edit10 of 12 checks passed

The spec is comprehensive and only needs a quick edit to add a trust check and rephrase one unsupported claim; no re-checking of evidence or redoing of reasoning is required.

Output

Experiment spec: Order batching before the peak-season freeze

Decision owner: COO · Sign-off: COO, CFO, Head of Operations, Courier Relations Lead · Copied: Head of Growth Status: Draft for sign-off, 30 September 2026

1. Decision

Should we switch batching on in all 14 cities before the code freeze on 6 November 2026? The switch works per city, so the result can also be "on in some cities."

2. What changed from the draft plan, and why

DraftProblemFix
Randomise orders 50/50Treatment and control orders share the same couriers. Batching frees courier time, and that spare capacity speeds up control orders, which shrinks the measured difference. Earnings per courier hour can't be split by order at all.Switchback test: randomise each city in 2-hour windows. Everything in a window is either on or off.
Runs in every city from day onePorto needs 14 days' written notice. The earliest a notice sent today can take effect is 14 October.Send notice today. Porto joins on 14 Oct.
"Late deliveries"UndefinedShare of orders delivered >45 min after ordering (baseline 7.5%)
30-day reorder rate as a success metricIt can't be measured before the freeze. In a switchback, each customer also sees both arms.Monitor it after rollout. It is not a decision metric.
No thresholds or stop rulesWe can't make a pass/fail callThresholds in §5, stop rules in §7
Two weeks, assuming 740 independent windowsWindows are not independent. Batches straddle switch times, couriers reposition, and windows on the same city-day are correlated.Exclude a washout period from each window, cluster the analysis, and run 21 days (~1,700 windows).

3. Timeline

DateMilestone
Wed 30 SepWritten notice to the Porto couriers' association. It covers both the test and a possible rollout. Legal to confirm wording today.
30 Sep – 7 OctEngineering builds the per-city switch. Analytics builds the randomisation schedule and dashboard.
Thu 8 OctTest starts in 13 cities
Wed 14 OctPorto joins (notice period complete)
Wed 28 OctTest ends
29 Oct – 1 NovAnalysis
Mon 2 NovDecision meeting
Tue 3 NovRollout (one day)
4–5 NovBuffer for rollback or fixes. Freeze on 6 Nov.

The dates are fixed by two constraints. The Porto notice cannot slip past today without losing Porto days. The test cannot run past 28 October without eating into the rollout buffer.

4. Design

  • Unit: one city in one 2-hour window. Each day has six windows (11–13, 13–15, 15–17, 17–19, 19–21, 21–23).
  • Assignment: randomised 50/50 within each city and time slot. Each city gets a balanced mix of on and off across lunch, dinner, weekdays and weekends.
  • Switch behaviour: new batches form only while the switch is on. Batches already in progress when it turns off complete normally.
  • Washout: the first 30 minutes of each window are excluded from analysis. This leaves 90 analysed minutes per window.
  • Sample: 13 cities × 6 windows × 21 days plus Porto × 6 × 15 days gives 1,728 windows. That is about 2.3× the analyst's 740. The extra covers the washout and the correlation between windows. Analytics to confirm the power calculation using the correlation seen in the first week's data. If power is short, we extend to 30 October and give up one buffer day.
  • Analysis: compare window-level metrics with adjustments for city, day and time slot. Treat all windows in the same city on the same day as one group when calculating uncertainty. Report the network result and each city's result.

5. Metrics and decision rule

MetricOwnerBaselineRule
Courier cost per orderCFO€7.40Primary. Pass if the estimated reduction is ≥5% (≥€0.37) and the 95% confidence interval excludes zero.
Share of orders delivered >45 minHead of Ops7.5%Guardrail. Fail if the increase is more than 1.0 percentage point (to 8.5%).
Courier earnings per active hourCourier Relations€13.20Guardrail. Fail if the drop is more than €0.40 (3%). In Porto, any drop fails.
Median delivery time, second order in a batchHead of Opsn/aDiagnostic. Simulation expects +3–6 min.
Complaints and refunds per 1,000 ordersHead of Opsn/aDiagnostic. This is our short-term stand-in for customer harm.
Share of orders batchedAll0%Diagnostic. Shows whether batching is actually happening.
30-day reorder rateGrowthn/aTracked after rollout. Not a decision metric.

Rollout rule: a city rolls out on 3 Nov if the network result passes the primary metric and both guardrails, and that city's own result shows no clear guardrail failure.

The 1.0pp and €0.40 margins are proposals. Signers must confirm or change them before 8 October. They cannot be changed after the test starts.

6. Tensions signers should know about

What "at least 5%" can mean. We can reliably detect whether batching cuts cost by 5% versus not at all. We cannot prove the true effect is above 5% if it is only 6%, which is the low end of the simulation. The decision rule therefore uses the point estimate. If the estimate lands between 4% and 6%, the COO decides with that uncertainty in view.

Cost savings and courier earnings pull in opposite directions. Couriers are paid per order plus distance, so a lower cost per order is lower pay per order. Earnings per hour only hold up if couriers complete more orders per hour.

  • During the test, the same couriers are logged in whether batching is on or off. That makes the test a fair early read on hourly earnings.
  • After rollout, the saving only stays a saving if we need fewer courier hours. Fewer hours affects couriers' income, even if the hourly rate is protected.
  • The CFO and Courier Relations Lead should agree on this before sign-off, not after results.

Porto. "Must not reduce average hourly earnings" is a contractual condition, not a statistical margin. Porto gives us only ~90 windows, too few to prove "no reduction."

  • If network earnings per hour fall at all, Porto stays off unless it comes with an earnings protection mechanism. One example is a per-batch supplement, which legal and Courier Relations should scope now.
  • If Porto rolls out, we measure its hourly earnings against the pre-batching baseline every week.

"Noticeably longer." We are defining this as the >45-minute late share plus the complaint rate. We are not using average delivery time, because a 3–6 minute delay on second orders is expected by design.

7. Stop rules (checked daily by the on-call analyst)

  • City level: the >45-minute late share in "on" windows exceeds 10% for two consecutive days. Batching goes off in that city pending review.
  • Network level:
  • Any courier safety incident linked to batching, or a formal complaint from a couriers' association: pause everywhere.
  • Earnings per hour in "on" windows below €12.50 for three consecutive days: pause everywhere.
  • Porto only: any clear drop in hourly earnings leads to an immediate pause and notification to Courier Relations.

8. What each signer is agreeing to

  • COO: the dates in §3, the rollout rule in §5, and making the call if the result falls between 4% and 6%.
  • CFO: cost per order as the primary metric, with the ≥5% rule. Also acknowledges that post-rollout savings depend on how many courier hours we buy (§6).
  • Head of Operations: the 1.0pp late-share margin and the stop rules.
  • Courier Relations Lead:
  • the €0.40 margin;
  • the separate rule for Porto;
  • sending the Porto notice today.

9. Open items

1. Legal, today: Porto notice wording, and whether one notice covers both the test and the rollout. 2. Analytics, by 7 Oct: - power calculation that accounts for correlated windows; - randomisation schedule; - dashboard. 3. Engineering, by 7 Oct: per-city switch with switch times logged, tested in one city after hours. 4. Courier Relations and Legal, by 2 Nov: Porto earnings protection option. 5. Signers, by 8 Oct: confirm the guardrail margins in §5.

Grades and run details

Decision model 88 · LLM judge 11 of 13 checks
Decision model checks
  • partialUses the supplied evidence correctly8%
  • passAddresses the actual decision43%
  • partialRespects explicit constraints18%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims22%
  • passProduces the required deliverable56%
  • passAvoids contamination between the arms100%
  • passA realistic plan that beats the freeze100%
  • passAn unambiguous primary metric33%
  • partialDecision rule written before the test27%
  • passSized from the real traffic41%
  • passGuardrails with thresholds99%
Run
Run
#1
Time to output
63 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 9

Uses the supplied evidence correctlyMixedRightRight
Gemini 3.8 Flash · API

All factual claims about the current situation are directly from the brief or supplied context.

Sonnet 5.5 · API

All statements about the current situation are taken directly from the brief or follow by arithmetic.

Opus 5.5 · Claude

All claims about the current situation are taken directly from the brief or supplied context.

Addresses the actual decisionRightMixedRight
Gemini 3.8 Flash · API

The spec commits to a conditional rollout (ship if all four conditions are met, otherwise do not) and states the exact conditions that would change the call.

Sonnet 5.5 · API

The spec commits to a conditional rollout decision with clear thresholds and states what would change the call.

Opus 5.5 · Claude

The spec includes a clear decision rule (roll out if primary metric passes and guardrails hold, COO decides if estimate between 4% and 6%) that answers the question of whether to roll out.

Respects explicit constraintsMixedRightRight
Gemini 3.8 Flash · API

The output is an experiment spec under 1,200 words, addresses each stakeholder's metric, respects the Porto notice period, and fits the timeline before the freeze.

Sonnet 5.5 · API

The spec is under 1,200 words, respects the Porto notice and earnings requirement, and fits the timeline before the freeze.

Opus 5.5 · Claude

The output is under 1,200 words, respects the Porto 14-day notice requirement, and addresses all named signers.

Identifies material uncertaintyWrongRightRight
Gemini 3.8 Flash · API

The spec does not name specific unknowns that could change the decision (e.g., whether the 7-day reorder rate proxies 30-day retention, or how courier adaptation might evolve beyond the test window) or say how they would be resolved.

Sonnet 5.5 · API

It names seasonality, carryover, courier adaptation, and Porto legal risk, and says how each would be resolved or monitored.

Opus 5.5 · Claude

It names the uncertainty around detecting a 5% vs 6% effect, the tension between cost savings and courier earnings, and Porto's small sample, and says how they would be resolved.

Avoids unsupported claimsRightRightMixed
Gemini 3.8 Flash · API

Interpretations like the SUTVA violation are presented as design reasoning, not as established facts about the current situation, and no other unsupported factual claims are made.

Sonnet 5.5 · API

Interpretations and forecasts are clearly labelled as such; no factual claim is presented as established without support.

Opus 5.5 · Claude

It presents as fact that the test cannot prove the effect is above 5% if it is 6% and that Porto's 90 windows are too few to prove no reduction, neither of which is supported by the supplied evidence.

A realistic plan that beats the freezeWrongRightRight
Gemini 3.8 Flash · API

The plan takes the 740 independent-window requirement at face value without adjusting for within-city clustering or weekly cycles, and proposes only 14 days, which is too short for courier adaptation and robust weekly coverage.

Sonnet 5.5 · API

It correctly treats 740 windows as a floor, accounts for dependence and weekly cycles with 21 days, and fits the build, Porto notice, and decision before 6 November.

Opus 5.5 · Claude

It correctly converts 740 windows to about 9 days, explains why correlation and weekly cycles require about 3 weeks, and gives dates that fit before the 6 November freeze with Porto notice respected.

An unambiguous primary metricWrongMixedMixed
Gemini 3.8 Flash · API

The spec names courier cost per order as the primary metric but does not include a planned trust check (e.g., sample ratio check or balance verification) to confirm the experiment's integrity.

Sonnet 5.5 · API

The spec names courier cost per order as the primary metric with a rationale, but it does not include a planned trust check such as a sample ratio check or verification that the arms receive the scheduled share of windows and similar order volumes.

Opus 5.5 · Claude

It does not include a planned trust check such as a sample ratio check to verify that the arms are balanced and the experiment ran as intended.

Decision rule written before the testWrongRightRight
Gemini 3.8 Flash · API

The decision rule only covers the case where all conditions are met; it does not state what action to take if the result is inconclusive (e.g., p-value borderline, guardrail ambiguous).

Sonnet 5.5 · API

The decision rule is stated before the test: ship if cost reduction ≥5% with CI excluding zero and guardrails hold; otherwise no rollout, including the inconclusive case.

Opus 5.5 · Claude

It states a rollout rule (pass primary and guardrails, city-level no clear guardrail failure) and says the COO decides if the estimate is between 4% and 6%, covering the inconclusive case.

Sized from the real trafficWrongRightRight
Gemini 3.8 Flash · API

The sample size calculation uses the analyst's 740 windows without adjusting for the non-independence of windows within the same city, and the 14-day duration does not account for the need for multiple full weekly cycles.

Sonnet 5.5 · API

Sample size is derived from the analyst's 740-window floor, increased for dependence, and the duration runs three full weeks using the actual traffic and operating hours.

Opus 5.5 · Claude

Sample size is derived from the analyst's 740 windows, increased for correlation, and runs full weeks (21 days).

All got right 3

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

The output is a complete experiment spec with decision framework, design, metrics, timeline, and sign-off, usable by the named stakeholders with light edits.

Sonnet 5.5 · API

The output is a complete experiment spec in the requested form, under the word limit, and usable by the sign-off group.

Opus 5.5 · Claude

The output is a complete experiment spec with timeline, metrics, decision rule, and sign-off sections, under the word limit, and usable by the signers.

Avoids contamination between the armsRightRightRight
Gemini 3.8 Flash · API

The spec explicitly identifies the shared-courier contamination problem, adopts a city-time switchback design, and includes a 15-minute buffer to prevent carry-over at switches.

Sonnet 5.5 · API

It explains the shared-courier contamination problem and uses a switchback design with a 30-minute buffer to handle carry-over.

Opus 5.5 · Claude

It explains the shared-courier contamination problem and uses a switchback design with a 30-minute washout to keep the arms separate.

Guardrails with thresholdsRightRightRight
Gemini 3.8 Flash · API

Guardrail metrics (late delivery share, courier earnings per hour, 7-day reorder rate, second-order ETA impact) are named with explicit thresholds that would block rollout.

Sonnet 5.5 · API

Guardrails (late deliveries, courier earnings) are named with explicit thresholds that would block rollout.

Opus 5.5 · Claude

It names late delivery share (>45 min) and courier earnings per hour as guardrails with thresholds (1.0pp increase, €0.40 drop, and any drop in Porto).

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review