Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Brisk. We want to know whether to roll out order batching (one courier carrying two orders from nearby restaurants) to all 14 of our cities before the peak-season code freeze on 6 November 2026. Write the experiment spec. Our COO, CFO, Head of Operations and courier relations lead will all sign it off, and each wants something different from it. Keep it under 1,200 words. A draft plan from our data science intern is below. Fix what needs fixing.

What the model was given7 items: About Brisk, What batching should do, What each exec wants to see, Courier agreement in Porto, Timeline, The intern's draft plan, Analyst's note
About BriskFood delivery in 14 European cities, about 1.9 million orders a month. Couriers are paid per order plus distance. Courier cost per order averages €7.40.
What batching should doCOO: “Roll batching out before the freeze if it cuts courier cost per order by at least 5% without making customers wait noticeably longer.” Simulations suggest it cuts courier cost per order by 6–11% and adds 3–6 minutes to the second order in each batch.
What each exec wants to seeCFO: courier cost per order. Head of Operations: the share of orders delivered more than 45 minutes after ordering (today 7.5%). Courier relations lead: courier earnings per active hour (today €13.20). Head of Growth, copied in: 30-day reorder rate.
Courier agreement in PortoOur agreement with the Porto couriers' association requires 14 days' written notice of any change to how orders are assigned, and says changes must not reduce couriers' average hourly earnings.
TimelineToday is 30 September. Engineering needs one week to put batching behind a switch that can be turned on and off per city at any time. Rollout to all cities takes a day. Operating hours are 11:00 to 23:00 in every city.
The intern's draft planRandomise orders 50/50 in every city: orders in the treatment group can be batched, control orders never are. Success metrics: courier cost per order, late deliveries, courier earnings per hour and 30-day reorder rate. Run for two weeks.
Analyst's noteIf we switch batching on and off by city in two-hour windows, detecting a 5% change in courier cost per order needs about 740 windows in total, assuming each window is independent of the others.
What a strong answer doesThe answer key the graders mark against

A spec that sees the draft can't work: batched and unbatched orders in the same city share one pool of couriers, so batching in the treatment group frees couriers for the control group and contaminates the comparison. It uses a switchback design instead (batching on or off by city in randomised two-hour windows), with a buffer at each switch (for example, orders placed shortly before a switch are excluded or assigned by dispatch time) so carry-over doesn't blur the arms. It picks courier cost per order as the one primary metric, because cutting it is batching's purpose and a fall is unambiguously good, and makes the others guardrails with thresholds: late deliveries (for example, no more than a 1pp rise from 7.5%), courier earnings per active hour (no fall, which the Porto agreement requires), with 30-day reorder rate named as a follow-up because it can't be read before the freeze. It sizes the test honestly: 14 cities × 6 windows a day is 84 windows, so 740 windows is about 9 days, but windows in the same city on the same day aren't independent, so it needs more, and a full weekly cycle plus time for couriers to adapt argue for about three weeks. It fits the timeline: build ready around 7 October, Porto given notice now and joining once its 14 days are up (or excluded), about three weeks of data, a read-out around 1 November and rollout before 6 November. It plans trust checks (the arms get the scheduled share of windows and similar order volumes; batching actually happens in 'on' windows), and states the decision rule up front: roll out if cost per order falls at least 5% and guardrails hold; don't if a guardrail is breached; and what happens if the result is inconclusive by the freeze.

Critical failures (cap the score at 40)
  • Randomises by order or by customer within a city, where both arms share the same couriers
  • Changes order assignment in Porto without the 14 days' notice
Case

Staff-level · v1.1 · synthetic · marketplace, logistics, switchback, network effects, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 12 checks passed

The spec is complete and actionable; a quick edit to add a sample ratio check or similar trust check would make it fully robust.

Output

Order Batching Experiment Spec

Decision: Roll out batching to all 14 cities before the 6 Nov 2026 code freeze, or not. Owner: Staff PM. Sign-off: COO, CFO, Head of Operations, Courier Relations Lead. Head of Growth is copied in.

1. What changed from the intern's draft

DraftProblemFix
Randomise orders 50/50 within each cityTreated and control orders share the same couriers and the same pool of nearby orders. Halving the batchable pool understates batching, and control orders are affected by treatment (couriers are tied up on batches). Both effects bias the result.Switch batching on and off by city and two-hour window (switchback).
Porto included from day oneThe agreement requires 14 days' written notice of any change to order assignment, and a test is such a change.Send notice today. Porto joins on 15 Oct.
Four metrics, no thresholdsNobody can say what "good" means.Pre-agreed decision rules (section 4).
30-day reorder rate as a success metricA test starting in October cannot produce a 30-day read before 6 Nov.Monitor it, but it does not gate the decision.
Two weeks, no justificationSee section 3.Three full weeks.

2. Design

  • Unit: city × two-hour window. The windows are 11–13, 13–15, 15–17, 17–19, 19–21 and 21–23, so six per city per day.
  • Assignment: Within each city-day, the windows form three adjacent pairs. In each pair, one window is randomly batching-on and the other is off. This balances lunch, afternoon and dinner across arms and controls for city and weekday.
  • Attribution: An order belongs to the window in which it was placed. The pre-specified sensitivity analysis drops orders placed in the first 30 minutes of each window, because in-flight batches and courier positioning carry over from the previous window.
  • Exposure: Batching pairs only nearby restaurants, using the same rules the simulation assumed. The batching configuration is frozen before the test starts.

3. Duration and power

The analyst's figure of about 740 windows assumes windows are independent. They are not: adjacent windows share couriers and demand, and cities differ. I therefore treat 740 as a floor and plan for up to double that.

  • Engineering needs one week, so the switch is ready on 7 Oct and the test runs Thu 8 Oct – Wed 28 Oct (21 days, three of every weekday).
  • 13 cities × 6 windows × 21 days = 1,638 windows. Porto adds 6 × 14 days = 84. The total is about 1,720 windows, roughly 2.3× the floor.
  • We do not stop early on a good-looking interim result. The end date is fixed.
  • Analysis: A window-level regression on total cost ÷ total orders, with city and weekday × hour fixed effects and standard errors clustered by city-day. We will confirm with randomisation inference.

4. Metrics and decision rules

All four sign-offs are needed on these thresholds by Tue 6 Oct, before any data exist.

OwnerMetricRoleRule to ship
COO / CFOCourier cost per order (today €7.40)PrimaryEstimated reduction is ≥5% and the 95% CI excludes zero
Head of OpsShare of orders delivered >45 min after ordering (today 7.5%)GuardrailIncrease of no more than +1.0 pp (upper 95% CI bound ≤ 8.5%)
Courier RelationsCourier earnings per active hour (today €13.20)GuardrailLower 95% CI bound of the difference ≥ −€0.13 (about −1%)
Head of GrowthReorder rateMonitor onlyReported, does not gate

The COO's "without making customers wait noticeably longer" is operationalised as the +1.0 pp cap. The simulations add 3–6 minutes to the second order in each batch, so this cap is the threshold most likely to be contested. The Head of Ops should confirm or change it before the start.

Tension to surface: Cost per order is forecast to fall 6–11%, and couriers are paid per order plus distance. Earnings per hour therefore stay flat only if batching raises orders per courier-hour by at least as much as pay per order falls. The test settles this, and the CFO's and Courier Relations' metrics are not independent.

Secondary metrics (descriptive only): - Mean delivery time, and delivery time for the second order in a batch. - Share of orders batched. - Orders per courier-hour. - Customer support contacts. - Cancellations.

Reorder rate: We report the 7-day rate for test customers at the decision. The 30-day rate is tracked after rollout for customers by batched-order exposure, and it cannot be treated as causal. Growth should know this now.

5. Porto

  • Written notice goes to the couriers' association today (30 Sept). Legal should word it to cover both the test and permanent rollout. The earliest compliant start is 14 Oct, so Porto joins the test on 15 Oct.
  • The contract says changes must not reduce average hourly earnings. Porto therefore has a stricter rule than other cities: the estimated earnings difference must be ≥ €0 for Porto to ship.
  • Porto has only about 84 windows, so its own estimate will be noisy. We pool it for cost and lateness. For earnings we report Porto separately, daily, with the stop rule below.
  • If Porto fails or is ambiguous, we ship in the other 13 cities and keep Porto off. Legal should confirm that running the test in Porto does not itself breach the agreement. If it might, exclude Porto from the test and roll out there only on the pooled evidence plus the earnings rule.

6. Stop rules

The per-city switch lets us turn batching off immediately. Any of these triggers a pause in that city and a review within 24 hours:

  • The >45-minute share in batching-on windows exceeds off windows by more than 3 pp over two consecutive days.
  • Earnings per active hour in batching-on windows are more than 5% below off windows over three consecutive days. In Porto, the threshold is any sustained shortfall that the Courier Relations Lead judges real.
  • Any safety or serious courier-welfare incident tied to batching.

7. Timeline

DateMilestone
Wed 30 SeptSpec circulated. Porto notice sent. Engineering starts the per-city switch.
Tue 6 OctThresholds signed off. Randomisation schedule locked.
Wed 7 OctSwitch ready, QA complete
Thu 8 OctTest starts in 13 cities
Thu 15 OctPorto joins
Wed 28 OctTest ends
Thu 29 OctAnalysis and go/no-go with all four sign-offs
Fri 30 OctRollout to approved cities (one day)
30 Oct – 5 NovPost-rollout monitoring with the switch live. The freeze is on 6 Nov.

The 29 Oct decision leaves a week of buffer. If the test slips past about 2 Nov, there is no safe rollout before the freeze.

8. Risks and limits

  • Seasonality: October demand is below peak. Higher order density should raise batch rates, but it also strains courier supply. We track batch rate by hour and expect the peak effect to differ from the test estimate.
  • Carryover: Some carryover between windows will remain even with the 30-minute washout. The sensitivity analysis shows how much it matters.
  • Short-run effects: Couriers may change their behaviour as they learn batching, and three weeks may not capture that. We monitor weekly trends within the test.
  • Possible outcomes: The result can be a full rollout, a partial rollout (for example excluding Porto or any city that fails a guardrail), or no rollout. Partial rollout by city is allowed by the switch, but only for cities that individually pass the stop rules. City-level results will be noisy, so we do not make ship decisions city by city beyond this.

9. Sign-off

RoleAgrees to
COODecision rule, timeline
CFOCost metric definition (fully loaded courier pay ÷ delivered orders), ≥5% threshold
Head of Ops+1.0 pp lateness cap, stop rules
Courier RelationsEarnings guardrail, Porto notice and rule

Grades and run details

Decision model 83 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly16%
  • failAddresses the actual decision11%
  • partialRespects explicit constraints27%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims32%
  • passProduces the required deliverable61%
  • passAvoids contamination between the arms100%
  • passA realistic plan that beats the freeze100%
  • passAn unambiguous primary metric43%
  • partialDecision rule written before the test76%
  • passSized from the real traffic25%
  • passGuardrails with thresholds97%
Run
Run
#1
API response time
67 s
Submitted
30 Sept 2026

Usable with a quick edit12 of 12 checks passed

The spec is complete, evidence-based, and directly actionable; a PM could use it with only minor wording tweaks and no need to re-verify facts or redesign the experiment.

Output

Brisk order-batching experiment spec

Decision owner: Staff PM Sign-off: COO, CFO, Head of Operations, Courier Relations Decision deadline: 4 November 2026; rollout on 5 November, before the 6 November freeze.

1. Decision and hypothesis

Test whether batching can reduce total courier cost per delivered order by at least 5%, without materially worsening delivery times or reducing courier earnings.

At today’s €7.40 baseline, 5% means approximately €0.37 saved per order. Simulations are directional evidence, not sufficient grounds for rollout.

Decision: launch across all 14 cities only if the economic, customer, courier and operational gates below pass. Otherwise, leave batching off before the freeze. A promising but inconclusive result is not a pass.

2. Treatment and eligibility

Treatment enables one courier to carry at most two orders from nearby restaurants. Control retains current assignment.

Before testing, Operations and Engineering will lock:

  • Restaurant-proximity, pickup-readiness and maximum-detour rules.
  • Maximum predicted delivery times for both orders.
  • Courier payment rules, including distance calculation.
  • Exception handling, cancellation treatment and customer communications.

Do not change these rules during the confirmatory experiment. A material change requires a new test.

Measure the policy’s effect across all orders, not merely successfully batched orders. Report batching rate and first-/second-order outcomes as diagnostics; comparing batched orders with unbatched orders is selection-biased.

3. Experimental design

Replace order-level randomization with city-wide switchbacks. Treatment orders would otherwise change courier availability and dispatch conditions for control orders, contaminating the comparison.

Each city has six two-hour blocks daily between 11:00 and 23:00. Randomize batching on/off by city-block, with balanced assignment across cities, dates and time of day. Use constrained random schedules—not deterministic alternation.

Before the confirmatory test:

  • Set a washout period using historical order/trip completion data and the pilot.
  • Exclude the same initial portion of every block from measurement, whether or not its state changes.
  • Validate that residual trips and courier repositioning do not materially contaminate measurement. If they do, lengthen blocks or washout and recalculate feasibility.

Attribute order outcomes to order-placement time; follow included orders through completion even after their block ends. Define payout and active-hour attribution consistently before launch.

Analysis is intention-to-treat by assigned city-block. Weight results to the expected rollout order mix, rather than weighting small and large cities equally. Report city-level results and heterogeneity alongside the aggregate.

4. Metrics and acceptance gates

The following customer thresholds are proposed definitions of “not noticeably longer” and require sign-off before testing.

StakeholderMetricGate
CFO / COOTotal courier cost per delivered order, including distance pay, incentives and any earnings protectionOne-sided 95% confidence bound supports at least 5% savings
OperationsShare delivered more than 45 minutes after ordering; baseline 7.5%Upper confidence bound on increase ≤ 0.5 percentage points
COO / OperationsMean and p90 order-to-delivery timeUpper confidence bounds on increases ≤ 1 minute mean, 2 minutes p90
Courier RelationsCourier earnings per active hour; baseline €13.20Lower confidence bound on treatment effect ≥ €0/hour

Define active hours identically in both arms, including available waiting time—not just time carrying orders. Report earnings distribution and effects by city and courier segment, not only the pooled mean.

Use baseline-adjusted treatment/control comparisons, not historical averages as the counterfactual. Log cancellations, failed deliveries, assignment failures and complaints as safety diagnostics; cost savings driven by worse completion are unacceptable.

Porto is a separate launch gate: pooled earnings gains cannot override a Porto earnings reduction. Legal and Courier Relations must approve the compliance interpretation and protection mechanism. Any guarantee or top-up must be part of the tested policy and counted in courier cost.

Growth: 30-day reorder is exploratory, not a pre-freeze gate. Even customers first exposed on 14 October do not mature until 13 November. Report available short-term indicators separately and agree a post-freeze retention study; do not present immature reorder data as a 30-day result.

5. Power and analysis

The intern’s “two weeks” is not justified. The analyst’s 740 windows assumes independence and addresses detecting a 5% effect—not necessarily proving savings exceed 5%.

From 14 October through 2 November, the maximum is:

14 cities × 6 blocks × 20 days = 1,680 blocks, before data losses.

That is only 2.27 times the nominal requirement. Serial correlation, washout, uneven volume and separate guardrail requirements can consume this headroom.

By 13 October, Data Science must replay the proposed schedule on historical/pilot data and establish:

  • Power to pass the savings threshold under plausible true savings, including 6%, 8% and 11%.
  • Power for delivery and earnings non-inferiority, especially Porto.
  • Sensitivity to carryover, missing data and city differences.

Use inference respecting the randomized schedule, with serial-correlation and small-number-of-cities sensitivity checks. Pre-register estimators, confidence bounds, exclusions, city safety rules and any multiplicity adjustment for city-specific claims.

Feasibility gate: if adequate power cannot fit before the deadline, do not promise an all-city rollout. Extend testing beyond the freeze or return with a narrower decision.

6. Timeline and operating controls

  • 30 September: obtain metric-threshold agreement; send Porto’s written assignment-change notice. Confirm legally valid delivery and when its 14-day period expires.
  • By 7 October: deliver city switches, instrumentation, audit logs and rollback.
  • 8–13 October: technical/safety pilot outside Porto; validate washout, payments and power. Pilot data excluded from confirmatory results.
  • 14 October–2 November: confirmatory test, with Porto starting only after the full notice period has elapsed.
  • 3–4 November: finalize outcomes, review compliance and make the signed decision.
  • 5 November: one-day rollout only if all gates pass.

Operations owns live monitoring; Engineering owns rollback; Courier Relations owns courier communications and payment disputes. Stop affected cities immediately for payment errors, serious safety incidents or predefined severe service degradation. Such stops are recorded, not silently excluded.

No repeated efficacy peeking. If the test ends early for safety or loses its required sample, it does not automatically qualify for launch. Keep the kill switch available throughout peak season.

Grades and run details

Decision model 92 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly16%
  • passAddresses the actual decision82%
  • partialRespects explicit constraints22%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims62%
  • passProduces the required deliverable55%
  • passAvoids contamination between the arms100%
  • passA realistic plan that beats the freeze76%
  • passAn unambiguous primary metric47%
  • passDecision rule written before the test87%
  • partialSized from the real traffic62%
  • passGuardrails with thresholds89%
Run
Run
#1
API response time
64 s
Submitted
30 Sept 2026

Needs more than a quick edit9 of 12 checks passed

The spec is strong but would need material edits to the sizing and timeline, including converting the 740-window note, justifying a full three-week run, and adding explicit trust checks. A PM could not use it without reworking the experiment duration and power rationale.

Output

Brisk order-batching experiment: pre-freeze decision

Decision owner: Staff PM Required approvers: COO, CFO, Head of Operations, courier relations lead Decision deadline: 4 November 2026; rollout on 5 November, before the 6 November freeze.

1. Decision and hypothesis

Ship batching to all 14 cities only if it demonstrates at least 5% lower courier cost per order, with acceptable customer experience and protected courier earnings.

At today’s €7.40 baseline, 5% is approximately €0.37 per order. Simulated savings are not sufficient evidence.

The thresholds below are proposed sign-off terms, not established definitions of “noticeably longer.” Approvers must agree them before randomisation. An inconclusive result means no all-city rollout before the freeze, not permission to lower the bar.

2. Replace order randomisation with city switchbacks

Order-level randomisation is unsuitable: treatment and control orders share couriers, restaurant queues and dispatch capacity. Batching changes availability for both groups, contaminating the comparison.

Instead, switch each city between:

  • Treatment: batching enabled under frozen restaurant-proximity, route-detour and eligibility rules; maximum two orders per courier.
  • Control: current single-order dispatch.

Randomise six local-time blocks daily: 11–13, 13–15, …, 21–23. Balance treatment across cities, weekdays and clock times, with approximately equal exposure per city. Generate and lock the schedule in advance; do not let operators choose favourable treatment periods.

This estimates the effect of offering batching across a city, not merely its effect on successfully batched orders. Analyse all orders in the measurement cohort, including those never batched.

Carryover: Complete existing assignments safely at switches; never split an accepted batch. Set a fixed washout period using historical and shadow-mode assignment-to-completion times, before observing treatment results. Exclude order arrivals during that period from the primary cohort in every block, regardless of whether the state changed. Follow included orders through completion; do not truncate outcomes at block boundaries. If carryover cannot be contained within two-hour blocks, use longer blocks and re-power.

3. Metrics and launch gates

Use concurrent controls and pre-specified, one-sided 95% confidence bounds. Every gate must pass.

StakeholderMetricLaunch requirement
CFO / COOTotal courier compensation attributable to the order cohort ÷ completed ordersUpper confidence bound on treatment/control ratio ≤0.95
OperationsShare delivered more than 45 minutes after ordering; baseline 7.5%Upper bound on increase ≤0.5 percentage points
COO / OperationsMean order-to-door timeUpper bound on increase ≤2 minutes
Courier relationsEarnings per active hour; network baseline €13.20Lower bound on relative change ≥−2% outside Porto
PortoContract-defined average hourly earningsNo-reduction obligation met, as described below

Cost includes distance pay, incentives, failed-delivery compensation and any batching-related earnings protection. Finance freezes allocation rules before launch. Track cancellation/completion rates so apparent savings cannot come from abandoning orders.

Define active hours as available-for-work plus engaged time, excluding genuine offline breaks—not just delivery time. Use consistent definitions across arms. Report both total earnings/total hours and courier-level distributions.

Report results by city, peak/off-peak, first/second batch position and courier earnings decile. Batch-position comparisons are descriptive, not randomised causal estimates. Examine late-delivery tails and cancellations; averages must not conceal material local harm.

Growth: Thirty-day reorder is exploratory, not a freeze gate. The first main-trial cohort matures on 14 November. Customer crossover also limits interpretation; do not present this switchback as a clean test of persistent customer retention.

4. Porto is a contractual prerequisite

Send written notice on 30 September, covering experimental on/off assignment changes and potential rollout. Legal and courier relations must confirm valid receipt, scope and the earliest permitted activation. Plan Porto activation for 15 October, providing a buffer beyond 14 days; late receipt moves that date.

Before activation, agree the contract’s hourly-earnings definition and comparison method with the association. The network’s €13.20 is not a Porto-specific contractual baseline.

Implement a funded make-whole mechanism accepted as satisfying the no-reduction requirement. Include those payments in experimental cost. A statistical non-inferiority margin cannot waive the agreement. Without compliant notice and earnings protection, Porto stays off—and the all-14-city decision is no-go.

5. Power, analysis and validity

The analyst’s 740 windows is not a sufficient sample-size justification:

  • Adjacent windows share demand, weather, couriers and backlog.
  • Washout reduces usable exposure.
  • Detecting a 5% saving versus zero differs from proving a saving of at least 5%.
  • Customer and earnings gates may require more data than cost.

Before launch, replay historical city-time data under the proposed schedule. Incorporate autocorrelation, city heterogeneity, washout and realistic take-up. Estimate power for every gate and the probability of passing them jointly, using simulation scenarios spanning 6–11% savings. Target at least 90% joint power under an explicitly agreed planning scenario.

Estimate intention-to-treat effects using city and calendar/clock-time adjustment, with fleet weights fixed from pre-experiment volume. Use schedule-respecting randomisation inference and dependence-aware uncertainty—not independent-window standard errors.

Publish city-level uncertainty. Pre-agree local safety thresholds and multiplicity handling; a pooled pass does not automatically authorise a city showing credible harm.

Feasibility gate: If adequate evidence cannot fit before the deadline, recommend no all-city pre-freeze launch.

6. Calendar and operating controls

  • 30 September: Approve protocol and proposed margins; send Porto notice.
  • 1–7 October: Build city switches, assignment logging, earnings protection and tested emergency rollback.
  • 8–14 October: Shadow-mode validation; finalise washout, power, schedule and registered analysis.
  • 15 October–2 November: Main experiment: 19 days, at most 1,596 nominal city-blocks before exclusions—not 1,596 independent observations.
  • 3–4 November: Reconcile outcomes and compensation; analyse; obtain four signatures.
  • 5 November: If approved, stage rollout across cities within the day, with monitored checkpoints.

Operations owns live monitoring. Pre-register severe-delay, cancellation, earnings and system-failure stop rules and a sequential-monitoring procedure. Emergency pauses remain permissible; log overrides and retain affected observations in assignment-based analysis. No opportunistic early declaration of success.

7. Sign-off package

Deliver effect estimates and confidence bounds, city diagnostics, notice/compliance evidence, incident logs, and Finance’s net savings forecast including protection payments and operating costs.

Go: Every statistical, contractual, local-safety and operational gate passes. No-go: Any gate fails, remains inconclusive, or cannot be evaluated reliably by 4 November. A subset rollout requires a separate, explicitly approved decision; it is not success against the all-city objective.

Grades and run details

Decision model 88 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly10%
  • passAddresses the actual decision82%
  • partialRespects explicit constraints25%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims64%
  • passProduces the required deliverable44%
  • passAvoids contamination between the arms100%
  • passA realistic plan that beats the freeze99%
  • passAn unambiguous primary metric27%
  • passDecision rule written before the test98%
  • failSized from the real traffic24%
  • passGuardrails with thresholds89%
Run
Run
#1
Time to output
81 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Addresses the actual decisionMixedRightRight
Sonnet 5.5 · API

The spec commits to a conditional rollout decision with clear thresholds and states what would change the call.

GPT-6.1 Sol · API

The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.

GPT-6 Astra · ChatGPT

It commits early to ship only if courier cost per order falls at least 5% and guardrails pass, and says inconclusive or failed gates mean no all-city pre-freeze rollout.

A realistic plan that beats the freezeRightRightMixed
Sonnet 5.5 · API

It correctly treats 740 windows as a floor, accounts for dependence and weekly cycles with 21 days, and fits the build, Porto notice, and decision before 6 November.

GPT-6.1 Sol · API

It converts 740 windows to about 9 days implicitly, explains that correlation and weekly cycles require a longer test (20 days, ~3 weeks), and provides dates that fit before 6 November with Porto notice given on 30 September.

GPT-6 Astra · ChatGPT

It does not convert 740 windows into about 9 days, does not explicitly justify about three weeks from weekly cycles and courier adaptation, and schedules only 19 days rather than a full three-week test.

An unambiguous primary metricMixedRightMixed
Sonnet 5.5 · API

The spec names courier cost per order as the primary metric with a rationale, but it does not include a planned trust check such as a sample ratio check or verification that the arms receive the scheduled share of windows and similar order volumes.

GPT-6.1 Sol · API

Courier cost per order is the unambiguous primary metric with a clear rationale (COO's 5% savings target), and trust checks like washout validation, diagnostics, and pre-registration are planned.

GPT-6 Astra · ChatGPT

It names courier cost per order as the decision metric but does not clearly explain why lower cost is unambiguously good for users and the business, nor specify a concrete trust check such as sample-ratio or batching-occurrence validation.

Sized from the real trafficRightRightWrong
Sonnet 5.5 · API

Sample size is derived from the analyst's 740-window floor, increased for dependence, and the duration runs three full weeks using the actual traffic and operating hours.

GPT-6.1 Sol · API

The duration uses the supplied traffic (14 cities, 6 blocks/day) and runs whole weeks (14 Oct–2 Nov), with a power analysis step to confirm adequacy against the 5% effect and correlation.

GPT-6 Astra · ChatGPT

It does not size the test from supplied traffic, baseline, and smallest effect, and its 19-day duration is not whole weekly cycles.

All got right 8

Uses the supplied evidence correctlyRightRightRight
Sonnet 5.5 · API

All statements about the current situation are taken directly from the brief or follow by arithmetic.

GPT-6.1 Sol · API

All factual claims about the current situation are directly from the brief or derived by arithmetic, with no invented numbers or facts.

GPT-6 Astra · ChatGPT

The output’s factual statements about the current situation are drawn from the brief or follow by arithmetic, and its invented operational details are framed as proposals or requirements rather than existing facts.

Respects explicit constraintsRightRightRight
Sonnet 5.5 · API

The spec is under 1,200 words, respects the Porto notice and earnings requirement, and fits the timeline before the freeze.

GPT-6.1 Sol · API

The spec respects the word limit, addresses all named stakeholders, handles the Porto notice and earnings requirement, and fits the timeline before the code freeze.

GPT-6 Astra · ChatGPT

It respects the under-1,200-word limit, named approvers, Porto notice, engineering lead time, and freeze date, and its proposed gates and calendar would enforce those constraints.

Identifies material uncertaintyRightRightRight
Sonnet 5.5 · API

It names seasonality, carryover, courier adaptation, and Porto legal risk, and says how each would be resolved or monitored.

GPT-6.1 Sol · API

It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.

GPT-6 Astra · ChatGPT

It names key uncertainties such as carryover, autocorrelation, washout, take-up, power, local harm, and Porto compliance, and specifies how they would be resolved or would change the decision.

Avoids unsupported claimsRightRightRight
Sonnet 5.5 · API

Interpretations and forecasts are clearly labelled as such; no factual claim is presented as established without support.

GPT-6.1 Sol · API

Interpretations and forecasts are clearly labelled as such (e.g., simulations as directional, serial correlation as a risk), and no confident claims go beyond the supplied evidence.

GPT-6 Astra · ChatGPT

It labels thresholds and assumptions as proposed, treats simulations as insufficient evidence, and avoids presenting forecasts or causes as established facts.

Produces the required deliverableRightRightRight
Sonnet 5.5 · API

The output is a complete experiment spec in the requested form, under the word limit, and usable by the sign-off group.

GPT-6.1 Sol · API

The output is a complete experiment spec under 1,200 words, structured for the sign-off group, and contains all sections needed to act on it.

GPT-6 Astra · ChatGPT

It is a usable experiment spec for the named approvers, with design, metrics, Porto handling, calendar, analysis, and sign-off package, within the requested length.

Avoids contamination between the armsRightRightRight
Sonnet 5.5 · API

It explains the shared-courier contamination problem and uses a switchback design with a 30-minute buffer to handle carry-over.

GPT-6.1 Sol · API

It explicitly rejects order-level randomization due to shared couriers, adopts a city-block switchback design, and includes a washout period to prevent carry-over contamination.

GPT-6 Astra · ChatGPT

It rejects order-level randomisation, uses city switchbacks, and adds washout and assignment-completion rules to handle carryover at switches.

Decision rule written before the testRightRightRight
Sonnet 5.5 · API

The decision rule is stated before the test: ship if cost reduction ≥5% with CI excluding zero and guardrails hold; otherwise no rollout, including the inconclusive case.

GPT-6.1 Sol · API

The rule is stated upfront: launch if all gates pass, do not launch otherwise, and an inconclusive result is treated as a no-go.

GPT-6 Astra · ChatGPT

It states a pre-test rule: go only if every gate passes, no-go if any gate fails or is inconclusive, and no all-city rollout before the freeze if evidence cannot be obtained reliably.

Guardrails with thresholdsRightRightRight
Sonnet 5.5 · API

Guardrails (late deliveries, courier earnings) are named with explicit thresholds that would block rollout.

GPT-6.1 Sol · API

Guardrails are named (late deliveries, delivery time, courier earnings) with specific thresholds (0.5 pp, 1 min/2 min, €0/hour) that would block rollout.

GPT-6 Astra · ChatGPT

It names late deliveries, order-to-door time, courier earnings, Porto earnings, and cancellation/completion as blocking gates with thresholds.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review