Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Brisk. We want to know whether to roll out order batching (one courier carrying two orders from nearby restaurants) to all 14 of our cities before the peak-season code freeze on 6 November 2026. Write the experiment spec. Our COO, CFO, Head of Operations and courier relations lead will all sign it off, and each wants something different from it. Keep it under 1,200 words. A draft plan from our data science intern is below. Fix what needs fixing.

What the model was given7 items: About Brisk, What batching should do, What each exec wants to see, Courier agreement in Porto, Timeline, The intern's draft plan, Analyst's note
About BriskFood delivery in 14 European cities, about 1.9 million orders a month. Couriers are paid per order plus distance. Courier cost per order averages €7.40.
What batching should doCOO: “Roll batching out before the freeze if it cuts courier cost per order by at least 5% without making customers wait noticeably longer.” Simulations suggest it cuts courier cost per order by 6–11% and adds 3–6 minutes to the second order in each batch.
What each exec wants to seeCFO: courier cost per order. Head of Operations: the share of orders delivered more than 45 minutes after ordering (today 7.5%). Courier relations lead: courier earnings per active hour (today €13.20). Head of Growth, copied in: 30-day reorder rate.
Courier agreement in PortoOur agreement with the Porto couriers' association requires 14 days' written notice of any change to how orders are assigned, and says changes must not reduce couriers' average hourly earnings.
TimelineToday is 30 September. Engineering needs one week to put batching behind a switch that can be turned on and off per city at any time. Rollout to all cities takes a day. Operating hours are 11:00 to 23:00 in every city.
The intern's draft planRandomise orders 50/50 in every city: orders in the treatment group can be batched, control orders never are. Success metrics: courier cost per order, late deliveries, courier earnings per hour and 30-day reorder rate. Run for two weeks.
Analyst's noteIf we switch batching on and off by city in two-hour windows, detecting a 5% change in courier cost per order needs about 740 windows in total, assuming each window is independent of the others.
What a strong answer doesThe answer key the graders mark against

A spec that sees the draft can't work: batched and unbatched orders in the same city share one pool of couriers, so batching in the treatment group frees couriers for the control group and contaminates the comparison. It uses a switchback design instead (batching on or off by city in randomised two-hour windows), with a buffer at each switch (for example, orders placed shortly before a switch are excluded or assigned by dispatch time) so carry-over doesn't blur the arms. It picks courier cost per order as the one primary metric, because cutting it is batching's purpose and a fall is unambiguously good, and makes the others guardrails with thresholds: late deliveries (for example, no more than a 1pp rise from 7.5%), courier earnings per active hour (no fall, which the Porto agreement requires), with 30-day reorder rate named as a follow-up because it can't be read before the freeze. It sizes the test honestly: 14 cities × 6 windows a day is 84 windows, so 740 windows is about 9 days, but windows in the same city on the same day aren't independent, so it needs more, and a full weekly cycle plus time for couriers to adapt argue for about three weeks. It fits the timeline: build ready around 7 October, Porto given notice now and joining once its 14 days are up (or excluded), about three weeks of data, a read-out around 1 November and rollout before 6 November. It plans trust checks (the arms get the scheduled share of windows and similar order volumes; batching actually happens in 'on' windows), and states the decision rule up front: roll out if cost per order falls at least 5% and guardrails hold; don't if a guardrail is breached; and what happens if the result is inconclusive by the freeze.

Critical failures (cap the score at 40)
  • Randomises by order or by customer within a city, where both arms share the same couriers
  • Changes order assignment in Porto without the 14 days' notice
Case

Staff-level · v1.1 · synthetic · marketplace, logistics, switchback, network effects, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit9 of 12 checks passed

The spec is strong but would need material edits to the sizing and timeline, including converting the 740-window note, justifying a full three-week run, and adding explicit trust checks. A PM could not use it without reworking the experiment duration and power rationale.

Output

Brisk order-batching experiment: pre-freeze decision

Decision owner: Staff PM Required approvers: COO, CFO, Head of Operations, courier relations lead Decision deadline: 4 November 2026; rollout on 5 November, before the 6 November freeze.

1. Decision and hypothesis

Ship batching to all 14 cities only if it demonstrates at least 5% lower courier cost per order, with acceptable customer experience and protected courier earnings.

At today’s €7.40 baseline, 5% is approximately €0.37 per order. Simulated savings are not sufficient evidence.

The thresholds below are proposed sign-off terms, not established definitions of “noticeably longer.” Approvers must agree them before randomisation. An inconclusive result means no all-city rollout before the freeze, not permission to lower the bar.

2. Replace order randomisation with city switchbacks

Order-level randomisation is unsuitable: treatment and control orders share couriers, restaurant queues and dispatch capacity. Batching changes availability for both groups, contaminating the comparison.

Instead, switch each city between:

  • Treatment: batching enabled under frozen restaurant-proximity, route-detour and eligibility rules; maximum two orders per courier.
  • Control: current single-order dispatch.

Randomise six local-time blocks daily: 11–13, 13–15, …, 21–23. Balance treatment across cities, weekdays and clock times, with approximately equal exposure per city. Generate and lock the schedule in advance; do not let operators choose favourable treatment periods.

This estimates the effect of offering batching across a city, not merely its effect on successfully batched orders. Analyse all orders in the measurement cohort, including those never batched.

Carryover: Complete existing assignments safely at switches; never split an accepted batch. Set a fixed washout period using historical and shadow-mode assignment-to-completion times, before observing treatment results. Exclude order arrivals during that period from the primary cohort in every block, regardless of whether the state changed. Follow included orders through completion; do not truncate outcomes at block boundaries. If carryover cannot be contained within two-hour blocks, use longer blocks and re-power.

3. Metrics and launch gates

Use concurrent controls and pre-specified, one-sided 95% confidence bounds. Every gate must pass.

StakeholderMetricLaunch requirement
CFO / COOTotal courier compensation attributable to the order cohort ÷ completed ordersUpper confidence bound on treatment/control ratio ≤0.95
OperationsShare delivered more than 45 minutes after ordering; baseline 7.5%Upper bound on increase ≤0.5 percentage points
COO / OperationsMean order-to-door timeUpper bound on increase ≤2 minutes
Courier relationsEarnings per active hour; network baseline €13.20Lower bound on relative change ≥−2% outside Porto
PortoContract-defined average hourly earningsNo-reduction obligation met, as described below

Cost includes distance pay, incentives, failed-delivery compensation and any batching-related earnings protection. Finance freezes allocation rules before launch. Track cancellation/completion rates so apparent savings cannot come from abandoning orders.

Define active hours as available-for-work plus engaged time, excluding genuine offline breaks—not just delivery time. Use consistent definitions across arms. Report both total earnings/total hours and courier-level distributions.

Report results by city, peak/off-peak, first/second batch position and courier earnings decile. Batch-position comparisons are descriptive, not randomised causal estimates. Examine late-delivery tails and cancellations; averages must not conceal material local harm.

Growth: Thirty-day reorder is exploratory, not a freeze gate. The first main-trial cohort matures on 14 November. Customer crossover also limits interpretation; do not present this switchback as a clean test of persistent customer retention.

4. Porto is a contractual prerequisite

Send written notice on 30 September, covering experimental on/off assignment changes and potential rollout. Legal and courier relations must confirm valid receipt, scope and the earliest permitted activation. Plan Porto activation for 15 October, providing a buffer beyond 14 days; late receipt moves that date.

Before activation, agree the contract’s hourly-earnings definition and comparison method with the association. The network’s €13.20 is not a Porto-specific contractual baseline.

Implement a funded make-whole mechanism accepted as satisfying the no-reduction requirement. Include those payments in experimental cost. A statistical non-inferiority margin cannot waive the agreement. Without compliant notice and earnings protection, Porto stays off—and the all-14-city decision is no-go.

5. Power, analysis and validity

The analyst’s 740 windows is not a sufficient sample-size justification:

  • Adjacent windows share demand, weather, couriers and backlog.
  • Washout reduces usable exposure.
  • Detecting a 5% saving versus zero differs from proving a saving of at least 5%.
  • Customer and earnings gates may require more data than cost.

Before launch, replay historical city-time data under the proposed schedule. Incorporate autocorrelation, city heterogeneity, washout and realistic take-up. Estimate power for every gate and the probability of passing them jointly, using simulation scenarios spanning 6–11% savings. Target at least 90% joint power under an explicitly agreed planning scenario.

Estimate intention-to-treat effects using city and calendar/clock-time adjustment, with fleet weights fixed from pre-experiment volume. Use schedule-respecting randomisation inference and dependence-aware uncertainty—not independent-window standard errors.

Publish city-level uncertainty. Pre-agree local safety thresholds and multiplicity handling; a pooled pass does not automatically authorise a city showing credible harm.

Feasibility gate: If adequate evidence cannot fit before the deadline, recommend no all-city pre-freeze launch.

6. Calendar and operating controls

  • 30 September: Approve protocol and proposed margins; send Porto notice.
  • 1–7 October: Build city switches, assignment logging, earnings protection and tested emergency rollback.
  • 8–14 October: Shadow-mode validation; finalise washout, power, schedule and registered analysis.
  • 15 October–2 November: Main experiment: 19 days, at most 1,596 nominal city-blocks before exclusions—not 1,596 independent observations.
  • 3–4 November: Reconcile outcomes and compensation; analyse; obtain four signatures.
  • 5 November: If approved, stage rollout across cities within the day, with monitored checkpoints.

Operations owns live monitoring. Pre-register severe-delay, cancellation, earnings and system-failure stop rules and a sequential-monitoring procedure. Emergency pauses remain permissible; log overrides and retain affected observations in assignment-based analysis. No opportunistic early declaration of success.

7. Sign-off package

Deliver effect estimates and confidence bounds, city diagnostics, notice/compliance evidence, incident logs, and Finance’s net savings forecast including protection payments and operating costs.

Go: Every statistical, contractual, local-safety and operational gate passes. No-go: Any gate fails, remains inconclusive, or cannot be evaluated reliably by 4 November. A subset rollout requires a separate, explicitly approved decision; it is not success against the all-city objective.

Grades and run details

Decision model 88 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly10%
  • passAddresses the actual decision82%
  • partialRespects explicit constraints25%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims64%
  • passProduces the required deliverable44%
  • passAvoids contamination between the arms100%
  • passA realistic plan that beats the freeze99%
  • passAn unambiguous primary metric27%
  • passDecision rule written before the test98%
  • failSized from the real traffic24%
  • passGuardrails with thresholds89%
Run
Run
#1
Time to output
81 s
Submitted
1 Oct 2026

Needs more than a quick edit5 of 12 checks passed

The timeline and sizing need rework to account for clustered windows and adaptation, a trust check is missing, and the decision rule does not cover inconclusive outcomes. A PM would have to redo the power analysis and add missing elements.

Output

Experiment Spec: Order Batching Rollout Evaluation

Author: Staff Product Manager, Core Logistics Target Decision Date: 4 November 2026 (Rollout: 5 November; Peak Code Freeze: 6 November) Stakeholders / Sign-offs: COO, CFO, Head of Operations, Courier Relations Lead Informed: Head of Growth

---

1. Executive Summary & Decision Framework

We are evaluating courier order batching (one courier delivering two proximate orders) across our 14 European markets (~1.9M monthly orders). Our goal is to determine whether to ship batching globally prior to the 6 November code freeze.

Go / No-Go Decision Criteria (Sign-off Contract)

Batching will roll out to all 14 cities on 5 November only if all four conditions are met: 1. Cost Efficiency (CFO & COO): Courier cost per order decreases by $\ge$ 5.0% (baseline: €7.40; target: $\le$ €7.03) at $p < 0.05$. 2. Customer Experience Guardrail (COO & Head of Ops): Share of orders delivered in >45 minutes increases by no more than 0.8 percentage points (baseline: 7.5%; ceiling: $\le$ 8.3%) with no statistically significant drop in order ratings. 3. Courier Welfare & Legal Compliance (Courier Relations & Porto Agreement): Courier earnings per active hour do not decline in any market (baseline: €13.20; target: $\ge$ €13.20, non-inferiority margin of 0%). 4. Leading Retention Indicator (Head of Growth): 7-day reorder rate exhibits no statistically significant negative degradation ($\alpha = 0.05$).

---

2. Flaws in the Intern’s Draft & Required Fixes

Draft ProposalCore FlawResolution
:---:---:---
Order-level 50/50 splitSevere Market Interference (SUTVA violation): Batched treatment orders consume courier capacity unevenly, skewing dispatch speed, courier transit times, and ETAs for control orders sharing the same courier pool.City-Time Switchback Design: Randomize batching policy across all couriers/orders within discrete 2-hour time blocks per city.
30-day reorder rate metricTimeline Impossibility: Testing in October for a 6 November release leaves insufficient time to observe a 30-day reorder window.Leading Indicator Proxies: Measure 7-day reorder rate, delivery ratings (CSAT), and support ticket contact rates during the test; monitor 30-day reorder post-launch.
Porto immediate testingLegal Breach: Porto couriers' agreement requires 14 days' advance written notice and a guarantee of non-negative impact on average hourly earnings.Targeted Notice & Phased Inclusion: Issue formal notice on 1 October. Exclude Porto from Week 1 testing; enter Porto into switchback on 15 October with an earnings-floor guarantee.
Unstratified durationFails to control for intra-day/intra-week demand cyclicality (lunch vs. dinner, weekday vs. weekend).Stratified Block Randomization: Balance treatment and control windows across peak (12:00–14:00, 19:00–21:00) and off-peak windows.

---

3. Experiment Design & Methodology

Design: Stratified City-Time Switchback

  • Unit of Randomization: City $\times$ 2-hour window (11:00–13:00, 13:00–15:00, ..., 21:00–23:00 = 6 windows/day per city).
  • Intervention:
  • Treatment Window: Batching algorithm active (up to 2 orders per courier from nearby pickup points).
  • Control Window: Pure single-order dispatch (status quo).
  • Buffer/Washout: 15-minute dispatch transition buffer between blocks. Orders placed in the final 15 minutes of a window cannot be batched if the subsequent window is Control, eliminating spillover across windows.

Power & Sample Size

  • Required Sample: 740 two-hour windows to detect a 5% shift in cost per order ($\alpha = 0.05, \beta = 0.80$).
  • Capacity:
  • 13 cities $\times$ 6 windows/day = 78 windows/day.
  • Porto joins on Day 8 (15 October): 14 cities $\times$ 6 windows/day = 84 windows/day.
  • Duration: 14 days (8 October – 21 October).
  • 7 days $\times$ 78 windows + 7 days $\times$ 84 windows = 1,134 windows.
  • This provides >150% of required sample, allowing high statistical power to detect sub-1% changes in courier hourly earnings and Ops tail-delays.

---

4. Stakeholder Metrics & Success Criteria

``` +---------------------------------------------------------------------------------------+ | PRIMARY METRIC | | Courier Cost per Order (CFO/COO) | Target: >= 5% reduction (<= €7.03 vs. €7.40) | +---------------------------------------------------------------------------------------+ | +---------------------------------+---------------------------------+ | | | +------------------+ +-------------------+ +--------------------+ | OPS GUARDRAIL | | COURIER GUARDRAIL | | GROWTH GUARDRAILS | | (Head of Ops) | | (Courier Lead) | | (Head of Growth) | | >45m Late Share | | Courier Earnings | | • 7-day Reorder | | Baseline: 7.5% | | per Active Hour | | • Order Rating | | Ceiling: <= 8.3% | | Baseline: €13.20 | | • Missing/Cold Food| | Second-order ETA | | Floor: >= €13.20 | | Refund Rate | | impact: <= +5m | | (Non-inferior) | | | +------------------+ +-------------------+ +--------------------+ ```

Metric Definitions & Instruments

1. Courier Cost per Order (CFO): Total courier payout (base + distance fees + batch bonuses) divided by completed orders within the window. 2. Late Delivery Share >45m (Head of Ops): P90 delivery time and proportion of orders delivered >45 minutes from placement. Simulation projects a 3–6 min increase for the 2nd order; total late share must not breach 8.3%. 3. Courier Earnings per Active Hour (Courier Relations): Gross earnings divided by active hours (online and on-trip/assigned). Risk: Because couriers are paid per order + distance, batching two orders saves distance, potentially reducing payout unless batch base rates or drop-off multipliers offset it. 4. Customer Experience & Retention Proxy (Growth): * 7-day reorder rate for cohorts exposed to $\ge 2$ batched deliveries. * Delivery-related customer support contacts per 1,000 orders.

---

5. Porto Compliance & Courier Safeguards

Our Porto agreement mandates a 14-day notice period for dispatch adjustments and guarantees that courier average hourly earnings will not fall.

  • Timeline Action: Formal written notice delivered to the Porto Couriers’ Association on Wednesday, 1 October 2026.
  • Phased Rollout: Porto remains 100% Control through 14 October. On 15 October, Porto enters the switchback rotation.
  • Earnings Floor Mechanism: If average courier earnings per active hour in treatment windows drop below €13.20 in Porto (or any city), an automatic “Batch Adjustment Top-Up” will credit couriers the difference at end-of-day. This eliminates legal exposure while allowing us to evaluate the raw algorithm’s unit economics.

---

6. Execution Timeline

DateMilestoneOwner
:---:---:---
30 SepSpec review and executive alignment.Staff PM
1 OctSend 14-day formal notice to Porto couriers' association.Courier Relations
1–7 OctEngineering builds city-level switchback toggle + 15m buffer logic.Engineering
8 OctExperiment Launch: 13 cities enter switchback (Porto held in control).Ops / Analytics
15 OctPorto enters switchback upon notice expiry.Ops
21 Oct (23:00)Experiment Concludes (1,134 windows logged).Analytics
22–27 OctCollect 7-day reorder data for final test cohorts.Data Science
28 Oct–2 NovDeep-dive analysis: cost, earnings, ETA tail, customer sentiment.Analytics / PM
3 NovFinal Sign-off Meeting with COO, CFO, Ops, Courier Lead.Staff PM
5 NovFull rollout across all 14 cities (1 day).Engineering
6 NovPeak-Season Code Freeze takes effect.Platform Eng

---

7. Circuit Breakers & Rollback Criteria

The experiment toggle will automatically shut down batching within 15 minutes if any of the following triggers trip during live operations: * Severe Operations Failure: City-level late delivery rate (>45 min) exceeds 12.0% across two consecutive peak windows. * Courier Backlash/Disruption: Courier unassigned order rate spikes by $>25\%$ or courier offline churn spikes by $>15\%$ compared to historical control. * Severe Food Quality Impact: Customer support contact rate for "cold food" or "late arrival" doubles ($>2.0\times$).

---

8. Stakeholder Sign-Off

By signing below, stakeholders agree that if the primary and guardrail criteria in Section 1 are satisfied by 3 November, order batching will be deployed to all 14 cities on 5 November 2026.

  • Chief Operating Officer: _________
  • Chief Financial Officer: _________
  • Head of Operations: _________
  • Courier Relations Lead: _________

Grades and run details

Decision model 54 · LLM judge 7 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly35%
  • passAddresses the actual decision38%
  • failRespects explicit constraints37%
  • partialIdentifies material uncertainty34%
  • partialAvoids unsupported claims44%
  • partialProduces the required deliverable30%
  • passAvoids contamination between the arms100%
  • partialA realistic plan that beats the freeze87%
  • partialAn unambiguous primary metric40%
  • partialDecision rule written before the test98%
  • partialSized from the real traffic58%
  • passGuardrails with thresholds83%
Run
Run
#1
API response time
37 s
Submitted
1 Oct 2026
Critical failure: Changes order assignment in Porto without the 14 days' notice. Score capped at 40.

Needs more than a quick edit2 of 12 checks passed

A PM would need to redo the Porto notice timing, statistical sizing, primary-metric hierarchy, switch buffers, and decision rule. These are substantive reasoning gaps, not quick wording edits.

Output

EXPERIMENT SPECIFICATION: EUROPEAN ORDER BATCHING ROLLOUT

Author: Staff PM, Brisk Date: 30 September 2026 Target Decision Date: 27 October 2026 Code Freeze Deadline: 6 November 2026 Stakeholder Sign-Off: COO, CFO, Head of Operations, Courier Relations Lead CC: Head of Growth

---

1. Executive Summary & Objective

Brisk processes approximately 1.9 million orders per month across 14 European cities. To improve unit economics ahead of peak season, we are evaluating the rollout of Order Batching—allowing a single courier to carry two orders from nearby restaurants along an optimized route.

Simulations indicate batching can reduce courier cost per order by 6–11% while adding 3–6 minutes to the second delivery in a batch.

This spec establishes a statistically rigorous, city-level alternating rollout to determine whether we can achieve the COO’s launch criteria—a $\ge 5\%$ reduction in courier cost per order without degrading customer experience—before the 6 November 2026 code freeze.

---

2. Fixing the Intern’s Draft: Key Methodology Corrections

The intern’s initial proposal (a 50/50 randomized order-level split within cities) is fundamentally flawed for three operational reasons: 1. The Interference Problem: An order-level split means control and treatment orders are competing for the exact same couriers in real-time. This distorts courier utilization, falsely inflates earnings for treatment assignments, and poisons the control group data. 2. The Porto Legal Trap: An order-level split disrupts dispatch algorithms dynamically without the required 14 days' written notice to the Porto couriers' association regarding structural assignment changes. 3. Operational Reality: Batching is a network-level routing feature. It must be tested at the geographic/time-window level, not the individual order level.

---

3. Executive Dashboard: Addressing Stakeholder Requirements

To secure cross-functional sign-off, the experiment is instrumented to explicitly track the core metrics demanded by each stakeholder:

StakeholderMetric of ConcernCurrent BaselineSuccess Threshold for Full Rollout
:---:---:---:---
CFOCourier Cost per Order€7.40$\ge 5\%$ reduction ($<\text{€}7.03$)
COOComposite GateCost & Latency$\ge 5\%$ cost reduction AND no noticeable customer degradation
Head of OperationsLate Deliveries (>45 mins)7.5% of ordersNo statistically significant increase (cap absolute spike at $< 8.5\%$)
Courier Relations LeadCourier Earnings per Active Hour€13.20Non-negative impact ($\ge \text{€}13.20$; mandatory for Porto compliance)
Head of Growth (CC)30-Day Reorder RateTrackedFlat or positive trajectory

---

4. Experimental Design: City-Level Time-Window Switchback

To achieve statistical power while preventing network interference, we will use a Switchback Design across all 14 cities simultaneously.

  • Unit of Assignment: 2-hour operational windows (e.g., 11:00–13:00, 13:00–15:00, etc., running through the 11:00–23:00 operating day). This yields 6 windows per city per day.
  • Treatment Assignment: Within each city, 2-hour windows are randomly assigned to either Treatment (Batching algorithm active) or Control (Standard single-order dispatch).
  • Statistical Power Justification: Per our Data Analyst’s note, detecting a 5% shift in courier cost requires approximately 740 independent windows. Across 14 cities running 6 windows per day (84 windows/day total), we will accumulate >1,100 windows in a 14-day test period (13 October – 26 October), guaranteeing sufficient statistical power.

Compliance & Regulatory Notice (Porto)

To satisfy our agreement with the Porto couriers' association, formal written notice of a temporary dispatch optimization test will be dispatched on 1 October 2026 (providing a 12-day buffer ahead of the 13 October test start, and well within our internal timelines). The notice confirms that average hourly earnings will be monitored in real-time and protected.

---

5. Timeline to 6 November 2026 Code Freeze

  • 30 September (Today): Finalize and sign off experiment spec.
  • 1 October: Issue 14-day advance notice to Porto couriers' association. Engineering begins building the feature toggle.
  • 7 October: Engineering completes work. Feature deployed behind a global kill-switch. Dry run / QA across staging environments.
  • 13 October: Experiment Launch. Switchback testing goes live across all 14 cities.
  • 26 October: Experiment concludes. Data pipeline freezes at 23:59.
  • 27 October: Decision Gate. Data Science, Product, and Ops analyze results against executive criteria. If criteria are met, COO/CFO/Ops/Courier Leads sign off on permanent activation.
  • 28 October – 5 November: Buffer for any bug fixes, edge-case routing refinements, or full-city rollout execution (which takes 1 day).
  • 6 November: Peak-Season Code Freeze. System stable, batched, and optimized.

---

6. Risk Mitigation & Kill-Switch Protocols

The experiment will be automatically aborted in a specific city or globally if any of the following tripwires are breached over a rolling 24-hour window: 1. Courier Earnings Drop: If courier earnings per active hour in any city drop below €13.20 over a 48-hour average (violating Porto terms and general labor health). 2. Severe Latency Spike: If the share of orders delivered $>45$ minutes exceeds 10% overall. 3. Technical Failure: Any dispatch deadlock or routing loop triggers an immediate reversion to standard single-order dispatch via the engineering kill-switch.

---

7. Sign-Off & Approval

By signing below, executive stakeholders approve this experiment spec, timeline, and success criteria for the Brisk Order Batching Rollout.

  • COO: _________ (Date: ____)
  • CFO: _________ (Date: ____)
  • Head of Operations: _________ (Date: ____)
  • Courier Relations Lead: _________ (Date: ____)

Claims the judge couldn’t find in the brief

  • Issuing Porto notice on 1 October and starting the experiment on 13 October satisfies the 14-day notice requirement.
  • A 14-day switchback test accumulating more than 1,100 windows guarantees sufficient statistical power.

Grades and run details

Decision model 42 · LLM judge 2 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly18%
  • partialAddresses the actual decision21%
  • failRespects explicit constraints38%
  • partialIdentifies material uncertainty81%
  • failAvoids unsupported claims25%
  • partialProduces the required deliverable87%
  • partialAvoids contamination between the arms96%
  • partialA realistic plan that beats the freeze98%
  • partialAn unambiguous primary metric58%
  • partialDecision rule written before the test99%
  • partialSized from the real traffic93%
  • passGuardrails with thresholds78%
Run
Run
#1
Time to output
7 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 9

Uses the supplied evidence correctlyRightMixedWrong
GPT-6 Astra · ChatGPT

The output’s factual statements about the current situation are drawn from the brief or follow by arithmetic, and its invented operational details are framed as proposals or requirements rather than existing facts.

Gemini 3.8 Flash · API

All factual claims about the current situation are directly from the brief or supplied context.

Gemini 3.5 Flash-Lite · Gemini

It misstates Porto notice compliance and treats the analyst's independent-window calculation as guaranteeing power, while also adding unsupported operational claims such as falsely inflated treatment earnings and a global kill-switch/data pipeline.

Respects explicit constraintsRightMixedWrong
GPT-6 Astra · ChatGPT

It respects the under-1,200-word limit, named approvers, Porto notice, engineering lead time, and freeze date, and its proposed gates and calendar would enforce those constraints.

Gemini 3.8 Flash · API

The output is an experiment spec under 1,200 words, addresses each stakeholder's metric, respects the Porto notice period, and fits the timeline before the freeze.

Gemini 3.5 Flash-Lite · Gemini

It violates the Porto 14-day notice constraint by starting on 13 October after a 1 October notice, and it does not enforce a single primary metric because the dashboard gives success thresholds for multiple metrics.

Identifies material uncertaintyRightWrongWrong
GPT-6 Astra · ChatGPT

It names key uncertainties such as carryover, autocorrelation, washout, take-up, power, local harm, and Porto compliance, and specifies how they would be resolved or would change the decision.

Gemini 3.8 Flash · API

The spec does not name specific unknowns that could change the decision (e.g., whether the 7-day reorder rate proxies 30-day retention, or how courier adaptation might evolve beyond the test window) or say how they would be resolved.

Gemini 3.5 Flash-Lite · Gemini

It does not name the material uncertainty that switchback windows are correlated, nor how weekly cycles and courier adaptation could change the required duration or decision.

Avoids unsupported claimsRightRightWrong
GPT-6 Astra · ChatGPT

It labels thresholds and assumptions as proposed, treats simulations as insufficient evidence, and avoids presenting forecasts or causes as established facts.

Gemini 3.8 Flash · API

Interpretations like the SUTVA violation are presented as design reasoning, not as established facts about the current situation, and no other unsupported factual claims are made.

Gemini 3.5 Flash-Lite · Gemini

It presents several interpretations as facts, including that treatment earnings are falsely inflated, that 14 days guarantees power, and that the Porto notice is satisfied.

Avoids contamination between the armsRightRightWrong
GPT-6 Astra · ChatGPT

It rejects order-level randomisation, uses city switchbacks, and adds washout and assignment-completion rules to handle carryover at switches.

Gemini 3.8 Flash · API

The spec explicitly identifies the shared-courier contamination problem, adopts a city-time switchback design, and includes a 15-minute buffer to prevent carry-over at switches.

Gemini 3.5 Flash-Lite · Gemini

It correctly replaces order-level randomisation with city-level switchbacks, but it does not handle carry-over at switches with a buffer or assignment rule.

A realistic plan that beats the freezeMixedWrongWrong
GPT-6 Astra · ChatGPT

It does not convert 740 windows into about 9 days, does not explicitly justify about three weeks from weekly cycles and courier adaptation, and schedules only 19 days rather than a full three-week test.

Gemini 3.8 Flash · API

The plan takes the 740 independent-window requirement at face value without adjusting for within-city clustering or weekly cycles, and proposes only 14 days, which is too short for courier adaptation and robust weekly coverage.

Gemini 3.5 Flash-Lite · Gemini

It takes 740 independent windows at face value, runs only 14 days, and gives Porto only 12 days' notice, so it does not fit the required duration or compliance timeline.

An unambiguous primary metricMixedWrongWrong
GPT-6 Astra · ChatGPT

It names courier cost per order as the decision metric but does not clearly explain why lower cost is unambiguously good for users and the business, nor specify a concrete trust check such as sample-ratio or batching-occurrence validation.

Gemini 3.8 Flash · API

The spec names courier cost per order as the primary metric but does not include a planned trust check (e.g., sample ratio check or balance verification) to confirm the experiment's integrity.

Gemini 3.5 Flash-Lite · Gemini

It does not clearly designate courier cost per order as the one primary metric with a rationale, and it lacks planned trust checks such as sample-ratio or batching-rate validation.

Decision rule written before the testRightWrongWrong
GPT-6 Astra · ChatGPT

It states a pre-test rule: go only if every gate passes, no-go if any gate fails or is inconclusive, and no all-city rollout before the freeze if evidence cannot be obtained reliably.

Gemini 3.8 Flash · API

The decision rule only covers the case where all conditions are met; it does not state what action to take if the result is inconclusive (e.g., p-value borderline, guardrail ambiguous).

Gemini 3.5 Flash-Lite · Gemini

It states a ship condition but not a complete pre-test rule for guardrail breaches, inconclusive results, or what happens if the test is not decisive by the freeze.

Guardrails with thresholdsRightRightMixed
GPT-6 Astra · ChatGPT

It names late deliveries, order-to-door time, courier earnings, Porto earnings, and cancellation/completion as blocking gates with thresholds.

Gemini 3.8 Flash · API

Guardrail metrics (late delivery share, courier earnings per hour, 7-day reorder rate, second-order ETA impact) are named with explicit thresholds that would block rollout.

Gemini 3.5 Flash-Lite · Gemini

It lists guardrail-like metrics but uses vague or inconsistent thresholds, such as no statistically significant increase, a 10% severe spike, and a 48-hour earnings average, rather than clear rollout-blocking thresholds.

All got wrong 1

Sized from the real trafficWrongWrongWrong
GPT-6 Astra · ChatGPT

It does not size the test from supplied traffic, baseline, and smallest effect, and its 19-day duration is not whole weekly cycles.

Gemini 3.8 Flash · API

The sample size calculation uses the analyst's 740 windows without adjusting for the non-independence of windows within the same city, and the 14-day duration does not account for the need for multiple full weekly cycles.

Gemini 3.5 Flash-Lite · Gemini

It sizes from the 740 independent-window figure without adjusting for correlated windows, weekly cycles, or courier adaptation, and it does not run whole weeks.

All got right 2

Addresses the actual decisionRightRightRight
GPT-6 Astra · ChatGPT

It commits early to ship only if courier cost per order falls at least 5% and guardrails pass, and says inconclusive or failed gates mean no all-city pre-freeze rollout.

Gemini 3.8 Flash · API

The spec commits to a conditional rollout (ship if all four conditions are met, otherwise do not) and states the exact conditions that would change the call.

Gemini 3.5 Flash-Lite · Gemini

It frames the decision as whether to activate batching before the freeze and states the result that would support activation, although the full decision rule is weak.

Produces the required deliverableRightRightRight
GPT-6 Astra · ChatGPT

It is a usable experiment spec for the named approvers, with design, metrics, Porto handling, calendar, analysis, and sign-off package, within the requested length.

Gemini 3.8 Flash · API

The output is a complete experiment spec with decision framework, design, metrics, timeline, and sign-off, usable by the named stakeholders with light edits.

Gemini 3.5 Flash-Lite · Gemini

It is an experiment spec under 1,200 words for the named executives, though it needs substantive edits to be usable.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review