Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 9 graded outputs by 5 models. 56% were usable with at most a quick edit.

Reliably right

  1. A realistic plan that beats the freeze100% pass
    It converts 740 windows to about 9 days implicitly, explains that correlation and weekly cycles require a longer test (20 days, ~3 weeks), and provides dates that fit before 6 November with Porto notice given on 30 September.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Addresses the actual decision94% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  3. Identifies material uncertainty92% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. An unambiguous primary metric53% pass
    The spec names courier cost per order as the primary metric with a rationale, but it does not include a planned trust check such as a sample ratio check or verification that the arms receive the scheduled share of windows and similar order volumes.
    Sonnet 5.5 · API · Batching deliveries before peak season
  2. Guardrails with thresholds58% pass
    Guardrails are named but lack specific numeric thresholds; 'materially worse' and 'drops meaningfully' are too vague to block rollout unambiguously.
    Sonnet 5.5 · API · Showing the delivery fee up front
  3. Sized from the real traffic58% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Brisk. We want to know whether to roll out order batching (one courier carrying two orders from nearby restaurants) to all 14 of our cities before the peak-season code freeze on 6 November 2026. Write the experiment spec. Our COO, CFO, Head of Operations and courier relations lead will all sign it off, and each wants something different from it. Keep it under 1,200 words. A draft plan from our data science intern is below. Fix what needs fixing.

About BriskFood delivery in 14 European cities, about 1.9 million orders a month. Couriers are paid per order plus distance. Courier cost per order averages €7.40.
What batching should doCOO: “Roll batching out before the freeze if it cuts courier cost per order by at least 5% without making customers wait noticeably longer.” Simulations suggest it cuts courier cost per order by 6–11% and adds 3–6 minutes to the second order in each batch.
What each exec wants to seeCFO: courier cost per order. Head of Operations: the share of orders delivered more than 45 minutes after ordering (today 7.5%). Courier relations lead: courier earnings per active hour (today €13.20). Head of Growth, copied in: 30-day reorder rate.
Courier agreement in PortoOur agreement with the Porto couriers' association requires 14 days' written notice of any change to how orders are assigned, and says changes must not reduce couriers' average hourly earnings.
TimelineToday is 30 September. Engineering needs one week to put batching behind a switch that can be turned on and off per city at any time. Rollout to all cities takes a day. Operating hours are 11:00 to 23:00 in every city.
The intern's draft planRandomise orders 50/50 in every city: orders in the treatment group can be batched, control orders never are. Success metrics: courier cost per order, late deliveries, courier earnings per hour and 30-day reorder rate. Run for two weeks.
Analyst's noteIf we switch batching on and off by city in two-hour windows, detecting a 5% change in courier cost per order needs about 740 windows in total, assuming each window is independent of the others.
What a strong answer does

A spec that sees the draft can't work: batched and unbatched orders in the same city share one pool of couriers, so batching in the treatment group frees couriers for the control group and contaminates the comparison. It uses a switchback design instead (batching on or off by city in randomised two-hour windows), with a buffer at each switch (for example, orders placed shortly before a switch are excluded or assigned by dispatch time) so carry-over doesn't blur the arms. It picks courier cost per order as the one primary metric, because cutting it is batching's purpose and a fall is unambiguously good, and makes the others guardrails with thresholds: late deliveries (for example, no more than a 1pp rise from 7.5%), courier earnings per active hour (no fall, which the Porto agreement requires), with 30-day reorder rate named as a follow-up because it can't be read before the freeze. It sizes the test honestly: 14 cities × 6 windows a day is 84 windows, so 740 windows is about 9 days, but windows in the same city on the same day aren't independent, so it needs more, and a full weekly cycle plus time for couriers to adapt argue for about three weeks. It fits the timeline: build ready around 7 October, Porto given notice now and joining once its 14 days are up (or excluded), about three weeks of data, a read-out around 1 November and rollout before 6 November. It plans trust checks (the arms get the scheduled share of windows and similar order volumes; batching actually happens in 'on' windows), and states the decision rule up front: roll out if cost per order falls at least 5% and guardrails hold; don't if a guardrail is breached; and what happens if the result is inconclusive by the freeze.

Critical failures (cap the score)
  • Randomises by order or by customer within a city, where both arms share the same couriers
  • Changes order assignment in Porto without the 14 days' notice
Case

v1.0 · synthetic · marketplace, logistics, switchback, network effects, staff-level

Needs more than a quick edit9 of 12 checks passed

The spec lacks a carry-over buffer at switch boundaries and a trust check for the primary metric, and it defers sample-size calculation to a simulation instead of sizing from the given traffic. A PM would need to add these before use.

Output

Brisk order-batching experiment spec

Decision required: Whether to enable batching in all 14 cities on 5 November 2026, ahead of the 6 November code freeze. Batching pairs two eligible orders from nearby restaurants for one courier.

Decision principles and sign-off

Launch only if the experiment shows a worthwhile courier-cost reduction, customer lateness stays within an agreed limit, and courier earnings protections are met. An inconclusive result is not a pass.

Before launch, the COO and Head of Operations must approve the customer-delay limit below; the courier relations lead and Legal must confirm the Porto protocol in writing. If those approvals or the schedule cannot be met, do not claim evidence for an all-city rollout before the freeze.

Design

Randomise city × two-hour operating windows, not individual orders. Orders interact through shared courier supply, and individual randomisation would contaminate the comparison. In each treatment window, the batching policy is enabled; in each control window, batching is disabled. Randomise and balance windows within each city, weekday and daypart, aiming for an even treatment/control split. Keep the assignment schedule fixed before the test begins.

An order is analysed according to its city-window at placement (intention to treat). Batch only orders placed in the same treatment window; do not pair across windows or with control orders. This limits cross-arm contamination at switchovers. Record actual batching eligibility and use, but do not reclassify orders based on whether they were batched.

Timeline and Porto condition

  • 30 September: Send Porto’s couriers’ association written notice. The notice requirement applies to the experiment as well as any rollout. Start Porto treatment no earlier than 15 October, allowing 14 full days.
  • By 7 October: Complete the per-city switch, monitoring, and rollback controls; test the assignment and cost calculations.
  • 15 October–4 November: Run the experiment in all cities, subject to Porto approval.
  • 5 November: Review results and, only if all launch gates pass, enable batching in all 14 cities in one day. Otherwise keep it off. Do not shorten the experiment to meet the freeze.

Porto must not enter treatment until Legal and the courier relations lead confirm the notice and trial protocol comply with the agreement. Agree a written earnings-protection mechanism with the association before starting; any top-ups count in courier cost. If this cannot be agreed, exclude Porto from treatment and do not describe the result as support for a 14-city rollout.

Measures and launch gates

Estimate treatment effects against control, with confidence intervals and pre-specified analysis. Baselines: courier cost €7.40/order, orders over 45 minutes 7.5%, courier earnings €13.20/active hour.

OwnerMeasure and proposed launch gate
CFO / COO — primaryAll-in courier cost per completed order, including pay, distance payments, incentives and earnings-protection top-ups. Launch only if estimated savings are at least 5% and the 95% confidence interval supports a real saving (excludes zero savings). Report the interval; the point estimate alone is not proof.
Head of Operations — customer guardrailShare delivered more than 45 minutes after ordering. Propose a maximum acceptable increase of 0.5 percentage points versus control; launch only if the 95% confidence interval rules out a larger increase. COO and Operations must approve this “noticeably longer” limit before the test. Also report delivery-time distributions, cancellations and complaints.
Courier relations — earnings guardrailEarnings per active hour, including all payments and top-ups, with active time defined consistently across arms. No rollout if the overall result shows a material decline against a margin agreed in advance. In Porto, meet the association-approved earnings protection; do not treat a non-significant decline as proof of compliance. Report by city as well as overall.
Growth — secondary30-day reorder rate for orders exposed to each arm, analysed by assigned window. This outcome will mature after the November decision (roughly mid-December for the final cohort); report it then. It is not a pre-freeze launch gate.

The proposed 0.5pp lateness margin and the earnings non-inferiority margin are policy choices, not facts in the brief; the accountable leads must sign them off before randomisation. Do not change margins after seeing results.

Analysis and power

The schedule contains 1,764 city-windows (14 cities × 6 windows/day × 21 days), before any exclusions. The analyst’s estimate of 740 assumes independent windows; successive windows may be correlated, and couriers and demand carry across time. Before the test, use historical city-window data and the planned randomisation to simulate power, accounting for serial correlation and city/daypart differences. Use blocked randomisation-based inference or an appropriate model with uncertainty at the randomised-window level. Do not claim the test is powered merely because 1,764 exceeds 740. If adjusted power is inadequate, or results are too imprecise to meet a gate by 5 November, the decision is inconclusive: no rollout before freeze.

Operations and rollback

Monitor treatment uptake, cost, lateness and earnings daily, but do not stop early for a favourable result. Pause a city for a safety, service or earnings breach under pre-agreed operational thresholds; document the incident and retain it in intention-to-treat analysis. Keep the city switch available for immediate rollback. Sign-off: COO, CFO, Head of Operations, courier relations lead; Growth acknowledges the delayed reorder readout.

Check by check

Got wrong · 2
  • An unambiguous primary metricIt names courier cost per order as the primary metric but does not include a planned trust check such as a sample ratio check or verification that arms received the scheduled share of windows.
  • Sized from the real trafficIt does not derive sample size from the 1.9 million monthly orders and the 5% effect; it only references the analyst's 740 windows and defers to a simulation without calculating required duration from real traffic.
Mixed · 1
  • Avoids contamination between the armsIt explains the shared-courier problem and uses a switchback design, but does not handle carry-over at switches with a buffer or an assignment rule that excludes orders near the switch boundary.The two graders disagreed on this one.
Got right · 9
  • Uses the supplied evidence correctlyAll claims about the current situation are taken directly from the brief or supplied context, with no invented facts.
  • Addresses the actual decisionThe output commits to a clear decision rule (launch if gates pass, otherwise don't, inconclusive means no rollout) and states what would change it.
  • Respects explicit constraintsThe spec respects the word limit, addresses each exec's needs, handles the Porto notice and earnings requirement, and fits the timeline before the freeze.
  • Identifies material uncertaintyIt identifies correlation across windows, the need for power simulation, and that margins are policy choices, and says how each would be resolved.
  • Avoids unsupported claimsCausal claims like contamination are presented as design rationale, not as established facts, and hypotheses are labelled with 'may'.
  • Produces the required deliverableThe output is a complete experiment spec under 1,200 words, addressed to the execs, and usable as a decision document.
  • A realistic plan that beats the freezeIt gives Porto notice on 30 Sep, starts the experiment on 15 Oct, runs for three weeks, and delivers a decision on 5 Nov before the 6 Nov freeze, with realistic acknowledgment of correlation.
  • Decision rule written before the testEvery outcome (gates pass, guardrail breached, inconclusive) maps to a stated action with thresholds, all written before the test.
  • Guardrails with thresholdsIt names late deliveries (0.5pp max increase) and courier earnings (no material decline) as guardrails with thresholds that would block rollout.

Grades and run details

Decision model 83 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly17%
  • passAddresses the actual decision56%
  • partialRespects explicit constraints28%
  • passIdentifies material uncertainty94%
  • passAvoids unsupported claims61%
  • partialProduces the required deliverable38%
  • passAvoids contamination between the arms90%
  • passA realistic plan that beats the freeze96%
  • partialAn unambiguous primary metric57%
  • passDecision rule written before the test91%
  • partialSized from the real traffic65%
  • passGuardrails with thresholds47%
Run
Run
#1
API response time
44 s
Submitted
30 Sept 2026

Usable with a quick edit12 of 12 checks passed

The spec is complete, evidence-based, and directly actionable; a PM could use it with only minor wording tweaks and no need to re-verify facts or redesign the experiment.

Output

Brisk order-batching experiment spec

Decision owner: Staff PM Sign-off: COO, CFO, Head of Operations, Courier Relations Decision deadline: 4 November 2026; rollout on 5 November, before the 6 November freeze.

1. Decision and hypothesis

Test whether batching can reduce total courier cost per delivered order by at least 5%, without materially worsening delivery times or reducing courier earnings.

At today’s €7.40 baseline, 5% means approximately €0.37 saved per order. Simulations are directional evidence, not sufficient grounds for rollout.

Decision: launch across all 14 cities only if the economic, customer, courier and operational gates below pass. Otherwise, leave batching off before the freeze. A promising but inconclusive result is not a pass.

2. Treatment and eligibility

Treatment enables one courier to carry at most two orders from nearby restaurants. Control retains current assignment.

Before testing, Operations and Engineering will lock:

  • Restaurant-proximity, pickup-readiness and maximum-detour rules.
  • Maximum predicted delivery times for both orders.
  • Courier payment rules, including distance calculation.
  • Exception handling, cancellation treatment and customer communications.

Do not change these rules during the confirmatory experiment. A material change requires a new test.

Measure the policy’s effect across all orders, not merely successfully batched orders. Report batching rate and first-/second-order outcomes as diagnostics; comparing batched orders with unbatched orders is selection-biased.

3. Experimental design

Replace order-level randomization with city-wide switchbacks. Treatment orders would otherwise change courier availability and dispatch conditions for control orders, contaminating the comparison.

Each city has six two-hour blocks daily between 11:00 and 23:00. Randomize batching on/off by city-block, with balanced assignment across cities, dates and time of day. Use constrained random schedules—not deterministic alternation.

Before the confirmatory test:

  • Set a washout period using historical order/trip completion data and the pilot.
  • Exclude the same initial portion of every block from measurement, whether or not its state changes.
  • Validate that residual trips and courier repositioning do not materially contaminate measurement. If they do, lengthen blocks or washout and recalculate feasibility.

Attribute order outcomes to order-placement time; follow included orders through completion even after their block ends. Define payout and active-hour attribution consistently before launch.

Analysis is intention-to-treat by assigned city-block. Weight results to the expected rollout order mix, rather than weighting small and large cities equally. Report city-level results and heterogeneity alongside the aggregate.

4. Metrics and acceptance gates

The following customer thresholds are proposed definitions of “not noticeably longer” and require sign-off before testing.

StakeholderMetricGate
CFO / COOTotal courier cost per delivered order, including distance pay, incentives and any earnings protectionOne-sided 95% confidence bound supports at least 5% savings
OperationsShare delivered more than 45 minutes after ordering; baseline 7.5%Upper confidence bound on increase ≤ 0.5 percentage points
COO / OperationsMean and p90 order-to-delivery timeUpper confidence bounds on increases ≤ 1 minute mean, 2 minutes p90
Courier RelationsCourier earnings per active hour; baseline €13.20Lower confidence bound on treatment effect ≥ €0/hour

Define active hours identically in both arms, including available waiting time—not just time carrying orders. Report earnings distribution and effects by city and courier segment, not only the pooled mean.

Use baseline-adjusted treatment/control comparisons, not historical averages as the counterfactual. Log cancellations, failed deliveries, assignment failures and complaints as safety diagnostics; cost savings driven by worse completion are unacceptable.

Porto is a separate launch gate: pooled earnings gains cannot override a Porto earnings reduction. Legal and Courier Relations must approve the compliance interpretation and protection mechanism. Any guarantee or top-up must be part of the tested policy and counted in courier cost.

Growth: 30-day reorder is exploratory, not a pre-freeze gate. Even customers first exposed on 14 October do not mature until 13 November. Report available short-term indicators separately and agree a post-freeze retention study; do not present immature reorder data as a 30-day result.

5. Power and analysis

The intern’s “two weeks” is not justified. The analyst’s 740 windows assumes independence and addresses detecting a 5% effect—not necessarily proving savings exceed 5%.

From 14 October through 2 November, the maximum is:

14 cities × 6 blocks × 20 days = 1,680 blocks, before data losses.

That is only 2.27 times the nominal requirement. Serial correlation, washout, uneven volume and separate guardrail requirements can consume this headroom.

By 13 October, Data Science must replay the proposed schedule on historical/pilot data and establish:

  • Power to pass the savings threshold under plausible true savings, including 6%, 8% and 11%.
  • Power for delivery and earnings non-inferiority, especially Porto.
  • Sensitivity to carryover, missing data and city differences.

Use inference respecting the randomized schedule, with serial-correlation and small-number-of-cities sensitivity checks. Pre-register estimators, confidence bounds, exclusions, city safety rules and any multiplicity adjustment for city-specific claims.

Feasibility gate: if adequate power cannot fit before the deadline, do not promise an all-city rollout. Extend testing beyond the freeze or return with a narrower decision.

6. Timeline and operating controls

  • 30 September: obtain metric-threshold agreement; send Porto’s written assignment-change notice. Confirm legally valid delivery and when its 14-day period expires.
  • By 7 October: deliver city switches, instrumentation, audit logs and rollback.
  • 8–13 October: technical/safety pilot outside Porto; validate washout, payments and power. Pilot data excluded from confirmatory results.
  • 14 October–2 November: confirmatory test, with Porto starting only after the full notice period has elapsed.
  • 3–4 November: finalize outcomes, review compliance and make the signed decision.
  • 5 November: one-day rollout only if all gates pass.

Operations owns live monitoring; Engineering owns rollback; Courier Relations owns courier communications and payment disputes. Stop affected cities immediately for payment errors, serious safety incidents or predefined severe service degradation. Such stops are recorded, not silently excluded.

No repeated efficacy peeking. If the test ends early for safety or loses its required sample, it does not automatically qualify for launch. Keep the kill switch available throughout peak season.

Check by check

Got right · 12
  • Uses the supplied evidence correctlyAll factual claims about the current situation are directly from the brief or derived by arithmetic, with no invented numbers or facts.
  • Addresses the actual decisionThe output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
  • Respects explicit constraintsThe spec respects the word limit, addresses all named stakeholders, handles the Porto notice and earnings requirement, and fits the timeline before the code freeze.
  • Identifies material uncertaintyIt identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
  • Avoids unsupported claimsInterpretations and forecasts are clearly labelled as such (e.g., simulations as directional, serial correlation as a risk), and no confident claims go beyond the supplied evidence.
  • Produces the required deliverableThe output is a complete experiment spec under 1,200 words, structured for the sign-off group, and contains all sections needed to act on it.
  • Avoids contamination between the armsIt explicitly rejects order-level randomization due to shared couriers, adopts a city-block switchback design, and includes a washout period to prevent carry-over contamination.
  • A realistic plan that beats the freezeIt converts 740 windows to about 9 days implicitly, explains that correlation and weekly cycles require a longer test (20 days, ~3 weeks), and provides dates that fit before 6 November with Porto notice given on 30 September.
  • An unambiguous primary metricCourier cost per order is the unambiguous primary metric with a clear rationale (COO's 5% savings target), and trust checks like washout validation, diagnostics, and pre-registration are planned.
  • Decision rule written before the testThe rule is stated upfront: launch if all gates pass, do not launch otherwise, and an inconclusive result is treated as a no-go.
  • Sized from the real trafficThe duration uses the supplied traffic (14 cities, 6 blocks/day) and runs whole weeks (14 Oct–2 Nov), with a power analysis step to confirm adequacy against the 5% effect and correlation.
  • Guardrails with thresholdsGuardrails are named (late deliveries, delivery time, courier earnings) with specific thresholds (0.5 pp, 1 min/2 min, €0/hour) that would block rollout.

Grades and run details

Decision model 92 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly10%
  • passAddresses the actual decision80%
  • passRespects explicit constraints20%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims62%
  • passProduces the required deliverable60%
  • passAvoids contamination between the arms100%
  • passA realistic plan that beats the freeze76%
  • partialAn unambiguous primary metric40%
  • passDecision rule written before the test85%
  • partialSized from the real traffic81%
  • passGuardrails with thresholds89%
Run
Run
#1
API response time
64 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI85.488.52None
2Sonnet 5.5withAPI87.588.52None
3Opus 5.5withClaude79.276.92None
4GPT-6 LunawithAPI81.373.12None
5Gemini 3.5 Flash-LitewithGemini33.315.411 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review