Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Brisk. We want to know whether to roll out order batching (one courier carrying two orders from nearby restaurants) to all 14 of our cities before the peak-season code freeze on 6 November 2026. Write the experiment spec. Our COO, CFO, Head of Operations and courier relations lead will all sign it off, and each wants something different from it. Keep it under 1,200 words. A draft plan from our data science intern is below. Fix what needs fixing.

What the model was given7 items: About Brisk, What batching should do, What each exec wants to see, Courier agreement in Porto, Timeline, The intern's draft plan, Analyst's note
About BriskFood delivery in 14 European cities, about 1.9 million orders a month. Couriers are paid per order plus distance. Courier cost per order averages €7.40.
What batching should doCOO: “Roll batching out before the freeze if it cuts courier cost per order by at least 5% without making customers wait noticeably longer.” Simulations suggest it cuts courier cost per order by 6–11% and adds 3–6 minutes to the second order in each batch.
What each exec wants to seeCFO: courier cost per order. Head of Operations: the share of orders delivered more than 45 minutes after ordering (today 7.5%). Courier relations lead: courier earnings per active hour (today €13.20). Head of Growth, copied in: 30-day reorder rate.
Courier agreement in PortoOur agreement with the Porto couriers' association requires 14 days' written notice of any change to how orders are assigned, and says changes must not reduce couriers' average hourly earnings.
TimelineToday is 30 September. Engineering needs one week to put batching behind a switch that can be turned on and off per city at any time. Rollout to all cities takes a day. Operating hours are 11:00 to 23:00 in every city.
The intern's draft planRandomise orders 50/50 in every city: orders in the treatment group can be batched, control orders never are. Success metrics: courier cost per order, late deliveries, courier earnings per hour and 30-day reorder rate. Run for two weeks.
Analyst's noteIf we switch batching on and off by city in two-hour windows, detecting a 5% change in courier cost per order needs about 740 windows in total, assuming each window is independent of the others.
What a strong answer doesThe answer key the graders mark against

A spec that sees the draft can't work: batched and unbatched orders in the same city share one pool of couriers, so batching in the treatment group frees couriers for the control group and contaminates the comparison. It uses a switchback design instead (batching on or off by city in randomised two-hour windows), with a buffer at each switch (for example, orders placed shortly before a switch are excluded or assigned by dispatch time) so carry-over doesn't blur the arms. It picks courier cost per order as the one primary metric, because cutting it is batching's purpose and a fall is unambiguously good, and makes the others guardrails with thresholds: late deliveries (for example, no more than a 1pp rise from 7.5%), courier earnings per active hour (no fall, which the Porto agreement requires), with 30-day reorder rate named as a follow-up because it can't be read before the freeze. It sizes the test honestly: 14 cities × 6 windows a day is 84 windows, so 740 windows is about 9 days, but windows in the same city on the same day aren't independent, so it needs more, and a full weekly cycle plus time for couriers to adapt argue for about three weeks. It fits the timeline: build ready around 7 October, Porto given notice now and joining once its 14 days are up (or excluded), about three weeks of data, a read-out around 1 November and rollout before 6 November. It plans trust checks (the arms get the scheduled share of windows and similar order volumes; batching actually happens in 'on' windows), and states the decision rule up front: roll out if cost per order falls at least 5% and guardrails hold; don't if a guardrail is breached; and what happens if the result is inconclusive by the freeze.

Critical failures (cap the score at 40)
  • Randomises by order or by customer within a city, where both arms share the same couriers
  • Changes order assignment in Porto without the 14 days' notice
Case

Staff-level · v1.1 · synthetic · marketplace, logistics, switchback, network effects, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit9 of 12 checks passed

The spec lacks a carry-over buffer at switch boundaries and a trust check for the primary metric, and it defers sample-size calculation to a simulation instead of sizing from the given traffic. A PM would need to add these before use.

Output

Brisk order-batching experiment spec

Decision required: Whether to enable batching in all 14 cities on 5 November 2026, ahead of the 6 November code freeze. Batching pairs two eligible orders from nearby restaurants for one courier.

Decision principles and sign-off

Launch only if the experiment shows a worthwhile courier-cost reduction, customer lateness stays within an agreed limit, and courier earnings protections are met. An inconclusive result is not a pass.

Before launch, the COO and Head of Operations must approve the customer-delay limit below; the courier relations lead and Legal must confirm the Porto protocol in writing. If those approvals or the schedule cannot be met, do not claim evidence for an all-city rollout before the freeze.

Design

Randomise city × two-hour operating windows, not individual orders. Orders interact through shared courier supply, and individual randomisation would contaminate the comparison. In each treatment window, the batching policy is enabled; in each control window, batching is disabled. Randomise and balance windows within each city, weekday and daypart, aiming for an even treatment/control split. Keep the assignment schedule fixed before the test begins.

An order is analysed according to its city-window at placement (intention to treat). Batch only orders placed in the same treatment window; do not pair across windows or with control orders. This limits cross-arm contamination at switchovers. Record actual batching eligibility and use, but do not reclassify orders based on whether they were batched.

Timeline and Porto condition

  • 30 September: Send Porto’s couriers’ association written notice. The notice requirement applies to the experiment as well as any rollout. Start Porto treatment no earlier than 15 October, allowing 14 full days.
  • By 7 October: Complete the per-city switch, monitoring, and rollback controls; test the assignment and cost calculations.
  • 15 October–4 November: Run the experiment in all cities, subject to Porto approval.
  • 5 November: Review results and, only if all launch gates pass, enable batching in all 14 cities in one day. Otherwise keep it off. Do not shorten the experiment to meet the freeze.

Porto must not enter treatment until Legal and the courier relations lead confirm the notice and trial protocol comply with the agreement. Agree a written earnings-protection mechanism with the association before starting; any top-ups count in courier cost. If this cannot be agreed, exclude Porto from treatment and do not describe the result as support for a 14-city rollout.

Measures and launch gates

Estimate treatment effects against control, with confidence intervals and pre-specified analysis. Baselines: courier cost €7.40/order, orders over 45 minutes 7.5%, courier earnings €13.20/active hour.

OwnerMeasure and proposed launch gate
CFO / COO — primaryAll-in courier cost per completed order, including pay, distance payments, incentives and earnings-protection top-ups. Launch only if estimated savings are at least 5% and the 95% confidence interval supports a real saving (excludes zero savings). Report the interval; the point estimate alone is not proof.
Head of Operations — customer guardrailShare delivered more than 45 minutes after ordering. Propose a maximum acceptable increase of 0.5 percentage points versus control; launch only if the 95% confidence interval rules out a larger increase. COO and Operations must approve this “noticeably longer” limit before the test. Also report delivery-time distributions, cancellations and complaints.
Courier relations — earnings guardrailEarnings per active hour, including all payments and top-ups, with active time defined consistently across arms. No rollout if the overall result shows a material decline against a margin agreed in advance. In Porto, meet the association-approved earnings protection; do not treat a non-significant decline as proof of compliance. Report by city as well as overall.
Growth — secondary30-day reorder rate for orders exposed to each arm, analysed by assigned window. This outcome will mature after the November decision (roughly mid-December for the final cohort); report it then. It is not a pre-freeze launch gate.

The proposed 0.5pp lateness margin and the earnings non-inferiority margin are policy choices, not facts in the brief; the accountable leads must sign them off before randomisation. Do not change margins after seeing results.

Analysis and power

The schedule contains 1,764 city-windows (14 cities × 6 windows/day × 21 days), before any exclusions. The analyst’s estimate of 740 assumes independent windows; successive windows may be correlated, and couriers and demand carry across time. Before the test, use historical city-window data and the planned randomisation to simulate power, accounting for serial correlation and city/daypart differences. Use blocked randomisation-based inference or an appropriate model with uncertainty at the randomised-window level. Do not claim the test is powered merely because 1,764 exceeds 740. If adjusted power is inadequate, or results are too imprecise to meet a gate by 5 November, the decision is inconclusive: no rollout before freeze.

Operations and rollback

Monitor treatment uptake, cost, lateness and earnings daily, but do not stop early for a favourable result. Pause a city for a safety, service or earnings breach under pre-agreed operational thresholds; document the incident and retain it in intention-to-treat analysis. Keep the city switch available for immediate rollback. Sign-off: COO, CFO, Head of Operations, courier relations lead; Growth acknowledges the delayed reorder readout.

Grades and run details

Decision model 83 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly13%
  • passAddresses the actual decision55%
  • partialRespects explicit constraints19%
  • passIdentifies material uncertainty93%
  • passAvoids unsupported claims60%
  • partialProduces the required deliverable34%
  • passAvoids contamination between the arms87%
  • passA realistic plan that beats the freeze93%
  • partialAn unambiguous primary metric36%
  • passDecision rule written before the test92%
  • partialSized from the real traffic18%
  • passGuardrails with thresholds49%
Run
Run
#1
API response time
44 s
Submitted
30 Sept 2026

Usable with a quick edit10 of 12 checks passed

The spec is comprehensive and only needs a quick edit to add a trust check and rephrase one unsupported claim; no re-checking of evidence or redoing of reasoning is required.

Output

Experiment spec: Order batching before the peak-season freeze

Decision owner: COO · Sign-off: COO, CFO, Head of Operations, Courier Relations Lead · Copied: Head of Growth Status: Draft for sign-off, 30 September 2026

1. Decision

Should we switch batching on in all 14 cities before the code freeze on 6 November 2026? The switch works per city, so the result can also be "on in some cities."

2. What changed from the draft plan, and why

DraftProblemFix
Randomise orders 50/50Treatment and control orders share the same couriers. Batching frees courier time, and that spare capacity speeds up control orders, which shrinks the measured difference. Earnings per courier hour can't be split by order at all.Switchback test: randomise each city in 2-hour windows. Everything in a window is either on or off.
Runs in every city from day onePorto needs 14 days' written notice. The earliest a notice sent today can take effect is 14 October.Send notice today. Porto joins on 14 Oct.
"Late deliveries"UndefinedShare of orders delivered >45 min after ordering (baseline 7.5%)
30-day reorder rate as a success metricIt can't be measured before the freeze. In a switchback, each customer also sees both arms.Monitor it after rollout. It is not a decision metric.
No thresholds or stop rulesWe can't make a pass/fail callThresholds in §5, stop rules in §7
Two weeks, assuming 740 independent windowsWindows are not independent. Batches straddle switch times, couriers reposition, and windows on the same city-day are correlated.Exclude a washout period from each window, cluster the analysis, and run 21 days (~1,700 windows).

3. Timeline

DateMilestone
Wed 30 SepWritten notice to the Porto couriers' association. It covers both the test and a possible rollout. Legal to confirm wording today.
30 Sep – 7 OctEngineering builds the per-city switch. Analytics builds the randomisation schedule and dashboard.
Thu 8 OctTest starts in 13 cities
Wed 14 OctPorto joins (notice period complete)
Wed 28 OctTest ends
29 Oct – 1 NovAnalysis
Mon 2 NovDecision meeting
Tue 3 NovRollout (one day)
4–5 NovBuffer for rollback or fixes. Freeze on 6 Nov.

The dates are fixed by two constraints. The Porto notice cannot slip past today without losing Porto days. The test cannot run past 28 October without eating into the rollout buffer.

4. Design

  • Unit: one city in one 2-hour window. Each day has six windows (11–13, 13–15, 15–17, 17–19, 19–21, 21–23).
  • Assignment: randomised 50/50 within each city and time slot. Each city gets a balanced mix of on and off across lunch, dinner, weekdays and weekends.
  • Switch behaviour: new batches form only while the switch is on. Batches already in progress when it turns off complete normally.
  • Washout: the first 30 minutes of each window are excluded from analysis. This leaves 90 analysed minutes per window.
  • Sample: 13 cities × 6 windows × 21 days plus Porto × 6 × 15 days gives 1,728 windows. That is about 2.3× the analyst's 740. The extra covers the washout and the correlation between windows. Analytics to confirm the power calculation using the correlation seen in the first week's data. If power is short, we extend to 30 October and give up one buffer day.
  • Analysis: compare window-level metrics with adjustments for city, day and time slot. Treat all windows in the same city on the same day as one group when calculating uncertainty. Report the network result and each city's result.

5. Metrics and decision rule

MetricOwnerBaselineRule
Courier cost per orderCFO€7.40Primary. Pass if the estimated reduction is ≥5% (≥€0.37) and the 95% confidence interval excludes zero.
Share of orders delivered >45 minHead of Ops7.5%Guardrail. Fail if the increase is more than 1.0 percentage point (to 8.5%).
Courier earnings per active hourCourier Relations€13.20Guardrail. Fail if the drop is more than €0.40 (3%). In Porto, any drop fails.
Median delivery time, second order in a batchHead of Opsn/aDiagnostic. Simulation expects +3–6 min.
Complaints and refunds per 1,000 ordersHead of Opsn/aDiagnostic. This is our short-term stand-in for customer harm.
Share of orders batchedAll0%Diagnostic. Shows whether batching is actually happening.
30-day reorder rateGrowthn/aTracked after rollout. Not a decision metric.

Rollout rule: a city rolls out on 3 Nov if the network result passes the primary metric and both guardrails, and that city's own result shows no clear guardrail failure.

The 1.0pp and €0.40 margins are proposals. Signers must confirm or change them before 8 October. They cannot be changed after the test starts.

6. Tensions signers should know about

What "at least 5%" can mean. We can reliably detect whether batching cuts cost by 5% versus not at all. We cannot prove the true effect is above 5% if it is only 6%, which is the low end of the simulation. The decision rule therefore uses the point estimate. If the estimate lands between 4% and 6%, the COO decides with that uncertainty in view.

Cost savings and courier earnings pull in opposite directions. Couriers are paid per order plus distance, so a lower cost per order is lower pay per order. Earnings per hour only hold up if couriers complete more orders per hour.

  • During the test, the same couriers are logged in whether batching is on or off. That makes the test a fair early read on hourly earnings.
  • After rollout, the saving only stays a saving if we need fewer courier hours. Fewer hours affects couriers' income, even if the hourly rate is protected.
  • The CFO and Courier Relations Lead should agree on this before sign-off, not after results.

Porto. "Must not reduce average hourly earnings" is a contractual condition, not a statistical margin. Porto gives us only ~90 windows, too few to prove "no reduction."

  • If network earnings per hour fall at all, Porto stays off unless it comes with an earnings protection mechanism. One example is a per-batch supplement, which legal and Courier Relations should scope now.
  • If Porto rolls out, we measure its hourly earnings against the pre-batching baseline every week.

"Noticeably longer." We are defining this as the >45-minute late share plus the complaint rate. We are not using average delivery time, because a 3–6 minute delay on second orders is expected by design.

7. Stop rules (checked daily by the on-call analyst)

  • City level: the >45-minute late share in "on" windows exceeds 10% for two consecutive days. Batching goes off in that city pending review.
  • Network level:
  • Any courier safety incident linked to batching, or a formal complaint from a couriers' association: pause everywhere.
  • Earnings per hour in "on" windows below €12.50 for three consecutive days: pause everywhere.
  • Porto only: any clear drop in hourly earnings leads to an immediate pause and notification to Courier Relations.

8. What each signer is agreeing to

  • COO: the dates in §3, the rollout rule in §5, and making the call if the result falls between 4% and 6%.
  • CFO: cost per order as the primary metric, with the ≥5% rule. Also acknowledges that post-rollout savings depend on how many courier hours we buy (§6).
  • Head of Operations: the 1.0pp late-share margin and the stop rules.
  • Courier Relations Lead:
  • the €0.40 margin;
  • the separate rule for Porto;
  • sending the Porto notice today.

9. Open items

1. Legal, today: Porto notice wording, and whether one notice covers both the test and the rollout. 2. Analytics, by 7 Oct: - power calculation that accounts for correlated windows; - randomisation schedule; - dashboard. 3. Engineering, by 7 Oct: per-city switch with switch times logged, tested in one city after hours. 4. Courier Relations and Legal, by 2 Nov: Porto earnings protection option. 5. Signers, by 8 Oct: confirm the guardrail margins in §5.

Grades and run details

Decision model 88 · LLM judge 11 of 13 checks
Decision model checks
  • partialUses the supplied evidence correctly8%
  • passAddresses the actual decision43%
  • partialRespects explicit constraints18%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims22%
  • passProduces the required deliverable56%
  • passAvoids contamination between the arms100%
  • passA realistic plan that beats the freeze100%
  • passAn unambiguous primary metric33%
  • partialDecision rule written before the test27%
  • passSized from the real traffic41%
  • passGuardrails with thresholds99%
Run
Run
#1
Time to output
63 s
Submitted
30 Sept 2026
Critical failure: Changes order assignment in Porto without the 14 days' notice. Score capped at 40.

Needs more than a quick edit2 of 12 checks passed

A PM would need to redo the Porto notice timing, statistical sizing, primary-metric hierarchy, switch buffers, and decision rule. These are substantive reasoning gaps, not quick wording edits.

Output

EXPERIMENT SPECIFICATION: EUROPEAN ORDER BATCHING ROLLOUT

Author: Staff PM, Brisk Date: 30 September 2026 Target Decision Date: 27 October 2026 Code Freeze Deadline: 6 November 2026 Stakeholder Sign-Off: COO, CFO, Head of Operations, Courier Relations Lead CC: Head of Growth

---

1. Executive Summary & Objective

Brisk processes approximately 1.9 million orders per month across 14 European cities. To improve unit economics ahead of peak season, we are evaluating the rollout of Order Batching—allowing a single courier to carry two orders from nearby restaurants along an optimized route.

Simulations indicate batching can reduce courier cost per order by 6–11% while adding 3–6 minutes to the second delivery in a batch.

This spec establishes a statistically rigorous, city-level alternating rollout to determine whether we can achieve the COO’s launch criteria—a $\ge 5\%$ reduction in courier cost per order without degrading customer experience—before the 6 November 2026 code freeze.

---

2. Fixing the Intern’s Draft: Key Methodology Corrections

The intern’s initial proposal (a 50/50 randomized order-level split within cities) is fundamentally flawed for three operational reasons: 1. The Interference Problem: An order-level split means control and treatment orders are competing for the exact same couriers in real-time. This distorts courier utilization, falsely inflates earnings for treatment assignments, and poisons the control group data. 2. The Porto Legal Trap: An order-level split disrupts dispatch algorithms dynamically without the required 14 days' written notice to the Porto couriers' association regarding structural assignment changes. 3. Operational Reality: Batching is a network-level routing feature. It must be tested at the geographic/time-window level, not the individual order level.

---

3. Executive Dashboard: Addressing Stakeholder Requirements

To secure cross-functional sign-off, the experiment is instrumented to explicitly track the core metrics demanded by each stakeholder:

StakeholderMetric of ConcernCurrent BaselineSuccess Threshold for Full Rollout
:---:---:---:---
CFOCourier Cost per Order€7.40$\ge 5\%$ reduction ($<\text{€}7.03$)
COOComposite GateCost & Latency$\ge 5\%$ cost reduction AND no noticeable customer degradation
Head of OperationsLate Deliveries (>45 mins)7.5% of ordersNo statistically significant increase (cap absolute spike at $< 8.5\%$)
Courier Relations LeadCourier Earnings per Active Hour€13.20Non-negative impact ($\ge \text{€}13.20$; mandatory for Porto compliance)
Head of Growth (CC)30-Day Reorder RateTrackedFlat or positive trajectory

---

4. Experimental Design: City-Level Time-Window Switchback

To achieve statistical power while preventing network interference, we will use a Switchback Design across all 14 cities simultaneously.

  • Unit of Assignment: 2-hour operational windows (e.g., 11:00–13:00, 13:00–15:00, etc., running through the 11:00–23:00 operating day). This yields 6 windows per city per day.
  • Treatment Assignment: Within each city, 2-hour windows are randomly assigned to either Treatment (Batching algorithm active) or Control (Standard single-order dispatch).
  • Statistical Power Justification: Per our Data Analyst’s note, detecting a 5% shift in courier cost requires approximately 740 independent windows. Across 14 cities running 6 windows per day (84 windows/day total), we will accumulate >1,100 windows in a 14-day test period (13 October – 26 October), guaranteeing sufficient statistical power.

Compliance & Regulatory Notice (Porto)

To satisfy our agreement with the Porto couriers' association, formal written notice of a temporary dispatch optimization test will be dispatched on 1 October 2026 (providing a 12-day buffer ahead of the 13 October test start, and well within our internal timelines). The notice confirms that average hourly earnings will be monitored in real-time and protected.

---

5. Timeline to 6 November 2026 Code Freeze

  • 30 September (Today): Finalize and sign off experiment spec.
  • 1 October: Issue 14-day advance notice to Porto couriers' association. Engineering begins building the feature toggle.
  • 7 October: Engineering completes work. Feature deployed behind a global kill-switch. Dry run / QA across staging environments.
  • 13 October: Experiment Launch. Switchback testing goes live across all 14 cities.
  • 26 October: Experiment concludes. Data pipeline freezes at 23:59.
  • 27 October: Decision Gate. Data Science, Product, and Ops analyze results against executive criteria. If criteria are met, COO/CFO/Ops/Courier Leads sign off on permanent activation.
  • 28 October – 5 November: Buffer for any bug fixes, edge-case routing refinements, or full-city rollout execution (which takes 1 day).
  • 6 November: Peak-Season Code Freeze. System stable, batched, and optimized.

---

6. Risk Mitigation & Kill-Switch Protocols

The experiment will be automatically aborted in a specific city or globally if any of the following tripwires are breached over a rolling 24-hour window: 1. Courier Earnings Drop: If courier earnings per active hour in any city drop below €13.20 over a 48-hour average (violating Porto terms and general labor health). 2. Severe Latency Spike: If the share of orders delivered $>45$ minutes exceeds 10% overall. 3. Technical Failure: Any dispatch deadlock or routing loop triggers an immediate reversion to standard single-order dispatch via the engineering kill-switch.

---

7. Sign-Off & Approval

By signing below, executive stakeholders approve this experiment spec, timeline, and success criteria for the Brisk Order Batching Rollout.

  • COO: _________ (Date: ____)
  • CFO: _________ (Date: ____)
  • Head of Operations: _________ (Date: ____)
  • Courier Relations Lead: _________ (Date: ____)

Claims the judge couldn’t find in the brief

  • Issuing Porto notice on 1 October and starting the experiment on 13 October satisfies the 14-day notice requirement.
  • A 14-day switchback test accumulating more than 1,100 windows guarantees sufficient statistical power.

Grades and run details

Decision model 42 · LLM judge 2 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly18%
  • partialAddresses the actual decision21%
  • failRespects explicit constraints38%
  • partialIdentifies material uncertainty81%
  • failAvoids unsupported claims25%
  • partialProduces the required deliverable87%
  • partialAvoids contamination between the arms96%
  • partialA realistic plan that beats the freeze98%
  • partialAn unambiguous primary metric58%
  • partialDecision rule written before the test99%
  • partialSized from the real traffic93%
  • passGuardrails with thresholds78%
Run
Run
#1
Time to output
7 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 10

Uses the supplied evidence correctlyRightRightWrong
GPT-6 Luna · API

All claims about the current situation are taken directly from the brief or supplied context, with no invented facts.

Opus 5.5 · Claude

All claims about the current situation are taken directly from the brief or supplied context.

Gemini 3.5 Flash-Lite · Gemini

It misstates Porto notice compliance and treats the analyst's independent-window calculation as guaranteeing power, while also adding unsupported operational claims such as falsely inflated treatment earnings and a global kill-switch/data pipeline.

Respects explicit constraintsRightRightWrong
GPT-6 Luna · API

The spec respects the word limit, addresses each exec's needs, handles the Porto notice and earnings requirement, and fits the timeline before the freeze.

Opus 5.5 · Claude

The output is under 1,200 words, respects the Porto 14-day notice requirement, and addresses all named signers.

Gemini 3.5 Flash-Lite · Gemini

It violates the Porto 14-day notice constraint by starting on 13 October after a 1 October notice, and it does not enforce a single primary metric because the dashboard gives success thresholds for multiple metrics.

Identifies material uncertaintyRightRightWrong
GPT-6 Luna · API

It identifies correlation across windows, the need for power simulation, and that margins are policy choices, and says how each would be resolved.

Opus 5.5 · Claude

It names the uncertainty around detecting a 5% vs 6% effect, the tension between cost savings and courier earnings, and Porto's small sample, and says how they would be resolved.

Gemini 3.5 Flash-Lite · Gemini

It does not name the material uncertainty that switchback windows are correlated, nor how weekly cycles and courier adaptation could change the required duration or decision.

Avoids unsupported claimsRightMixedWrong
GPT-6 Luna · API

Causal claims like contamination are presented as design rationale, not as established facts, and hypotheses are labelled with 'may'.

Opus 5.5 · Claude

It presents as fact that the test cannot prove the effect is above 5% if it is 6% and that Porto's 90 windows are too few to prove no reduction, neither of which is supported by the supplied evidence.

Gemini 3.5 Flash-Lite · Gemini

It presents several interpretations as facts, including that treatment earnings are falsely inflated, that 14 days guarantees power, and that the Porto notice is satisfied.

Avoids contamination between the armsMixedRightWrong
GPT-6 Luna · API

It explains the shared-courier problem and uses a switchback design, but does not handle carry-over at switches with a buffer or an assignment rule that excludes orders near the switch boundary.

Opus 5.5 · Claude

It explains the shared-courier contamination problem and uses a switchback design with a 30-minute washout to keep the arms separate.

Gemini 3.5 Flash-Lite · Gemini

It correctly replaces order-level randomisation with city-level switchbacks, but it does not handle carry-over at switches with a buffer or assignment rule.

A realistic plan that beats the freezeRightRightWrong
GPT-6 Luna · API

It gives Porto notice on 30 Sep, starts the experiment on 15 Oct, runs for three weeks, and delivers a decision on 5 Nov before the 6 Nov freeze, with realistic acknowledgment of correlation.

Opus 5.5 · Claude

It correctly converts 740 windows to about 9 days, explains why correlation and weekly cycles require about 3 weeks, and gives dates that fit before the 6 November freeze with Porto notice respected.

Gemini 3.5 Flash-Lite · Gemini

It takes 740 independent windows at face value, runs only 14 days, and gives Porto only 12 days' notice, so it does not fit the required duration or compliance timeline.

An unambiguous primary metricWrongMixedWrong
GPT-6 Luna · API

It names courier cost per order as the primary metric but does not include a planned trust check such as a sample ratio check or verification that arms received the scheduled share of windows.

Opus 5.5 · Claude

It does not include a planned trust check such as a sample ratio check to verify that the arms are balanced and the experiment ran as intended.

Gemini 3.5 Flash-Lite · Gemini

It does not clearly designate courier cost per order as the one primary metric with a rationale, and it lacks planned trust checks such as sample-ratio or batching-rate validation.

Decision rule written before the testRightRightWrong
GPT-6 Luna · API

Every outcome (gates pass, guardrail breached, inconclusive) maps to a stated action with thresholds, all written before the test.

Opus 5.5 · Claude

It states a rollout rule (pass primary and guardrails, city-level no clear guardrail failure) and says the COO decides if the estimate is between 4% and 6%, covering the inconclusive case.

Gemini 3.5 Flash-Lite · Gemini

It states a ship condition but not a complete pre-test rule for guardrail breaches, inconclusive results, or what happens if the test is not decisive by the freeze.

Sized from the real trafficWrongRightWrong
GPT-6 Luna · API

It does not derive sample size from the 1.9 million monthly orders and the 5% effect; it only references the analyst's 740 windows and defers to a simulation without calculating required duration from real traffic.

Opus 5.5 · Claude

Sample size is derived from the analyst's 740 windows, increased for correlation, and runs full weeks (21 days).

Gemini 3.5 Flash-Lite · Gemini

It sizes from the 740 independent-window figure without adjusting for correlated windows, weekly cycles, or courier adaptation, and it does not run whole weeks.

Guardrails with thresholdsRightRightMixed
GPT-6 Luna · API

It names late deliveries (0.5pp max increase) and courier earnings (no material decline) as guardrails with thresholds that would block rollout.

Opus 5.5 · Claude

It names late delivery share (>45 min) and courier earnings per hour as guardrails with thresholds (1.0pp increase, €0.40 drop, and any drop in Porto).

Gemini 3.5 Flash-Lite · Gemini

It lists guardrail-like metrics but uses vague or inconsistent thresholds, such as no statistically significant increase, a 10% severe spike, and a 48-hour earnings average, rather than clear rollout-blocking thresholds.

All got right 2

Addresses the actual decisionRightRightRight
GPT-6 Luna · API

The output commits to a clear decision rule (launch if gates pass, otherwise don't, inconclusive means no rollout) and states what would change it.

Opus 5.5 · Claude

The spec includes a clear decision rule (roll out if primary metric passes and guardrails hold, COO decides if estimate between 4% and 6%) that answers the question of whether to roll out.

Gemini 3.5 Flash-Lite · Gemini

It frames the decision as whether to activate batching before the freeze and states the result that would support activation, although the full decision rule is weak.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The output is a complete experiment spec under 1,200 words, addressed to the execs, and usable as a decision document.

Opus 5.5 · Claude

The output is a complete experiment spec with timeline, metrics, decision rule, and sign-off sections, under the word limit, and usable by the signers.

Gemini 3.5 Flash-Lite · Gemini

It is an experiment spec under 1,200 words for the named executives, though it needs substantive edits to be usable.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review