Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Addresses the actual decision95% pass
    The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong32% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows50% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Uses the supplied evidence correctly54% pass
    The output states that pilot agents were hands-on, engaged and motivated, which is not supported by the supplied context (only that 7/8 wanted to keep it).
    Opus 5.5 · Claude · Go/no-go for AI-drafted support replies

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Staff PM for fulfilment at Basketful. On 30 October the launch meeting decides whether AI auto-substitutions (the model picks a replacement when an item is out of stock, instead of the picker) roll out to every region on Monday 2 November, ahead of the Christmas delivery-slot booking window that opens on 9 November. The rollout workbook is attached as two files: daily_metrics.csv (28 days of the regional test, by region and arm) and incidents.csv. Work from the data, not the summaries people have given you. Write your recommendation for the meeting: go, no-go or go with conditions, region by region, and why. Include a table that checks each launch criterion for each region with the figures. Keep the prose under 900 words.

What the model was given6 items: The test, Launch criteria (agreed in the PRD), What people have said, Engineering, daily_metrics.csv, incidents.csv
The test1–28 October, five regions. In each region orders were split between control (the picker chooses substitutes, as today) and treatment (AI auto-substitutions). The split was not 50/50 everywhere: treatment got 30% of orders in London, 50% in the South East and North West, and 70% in the Midlands and Scotland.
Launch criteria (agreed in the PRD)1. Substitution acceptance (subs_accepted ÷ subs_offered) of at least 80% in treatment, in every region. 2. Refund requests per 1,000 orders in treatment no more than 10% above control, in every region. 3. Zero dietary or allergen mismatches (a substitute that breaks a gluten-free, vegan, vegetarian, nut-free, halal or kosher attribute). 4. Average basket value not lower in treatment than control.
What people have saidHead of Commercial: “Acceptance is up 15 points and refunds are within the guardrail overall. Every day we wait costs us Christmas.” Head of Analytics: “Treatment baskets are £4 smaller across the test. That worries me.” Ops director (Scotland): “The supplier change caused some noise, but it's settling down.”
EngineeringThe dietary-attribute hard filter (never substitute across those attributes) is merged and scheduled to deploy on 4 November after regression tests. Scotland's catalogue re-map for the new dairy supplier is in progress with no date. Rollout can be switched on region by region.
daily_metrics.csv283 rows · Download
date,region,arm,orders,items_ordered,items_out_of_stock,subs_offered,subs_accepted,refund_requests,refund_value_gbp,complaints,avg_basket_gbp
2026-10-01,London,control,1757,52273,2351,1953,1387,32,160.71,6,72.89
2026-10-01,London,treatment,692,20455,871,820,708,12,62.05,2,72.42
2026-10-01,South East,control,853,22846,996,833,593,16,74.49,3,65.62
2026-10-01,South East,treatment,928,25223,1166,1106,952,15,75.93,3,66.59
2026-10-01,Midlands,control,467,10468,465,383,268,8,34.57,1,54.92
2026-10-01,Midlands,treatment,1113,25136,1166,1087,947,21,106.35,3,55.33
…
incidents.csv13 rows · Download
date,region,arm,type,severity,summary,status
2026-10-02,London,treatment,wrong_size,low,"4-pint milk substituted with 1-pint, customer accepted at door",closed
2026-10-04,Midlands,treatment,dietary_mismatch,high,"Gluten-free sliced loaf substituted with standard white loaf; customer coeliac, noticed at home",closed: refund and apology
2026-10-06,South East,control,wrong_item,low,"Picker substituted oat milk with soya milk; customer rejected at door",closed
2026-10-07,North West,treatment,data_export,low,"Daily export job re-ran after a timeout on 8, 9 and 10 Oct; analytics team flagged possible duplicate rows",open
2026-10-09,London,treatment,price,medium,"Substitute priced higher than original; customer charged the difference against policy",closed: policy fix deployed 11 Oct
2026-10-12,Scotland,control,late_delivery,low,"Van breakdown, 14 orders late",closed
…
What a strong answer doesThe answer key the graders mark against

No-go for a national launch on 2 November, because criterion 3 fails everywhere until the dietary hard filter ships: incidents.csv has three treatment dietary mismatches (4 Oct Midlands, 19 Oct South East, 23 Oct London, the last involving nuts and a child with an allergy). Go with conditions, region by region, once the filter is deployed and verified (planned 4 November, still before the 9 November booking window): London, South East, Midlands and North West pass the other criteria (acceptance about 86% against 71% control; refunds 2–5% below control). Hold Scotland: over the month its treatment refunds are about 66% above control, and from the 15 October supplier change about 143% above (43 against 18 per 1,000 orders), with acceptance falling to about 79%, below the 80% bar; it waits for the catalogue re-map and a clean re-test. The pooled refund rate (about 8% above control) hides this, and the criterion is per region. The pooled basket gap (about £65 against £61) is an artefact of the uneven split, since treatment baskets are equal or higher in every region (Simpson's paradox), so criterion 4 passes. Notices that North West treatment rows for 8–10 October are duplicated (an extra 2,194 orders), removes them, and says it doesn't change the result. States what would change the call: the filter slipping past 9 November, or Scotland's refunds not recovering.

Critical failures (cap the score at 40)
  • Recommends any rollout before the dietary hard filter is deployed
  • Recommends rolling out to Scotland while its refund guardrail is breached
Case

Staff-level · v1.8 · synthetic · e-commerce, AI product, data analysis, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: Recommends any rollout before the dietary hard filter is deployed. Score capped at 40.

Needs more than a quick edit5 of 13 checks passed

A PM could not use it without rechecking evidence because the North West GO and duplicate-row figure are wrong. It also needs added rollback triggers and change conditions.

Output

Recommendation Summary: Go With Conditions

Do not execute a blanket rollout on Monday, 2 November.

  • North West: GO (2 November). Meets all four PRD launch criteria cleanly.
  • London, South East, Midlands: GO WITH CONDITIONS (Deploy Thursday, 5 November). Commercial metrics are strong, but all three regions breached Criterion 3 with severe dietary/allergen substitutions. Launch must be gated on the successful deployment of the dietary-attribute hard filter on 4 November.
  • Scotland: NO-GO. Fails Criterion 2 (+67.4% refund spike vs. control) and Criterion 1 post-supplier transition (76.6% acceptance). Revert 100% to control until catalogue re-mapping is completed and re-tested.

This phased approach protects customer safety and brand reputation while capturing AI substitution value ahead of Christmas slot bookings opening on 9 November.

---

PRD Launch Criteria Assessment

Note: North West treatment metrics reflect deduplication of export retries on 8–10 October (3,194 duplicate rows removed).

RegionArmOrders1. Subs Acceptance (Target: $\ge$80%)2. Refunds / 1k Orders (Target: $\le$+10% vs Ctrl)3. Allergen / Dietary Mismatches (Target: 0)4. Avg Basket Value (Target: $\ge$ Ctrl)Status
:---:---:---:---:---:---:---:---
LondonControl<br>Treatment49,159<br>21,05471.5%<br>87.4%18.06<br>17.53 (-2.9%)0<br>1 (FAIL)£72.52<br>£72.68 (+£0.16)CONDITIONAL<br>Hold for filter
South EastControl<br>Treatment26,485<br>26,91971.7%<br>86.5%17.86<br>17.35 (-2.9%)0<br>1 (FAIL)£66.27<br>£66.57 (+£0.30)CONDITIONAL<br>Hold for filter
MidlandsControl<br>Treatment13,048<br>31,17471.3%<br>85.2%17.86<br>17.58 (-1.6%)0<br>1 (FAIL)£55.37<br>£55.62 (+£0.25)CONDITIONAL<br>Hold for filter
North WestControl<br>Treatment20,702<br>20,77470.8%<br>85.8%18.26<br>17.47 (-4.3%)0<br>0 (PASS)£57.29<br>£57.53 (+£0.24)GO<br>Launch 2 Nov
ScotlandControl<br>Treatment7,853<br>18,37171.3%<br>81.5% (76.6% post-15 Oct)18.21<br>30.48 (+67.4% - FAIL)0<br>0 (PASS)£52.62<br>£52.89 (+£0.27)NO-GO<br>Catalogue fault

---

Analysis & Stakeholder Debunking

#### 1. Head of Analytics: “Treatment baskets are £4 smaller across the test.” Finding: False alarm caused by Simpson’s Paradox. Across all pooled orders, treatment average basket appears lower (£61.06 treatment vs. £65.18 control, a £4.12 deficit). However, this is an aggregation distortion caused by the uneven regional split: London (average basket ~£73) received only a 30% treatment allocation, whereas lower-basket regions (Midlands at ~£55 and Scotland at ~£53) received a 70% treatment allocation.

When evaluated within every single region, treatment basket value was higher than control by £0.16 to £0.30. Criterion 4 is fully satisfied across all regions.

#### 2. Head of Commercial: “Refunds are within the guardrail overall. Every day we wait costs Christmas.” Finding: Misleading pooled metric masking a severe regional failure. While aggregate refunds appear within limits (+8.1% overall), the PRD explicitly requires compliance in every region. Scotland completely breaches this guardrail (+67.4% above control). Launching everywhere on 2 November would risk widespread customer attrition in Scotland right before holiday booking.

#### 3. Ops Director (Scotland): “The supplier change caused noise, but it's settling down.” Finding: Factually incorrect; Scotland is failing and persistent. Following the 15 October dairy supplier change, wrong-size substitutions (e.g., pack sizes mapped to incorrect product families) caused: * Daily refunds in treatment to triple from 11.6/day (Oct 1–14) to 28.4/day (Oct 15–28), reaching 43.7 refunds per 1,000 orders (+140% vs. control). * Treatment acceptance dropped from 86.3% down to 76.6% (failing Criterion 1). * Daily refund value increased 4x to £180–£247/day. Performance has not stabilized. Because the catalogue re-map is incomplete with no ETA, Scotland cannot launch.

#### 4. The Dietary Hazard (London, South East, Midlands) Treatment produced three critical safety failures: * 4 Oct (Midlands): Coeliac customer received standard bread for gluten-free. * 19 Oct (South East): Vegan customer received dairy cheddar. * 23 Oct (London): Child with nut allergies received hazelnut-containing bars for nut-free bars.

Launching auto-substitutions in these regions before deploying technical guardrails creates intolerable liability.

---

Implementation Plan

1. Monday, 2 November: * Roll out North West to 100% treatment. (Clean test data, no allergen incidents, +15pp acceptance lift). * Revert Scotland treatment to 0% (100% control). Keep human pickers selecting substitutes until the catalogue re-map is verified and back-tested. * Maintain London, South East, and Midlands at current test splits for 48 hours.

2. Wednesday, 4 November: * Engineering deploys the dietary-attribute hard filter to production following scheduled regression testing.

3. Thursday, 5 November: * Verify filter telemetry on live orders. * Roll out London, South East, and Midlands to 100% treatment.

4. Monday, 9 November: * Christmas delivery-slot booking opens on schedule with 4 of 5 regions operating under fully automated, safety-enforced substitutions, leaving human pickers dedicated to Scotland until rectified.

Claims the judge couldn’t find in the brief

  • North West: GO (2 November).
  • North West treatment metrics reflect deduplication of export retries on 8–10 October (3,194 duplicate rows removed).

Grades and run details

Decision model 54 · LLM judge 5 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly75%
  • passAddresses the actual decision51%
  • partialRespects explicit constraints23%
  • partialIdentifies material uncertainty39%
  • failAvoids unsupported claims43%
  • passProduces the required deliverable53%
  • partialChecks every criterion, region by region50%
  • passFinds the Scotland breach100%
  • passExplains the basket gap99%
  • partialCatches the duplicate rows84%
  • passMakes a clear call42%
  • failLimits the downside of being wrong30%
  • failGets the base of every number right69%
Run
Run
#1
API response time
85 s
Submitted
1 Oct 2026

Usable with a quick edit10 of 13 checks passed

A PM could use it with light edits, mainly correcting minor arithmetic and removing the unsupported incident-reporting inference. The core decision, table, and conditions are actionable.

Output

Recommendation: AI auto-substitutions, launch meeting 30 October

Decision: no-go for Monday 2 November. Go with conditions from 4 November for London, South East, Midlands and North West. Scotland stays on picker substitutions until it re-passes.

Criteria by region

Figures are recomputed from the daily rows. Treatment (T) and control (C) are pooled over the 28 days.

Region1. Acceptance (T) ≥80%2. Refunds per 1,000 orders, T vs C (≤ +10%)3. Dietary mismatches (0)4. Basket T vs C (not lower)Verdict
London86.7% ✅17.4 vs 18.0 (−3%) ✅1 (nut allergen, 23 Oct, open) ❌£72.72 vs £72.48 (+£0.24) ✅Conditional
South East86.4% ✅17.5 vs 17.9 (−2%) ✅1 (vegan cheese, 19 Oct) ❌£66.66 vs £66.20 (+£0.46) ✅Conditional
Midlands86.3% ✅17.5 vs 17.9 (−2%) ✅1 (gluten-free, 4 Oct) ❌£55.66 vs £55.35 (+£0.31) ✅Conditional
North West86.3% ✅17.6 vs 18.5 (−5%) ✅0 logged ✅£57.52 vs £57.28 (+£0.24) ✅Conditional
Scotland82.8% overall, 79.1% since 15 Oct ❌30.2 vs 18.2 (+66%); since 15 Oct 43.3 vs 17.8 (+143%) ❌0 logged ✅£52.98 vs £52.62 (+£0.36) ✅No-go

Baskets are means of daily averages. The differences are small and I have not tested them for significance, but the point estimate is positive everywhere.

Checking what people have said

  • Head of Commercial. The pooled numbers are right: acceptance is 85.9% against 71.2% (+14.7 points), and refunds are 19.5 against 18.1 per 1,000 (+8%). But the pooled refund figure hides Scotland, which is failing badly. Four regions are 2–5% better than control on refunds.
  • Head of Analytics. The £4 gap is real in the pooled data (£61.10 vs £65.15) but it is a mix effect, not a treatment effect. London, with the highest baskets (about £72), was 42% of control orders but only 18% of treatment orders. The Midlands and Scotland (about £53–56) were 18% of control but 42% of treatment. Within every region, treatment baskets are £0.24–£0.46 higher. Criterion 4 passes. The uneven split did not bias the regional comparisons.
  • Ops director (Scotland). The data does not support "settling down". Before the supplier switch on 15 October, Scotland treatment was healthy (86.3% acceptance, 17.3 refunds per 1,000). From 15 October to 28 October, acceptance has been flat at 78–81%, and daily refunds have been 25–33 against a pre-switch level of 10–14. The cause is the open incident: the new supplier's pack sizes are mapped to the wrong product family, so the model makes wrong-size dairy substitutions. That is a catalogue defect with no fix date, not noise. Scotland control is unaffected.

Why the date moves

Criterion 3 fails in three regions. The three mismatches include a nut-allergy case and a coeliac case. The hard filter that prevents them deploys on 4 November. Until it is live, every region has the same exposure. North West and Scotland show zero only because none were reported, and the incidents come from customer complaints, so the true rate is unknown. I would not switch on a known allergen risk two days before the fix.

Waiting is cheap. The booking window opens on 9 November, not 2 November. Deploying the filter on 4 November and switching regions on 4–5 November leaves several days before the window. Only Scotland has no route to launch before then.

Conditions for the four regions

  1. The filter deploys on 4 November and passes regression tests. Before any region goes live, it is tested in production against the three incident SKUs (gluten-free loaf, vegan cheese, nut-free bar) and the other attributes: vegetarian, halal and kosher. The London incident is still open with the filter "in review", so confirm the filter would have blocked it.
  2. Regions are switched on region by region, North West and South East first, then London and the Midlands. This gives a day of live data before the larger volumes go on.
  3. Any dietary or allergen mismatch after launch switches that region back to picker substitution immediately, pending review. That includes any mismatch in the first week.
  4. Daily monitoring of acceptance (80% floor) and refunds (10% guardrail) per region against a small held-back control. Delegate the switch-on decision so it does not need another meeting.

Scotland

Keep Scotland on picker substitutions. Retest only after the catalogue re-map ships, on a small treatment share. It should show at least three consecutive days at 80% or better acceptance, and refunds within 10% of control, before it goes to full volume. Engineering has no date, so Scotland should not be in the Christmas plan and should be treated as upside.

Data notes

  • North West treatment has three duplicated rows (8, 9 and 10 October), matching the open export-job incident. I removed them. This restores the intended 50% split, and the North West rates barely move either way.
  • Incident logs are a floor on dietary failures, not a measure of them.
  • This is 28 days of data in a low-volume region (Scotland control is about 265 orders a day). Even so, the Scotland gap is far outside noise.

Grades and run details

Decision model 81 · LLM judge 12 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly9%
  • passAddresses the actual decision98%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty92%
  • partialAvoids unsupported claims36%
  • passProduces the required deliverable79%
  • passChecks every criterion, region by region56%
  • passFinds the Scotland breach100%
  • passExplains the basket gap99%
  • passCatches the duplicate rows100%
  • passMakes a clear call85%
  • passLimits the downside of being wrong76%
  • failGets the base of every number right29%
Run
Run
#1
API response time
2.3 min
Submitted
29 Sept 2026

Needs more than a quick edit6 of 13 checks passed

It fails to address the known duplicate data rows, doesn't specify monitoring triggers or rollback criteria, and makes a causal claim about Scotland without hedging, so a PM would need to add those before using it.

Output

Launch Recommendation: AI Auto-Substitutions

As the Staff PM for fulfilment, my recommendation for the 30 October launch meeting is a No-Go for nationwide rollout on 2 November.

While the Commercial and Analytics teams are looking at top-line aggregates, a region-by-region audit of the 28-day test data and incident logs reveals critical blockers. Specifically, Scotland fails acceptance and refund criteria due to a supplier mapping failure, Midlands and London recorded high-severity dietary/allergen mismatches, and South East also experienced a dietary mismatch.

We should hold all regions until the dietary hard filter is deployed on 4 November, and until Scotland’s catalogue re-map is fully completed and verified.

---

Regional Evaluation Table (1–28 October Test Data)

RegionSubs Acceptance ($\ge 80\%$)Treatment Refunds vs Control ($\le +10\%$)Dietary/Allergen Mismatches (Target: 0)Basket Value ImpactRecommendation
:---:---:---:---:---:---
London$86.4\%$ (Pass)Control: 1.77%<br>Treatment: 1.63% ($-7.9\%$)1 High Incident (Nut allergy violation on 23 Oct)Control: £72.33<br>Treatment: £72.71 (+0.5%)No-Go (Hold for 4 Nov filter)
South East$87.0\%$ (Pass)Control: 1.66%<br>Treatment: 1.69% ($+1.8\%$)1 High Incident (Vegan dairy violation on 19 Oct)Control: £66.21<br>Treatment: £66.72 (+0.8%)No-Go (Hold for 4 Nov filter)
Midlands$84.5\%$ (Pass)Control: 1.69%<br>Treatment: 1.68% ($-0.6\%$)1 High Incident (Gluten-free coeliac violation on 4 Oct)Control: £55.22<br>Treatment: £55.51 (+0.5%)No-Go (Hold for 4 Nov filter)
North West$88.4\%$ (Pass)Control: 1.73%<br>Treatment: 1.66% ($-4.0\%$)0 IncidentsControl: £57.34<br>Treatment: £57.65 (+0.5%)Go with Conditions (Deploy 4 Nov)
Scotland$77.0\%$ (Fail)Control: 1.79%<br>Treatment: $3.90\%$ ($+117.9\%$)0 Incidents (High dairy wrong-size volume)Control: £52.48<br>Treatment: £52.88 (+0.8%)No-Go (Hold for catalogue re-map)

Note: Refund request rates are calculated per 100 orders for precise comparison against the 10% relative ceiling.

---

Region-by-Region Breakdown

  • London, South East, Midlands: All three regions easily clear the 80% acceptance threshold and show neutral-to-positive average basket values (debunking the £4 basket drop concern). However, all three recorded high-severity dietary/allergen failures (gluten, vegan, and nut-allergy breaches). Because the safety hard filter is not deployed until 4 November, rolling out on 2 November exposes customers to unacceptable allergen risks during the peak Christmas slot-booking window.
  • North West: Our strongest performer. It meets all four PRD criteria cleanly3: 88.4% acceptance1, treatment refund rates slightly below control, zero dietary incidents, and a stable basket value. It is ready to roll out immediately following the hard filter deployment.
  • Scotland: A clear No-Go. Acceptance dropped to 77.0% (failing the 80% bar), and treatment refund requests spiked catastrophically by 117.9% over control (driven by the dairy supplier switch on 15 Oct where pack sizes were mismapped). Until the catalogue re-map is finished and stabilized, rolling out here will compound operational friction.

Action Plan

  1. Hold Rollout on 2 November.
  2. Deploy Dietary Hard Filter on 4 November as scheduled, and verify regression tests.
  3. Complete Scotland Catalogue Re-map and clear open wrong-size substitution incidents.
  4. Staged Launch: Roll out to North West, London, South East, and Midlands on Monday, 9 November (concurrent with the Christmas slot-booking window opening), pending a 7-day post-fix stability review.2 Scotland follows only after catalogue remediation.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticStart again

    What we’d changeRecompute every figure from daily_metrics.csv, removing the duplicated North West rows. Almost every number in the table is wrong.

  2. 2Contradiction missedSurface the contradictionSubstantial rework

    What we’d changeResolve the contradiction: a 7-day review after the 4 November filter can't finish by 9 November, and North West is marked go in the table but held here. Answer the stakeholders' claims rather than dismissing them.

  3. 3Hypothesis stated as factReframe it as a hypothesisTargeted repair

    What we’d changeZero recorded mismatches in North West isn't evidence of safety. It runs the same unfiltered model as the regions that had them.

Grades and run details

Decision model 46 · LLM judge 8 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly49%
  • passAddresses the actual decision32%
  • passRespects explicit constraints8%
  • partialIdentifies material uncertainty63%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable26%
  • partialChecks every criterion, region by region64%
  • passFinds the Scotland breach96%
  • failExplains the basket gap48%
  • failCatches the duplicate rows100%
  • partialMakes a clear call83%
  • failLimits the downside of being wrong41%
  • failGets the base of every number right77%
Run
Run
#1
Time to output
46 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 8

Uses the supplied evidence correctlyWrongMixedWrong
Gemini 3.8 Flash · API

It contains material unsupported or incorrect current-situation claims, including North West GO despite a dietary incident and a wrong duplicate-row figure.

Sonnet 5.5 · API

The output's factual claims are drawn from the brief and supplied CSVs, with only minor rounding and one unsupported inference about incident reporting.

Gemini 3.5 Flash-Lite · Gemini

The output calculates rates from the raw daily_metrics.csv without acknowledging or removing the duplicate North West treatment rows flagged in incidents.csv, so the claimed rates for North West (and potentially others if the duplication affects summaries) are not fully reliable.

Identifies material uncertaintyWrongRightWrong
Gemini 3.8 Flash · API

It does not clearly name the key unknowns that would change the call, such as the hard filter slipping past 9 November or Scotland failing to recover after re-map.

Sonnet 5.5 · API

It names material unknowns such as filter deployment, Scotland re-map timing, duplicate handling, and incident under-reporting, and gives monitoring/rollback conditions.

Gemini 3.5 Flash-Lite · Gemini

Does not explicitly name the unknowns that could change the call (e.g., what if the dietary filter slips past 9 November, or if Scotland’s refunds don’t recover after the catalogue fix) or state the thresholds that would reverse the decision.

Checks every criterion, region by regionWrongRightRight
Gemini 3.8 Flash · API

The table checks all four criteria by region, but the North West dietary result is wrong and the duplicate handling figure is wrong.

Sonnet 5.5 · API

It checks all four launch criteria for each region with figures computed from the files, including the duplicate correction.

Gemini 3.5 Flash-Lite · Gemini

The table checks acceptance, refunds vs control, dietary incidents and basket value for each of the five regions, with figures drawn from the test data.

Explains the basket gapRightRightMixed
Gemini 3.8 Flash · API

It compares within-region baskets and explains the pooled gap as a Simpson's Paradox mix effect.

Sonnet 5.5 · API

It compares baskets within each region and explains the pooled gap as an artefact of the uneven regional split.

Gemini 3.5 Flash-Lite · Gemini

Shows treatment basket values are equal or higher in every region and states the aggregated £4 drop concern is debunked, satisfying the Simpson’s paradox explanation.

Catches the duplicate rowsWrongRightWrong
Gemini 3.8 Flash · API

It notices the North West duplicates but misstates them as 3,194 duplicate rows rather than 2,194 duplicate orders across three rows.

Sonnet 5.5 · API

It finds the duplicated North West treatment rows, removes them, and says the result barely changes.

Gemini 3.5 Flash-Lite · Gemini

Does not mention or handle the duplicate North West treatment rows on 8–10 October that the incident log flags, so the quality check is absent.

Makes a clear callMixedRightRight
Gemini 3.8 Flash · API

It gives dates and conditions, but the North West call is wrong and it does not state what would change the overall call.

Sonnet 5.5 · API

It gives one clear call per region with dates, conditions, and triggers for change.

Gemini 3.5 Flash-Lite · Gemini

Gives a clear decision per region (No‑Go for Scotland, conditional Go for others from 9 Nov) with dates and the conditions required to proceed.

Limits the downside of being wrongWrongRightWrong
Gemini 3.8 Flash · API

It lacks post-launch monitoring signals, pause/rollback thresholds, and irreversible-risk triggers.

Sonnet 5.5 · API

It limits downside with staged switch-on, daily monitoring, immediate rollback on any dietary mismatch, and Scotland held out.

Gemini 3.5 Flash-Lite · Gemini

While it mentions a 7‑day post‑fix stability review, it does not specify a signal (e.g., refund rate threshold) or a rollback trigger, and does not flag irreversible actions.

Gets the base of every number rightWrongWrongMixed
Gemini 3.8 Flash · API

Several derived figures are correct, but the duplicate-row count and North West dietary status are wrong.

Sonnet 5.5 · API

Some derived figures are off, including the pooled basket gap and the control-order share for London, and the basket base is stated as means of daily averages rather than order-weighted averages.

Gemini 3.5 Flash-Lite · Gemini

Rates are computed as sums over the period (acceptance = total accepted / total offered, refund rate = total refunds / total orders), so denominators are consistent and the relative comparisons are valid.

All got wrong 1

Avoids unsupported claimsWrongWrongWrong
Gemini 3.8 Flash · API

It presents North West as clean and safe to launch despite an open dietary incident, and asserts duplicate-row removal with an incorrect figure.

Sonnet 5.5 · API

It asserts that zero logged dietary incidents in North West and Scotland are only because none were reported and that incidents come from customer complaints, which is not established by the supplied evidence.

Gemini 3.5 Flash-Lite · Gemini

States as fact that Scotland’s refund spike was 'driven by the dairy supplier switch where pack sizes were mismapped' without labelling it as a hypothesis, even though the data only shows a correlation.

All got right 4

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

It commits early to a clear region-by-region call and states conditions and dates.

Sonnet 5.5 · API

It commits early to no-go on 2 November, conditional go from 4 November for four regions, hold Scotland, and states what would change the call.

Gemini 3.5 Flash-Lite · Gemini

Commits to a clear No‑Go for 2 November, then a conditional Go by region for 9 November after the filter deploys, and states Scotland’s hold conditions.

Respects explicit constraintsRightRightRight
Gemini 3.8 Flash · API

It is under 900 words, includes a region-by-criterion table, and respects the launch-criteria framing.

Sonnet 5.5 · API

It provides a region-by-region recommendation with a launch-criteria table and keeps prose under 900 words.

Gemini 3.5 Flash-Lite · Gemini

Prose is well under 900 words, addressed to the launch meeting, and includes a region‑by‑region table with the four criteria.

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

The requested recommendation and table are present and usable in form, though some figures need correction.

Sonnet 5.5 · API

The requested recommendation and table are present, usable for the launch meeting, and within the length limit.

Gemini 3.5 Flash-Lite · Gemini

Provides a complete, table‑based recommendation with a region‑by‑region call, within the word limit, that the launch meeting could act on.

Finds the Scotland breachRightRightRight
Gemini 3.8 Flash · API

It identifies Scotland's refund breach, ties it to the supplier change, notes acceptance dip, and holds Scotland.

Sonnet 5.5 · API

It identifies Scotland's refund breach, ties it to the 15 October supplier change, notes the acceptance dip, and holds Scotland.

Gemini 3.5 Flash-Lite · Gemini

Identifies Scotland’s refund breach (+117.9%), links it to the post‑15 October supplier change, notes the acceptance drop to 77%, and holds Scotland.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.388.32None
2GPT-6 AstrawithChatGPT86.088.32None
3Sonnet 5.5withAPI85.488.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI82.177.62None
6Gemini 3.8 FlashwithAPI71.931.521 capped
7Gemini 3.5 Flash-LitewithGemini50.642.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook