Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Respects explicit constraints98% pass
    The output is a launch recommendation, addresses the meeting, and stays within the 400-word limit (345 words).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong35% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows54% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Gets the base of every number right67% pass
    Some derived figures are off, including the pooled basket gap and the control-order share for London, and the basket base is stated as means of daily averages rather than order-weighted averages.
    Sonnet 5.5 · API · Go/no-go for AI substitutions, from the rollout data

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Staff PM for fulfilment at Basketful. On 30 October the launch meeting decides whether AI auto-substitutions (the model picks a replacement when an item is out of stock, instead of the picker) roll out to every region on Monday 2 November, ahead of the Christmas delivery-slot booking window that opens on 9 November. The rollout workbook is attached as two files: daily_metrics.csv (28 days of the regional test, by region and arm) and incidents.csv. Work from the data, not the summaries people have given you. Write your recommendation for the meeting: go, no-go or go with conditions, region by region, and why. Include a table that checks each launch criterion for each region with the figures. Keep the prose under 900 words.

The test1–28 October, five regions. In each region orders were split between control (the picker chooses substitutes, as today) and treatment (AI auto-substitutions). The split was not 50/50 everywhere: treatment got 30% of orders in London, 50% in the South East and North West, and 70% in the Midlands and Scotland.
Launch criteria (agreed in the PRD)1. Substitution acceptance (subs_accepted ÷ subs_offered) of at least 80% in treatment, in every region. 2. Refund requests per 1,000 orders in treatment no more than 10% above control, in every region. 3. Zero dietary or allergen mismatches (a substitute that breaks a gluten-free, vegan, vegetarian, nut-free, halal or kosher attribute). 4. Average basket value not lower in treatment than control.
What people have saidHead of Commercial: “Acceptance is up 15 points and refunds are within the guardrail overall. Every day we wait costs us Christmas.” Head of Analytics: “Treatment baskets are £4 smaller across the test. That worries me.” Ops director (Scotland): “The supplier change caused some noise, but it's settling down.”
EngineeringThe dietary-attribute hard filter (never substitute across those attributes) is merged and scheduled to deploy on 4 November after regression tests. Scotland's catalogue re-map for the new dairy supplier is in progress with no date. Rollout can be switched on region by region.
daily_metrics.csv283 rows · Download
date,region,arm,orders,items_ordered,items_out_of_stock,subs_offered,subs_accepted,refund_requests,refund_value_gbp,complaints,avg_basket_gbp
2026-10-01,London,control,1757,52273,2351,1953,1387,32,160.71,6,72.89
2026-10-01,London,treatment,692,20455,871,820,708,12,62.05,2,72.42
2026-10-01,South East,control,853,22846,996,833,593,16,74.49,3,65.62
2026-10-01,South East,treatment,928,25223,1166,1106,952,15,75.93,3,66.59
2026-10-01,Midlands,control,467,10468,465,383,268,8,34.57,1,54.92
2026-10-01,Midlands,treatment,1113,25136,1166,1087,947,21,106.35,3,55.33
…
incidents.csv13 rows · Download
date,region,arm,type,severity,summary,status
2026-10-02,London,treatment,wrong_size,low,"4-pint milk substituted with 1-pint, customer accepted at door",closed
2026-10-04,Midlands,treatment,dietary_mismatch,high,"Gluten-free sliced loaf substituted with standard white loaf; customer coeliac, noticed at home",closed: refund and apology
2026-10-06,South East,control,wrong_item,low,"Picker substituted oat milk with soya milk; customer rejected at door",closed
2026-10-07,North West,treatment,data_export,low,"Daily export job re-ran after a timeout on 8, 9 and 10 Oct; analytics team flagged possible duplicate rows",open
2026-10-09,London,treatment,price,medium,"Substitute priced higher than original; customer charged the difference against policy",closed: policy fix deployed 11 Oct
2026-10-12,Scotland,control,late_delivery,low,"Van breakdown, 14 orders late",closed
…
What a strong answer does

No-go for a national launch on 2 November, because criterion 3 fails everywhere until the dietary hard filter ships: incidents.csv has three treatment dietary mismatches (4 Oct Midlands, 19 Oct South East, 23 Oct London, the last involving nuts and a child with an allergy). Go with conditions, region by region, once the filter is deployed and verified (planned 4 November, still before the 9 November booking window): London, South East, Midlands and North West pass the other criteria (acceptance about 86% against 71% control; refunds 2–5% below control). Hold Scotland: over the month its treatment refunds are about 66% above control, and from the 15 October supplier change about 143% above (43 against 18 per 1,000 orders), with acceptance falling to about 79%, below the 80% bar; it waits for the catalogue re-map and a clean re-test. The pooled refund rate (about 8% above control) hides this, and the criterion is per region. The pooled basket gap (about £65 against £61) is an artefact of the uneven split, since treatment baskets are equal or higher in every region (Simpson's paradox), so criterion 4 passes. Notices that North West treatment rows for 8–10 October are duplicated (an extra 2,194 orders), removes them, and says it doesn't change the result. States what would change the call: the filter slipping past 9 November, or Scotland's refunds not recovering.

Critical failures (cap the score)
  • Recommends any rollout before the dietary hard filter is deployed
  • Recommends rolling out to Scotland while its refund guardrail is breached
Case

v1.6 · synthetic · e-commerce, AI product, data analysis, staff-level

Usable with a quick edit11 of 13 checks passed

The output is evidence-based, fully meets the brief, and can be used directly by a product manager without re-checking data or redoing analysis.

Output

AI auto-substitutions: launch recommendation

Oct 29, 2026 · @Martin Slaney · For the 30 October launch meeting

Recommendation

No-go for Monday 2 November in any region. The dietary hard filter isn't live until 4 November, and the test produced three allergen and dietary mismatches.

Go with conditions for London, South East, Midlands and North West from Thursday 5 November, once the filter is live and verified. That is still four days before the Christmas booking window opens on 9 November.

No-go for Scotland until the dairy catalogue re-map ships and a fresh week of test data passes the criteria.

Launch criteria by region

Figures are for 1–28 October from daily\_metrics.csv, with duplicate rows removed. T = treatment, C = control.

Region1. Acceptance (T), ≥80%2. Refunds per 1k orders, T vs C, ≤+10%3. Dietary mismatches (T), 04. Avg basket, T vs CVerdict
London86.7% ✅17.4 vs 18.0 (−3%) ✅1: nut-free to hazelnut, 23 Oct, still open ❌£72.70 vs £72.49 ✅Go 5 Nov, after filter
South East86.4% ✅17.5 vs 17.9 (−2%) ✅1: vegan to dairy, 19 Oct ❌£66.64 vs £66.20 ✅Go 5 Nov, after filter
Midlands86.4% ✅17.5 vs 17.9 (−2%) ✅1: gluten-free to standard, 4 Oct ❌£55.66 vs £55.35 ✅Go 5 Nov, after filter
North West86.3% ✅17.6 vs 18.5 (−5%) ✅0 ✅ (same model, so not evidence of safety)£57.51 vs £57.29 ✅Go 5 Nov, after filter
Scotland82.8% overall; 79.1% since 15 Oct ❌30.2 vs 18.2 (+66%); since 15 Oct 43.2 vs 17.8 (+143%) ❌0 ✅£52.98 vs £52.62 ✅No-go until re-map and re-test

Why not 2 November

Criterion 3 requires zero dietary or allergen mismatches. Treatment had three, all rated high severity: a gluten-free loaf swapped for a standard one (the customer is coeliac), vegan cheese swapped for dairy cheddar (the customer complained on social media), and a nut-free multipack swapped for one containing hazelnuts, sent to a child with a nut allergy. The nut incident is still open. Control had none.

The fix is the dietary hard filter, which is merged and deploys on 4 November after regression tests. Launching on 2 November would put every order in every region onto the unfiltered model for two days. North West's zero is luck, because it runs the same model.

Strictly, all four "go" regions fail criterion 3 as it's written. I'm asking the meeting to accept the filter as the remedy, on condition that we show it would have blocked all three incidents. If the meeting won't accept that, the honest answer is no-go everywhere.

Waiting three days costs nothing for Christmas, because bookings don't open until 9 November.

What the summaries get wrong

Head of Commercial: "refunds within the guardrail overall." Pooled, treatment refunds are 7.9% above control (19.5 vs 18.1 per 1,000 orders), which is inside 10%. But the criterion applies to every region, and the pooled figure hides Scotland at +66%. In the other four regions, treatment refunds are below control. The 15-point acceptance lift is real in those four regions (+15.0 to +15.4 points). In Scotland it's +11.7 overall and falling.

Head of Analytics: "treatment baskets are £4 smaller." Pooled, the gap is £61.09 vs £65.15, but it comes from how orders were split, not from the feature. Treatment got 70% of orders in Midlands and Scotland, which have the smallest baskets (£52–56), and only 30% in London, which has the largest (£72). Within every region, treatment baskets are £0.21–£0.44 larger, so criterion 4 passes everywhere.

Ops director, Scotland: "it's settling down." It isn't. Before the supplier switch on 15 October, Scotland's treatment arm matched the other regions (86.3% acceptance, 17.3 refunds per 1,000 vs 18.6 in control). Since then, treatment refunds have run at 25–33 a day with no downward trend: 28.1 a day on average from 15 to 21 October and 28.7 from 22 to 28 October. Refund value per order has tripled, from £0.09 to £0.30. Scotland's control arm is unaffected at 17.8 per 1,000. So the problem is specific to the model reading a mis-mapped catalogue: pickers can see the real pack sizes, and the model can't. The re-map has no date.

Data quality

The North West treatment rows for 8, 9 and 10 October each appear twice, because the export job re-ran (the open incident from 7 October). I removed the duplicates. They don't change any verdict: North West acceptance moves from 86.2% to 86.3%.

Conditions for the 5 November go

  1. The dietary filter deploys on 4 November and passes regression tests. We replay the three incident orders and confirm the filter blocks each of them.2
  2. The London nut-allergen incident is closed with a root cause.
  3. We switch on region by region and keep a 10% control holdout in each region through Christmas, so criteria 2 and 4 stay measurable.
  4. Kill switch: any single dietary or allergen mismatch turns the region off, as does a treatment refund rate more than 10% above its holdout over a rolling 3 days.1
  5. Scotland: ship the catalogue re-map, then run 7 days of test and re-check the criteria. If it can't be ready by 9 November, Scotland takes Christmas orders with picker substitutions.

What a PM had to fix

  1. 1Test or gate too weakTighten the testQuick edit

    What we’d changeSet the refund trigger outside day-to-day noise, with a bigger gap or a longer window, or it will switch regions off for no reason.

  2. 2Test or gate too weakTighten the testQuick edit

    What we’d changeThree replays is thin evidence the filter works. Test it across every protected attribute, including missing or conflicting ones.

Check by check

Mixed · 2
  • Uses the supplied evidence correctlyEvery factual claim about the current situation is drawn directly from the supplied files or computed correctly from them.The two graders disagreed on this one.
  • Gets the base of every number rightAll derived figures (acceptance, refunds per 1k, basket averages) are computed correctly from the supplied data with appropriate denominators.The two graders disagreed on this one.
Got right · 11
  • Addresses the actual decisionCommits unequivocally to a no-go for 2 November, go with conditions from 5 November for four regions, and no-go for Scotland, and states what would change each call.
  • Respects explicit constraintsRespects the 900-word prose limit, includes the required region-by-criterion table, and is addressed to the launch meeting as requested.
  • Identifies material uncertaintyNames the unknowns (filter deployment date, Scotland re-map readiness, acceptance of filter as remedy) and says each would change the decision.
  • Avoids unsupported claimsHypotheses (e.g. 'luck' for NW zero) are clearly labelled as interpretation; otherwise claims are grounded in the data.
  • Produces the required deliverableThe output is a complete, actionable recommendation with a criterion-check table, suitable for the meeting with minimal editing.
  • Checks every criterion, region by regionAll four launch criteria are checked for each region with correct computed figures from the files.
  • Finds the Scotland breachClearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
  • Explains the basket gapExplicitly explains the pooled basket gap as a traffic-mix artefact from the uneven regional split, with within-region comparisons showing treatment baskets are higher.
  • Catches the duplicate rowsNotices the duplicated North West treatment rows for 8–10 October, removes them, and confirms the verdict doesn't change.
  • Makes a clear callGives a clear decision per region with dates (no-go 2 Nov, go with conditions 5 Nov for four regions, no-go for Scotland) and states what would change it.
  • Limits the downside of being wrongSpecifies post-launch kill switch (any dietary mismatch or refund rate >10% above holdout over rolling 3 days), a 10% control holdout, and conditions for Scotland re-test.

Grades and run details

Decision model 81 · LLM judge 14 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly12%
  • passAddresses the actual decision83%
  • passRespects explicit constraints23%
  • passIdentifies material uncertainty98%
  • partialAvoids unsupported claims23%
  • passProduces the required deliverable76%
  • passChecks every criterion, region by region52%
  • passFinds the Scotland breach100%
  • passExplains the basket gap99%
  • passCatches the duplicate rows100%
  • passMakes a clear call95%
  • passLimits the downside of being wrong84%
  • failGets the base of every number right33%
Run
Run
#1
Time to output
2.6 min
Submitted
25 Sept 2026

Usable with a quick edit10 of 13 checks passed

The recommendation, table, and calculations are solid; only minor wording additions to state duplicate impact and flag irreversible dietary harm are needed.

Output

Recommendation for 30 October: go with conditions in four regions; no-go in Scotland

Do not expand any region on Monday 2 November. London, South East, Midlands and North West are candidates for a gated rollout after the dietary hard filter is deployed and validated—not before its scheduled 4 November deployment. Scotland must additionally resolve and re-test its catalogue mapping.

Christmas urgency does not justify launching ahead of a known safety fix. The 9 November booking window should guide sequencing, not override the PRD.

Method and launch-criterion checks

I removed three exact duplicate North West treatment rows, dated 8–10 October, consistent with the export incident. This leaves 280 unique date–region–arm records.

Acceptance and refunds below use ratios of summed counts, not averages of daily rates. Basket values are order-weighted. T = treatment; C = control. Dietary counts are recorded treatment incidents; zero recorded incidents is not proof that none occurred.

RegionTreatment acceptance ≥80%Refunds/1,000 orders: T vs C; relative change ≤+10%Zero dietary/allergen mismatchesAverage basket: T vs C; difference ≥£0Recommendation
LondonPass: 86.7% (23,607/27,234)Pass: 17.44 vs 18.04; −3.3%Fail: 1, nut-free → hazelnuts, 23 OctPass: £72.70 vs £72.49; +£0.21Go with conditions; not 2 Nov
South EastPass: 86.4% (26,586/30,770)Pass: 17.54 vs 17.89; −1.9%Fail: 1, vegan → dairy, 19 OctPass: £66.64 vs £66.20; +£0.44Go with conditions; not 2 Nov
MidlandsPass: 86.3% (26,732/30,958)Pass: 17.48 vs 17.89; −2.3%Fail: 1, gluten-free → standard bread, 4 OctPass: £55.66 vs £55.35; +£0.31Go with conditions; not 2 Nov
North WestPass: 86.3% (18,153/21,039)Pass: 17.58 vs 18.47; −4.8%Pass on incident log: 0 recordedPass: £57.51 vs £57.29; +£0.22Go with conditions; not 2 Nov
ScotlandPass over 28 days: 82.8% (14,333/17,319)Fail: 30.16 vs 18.20; +65.7%Pass on incident log: 0 recordedPass: £52.98 vs £52.62; +£0.36No-go

What the data actually say

Commercial’s overall refund claim hides a regional failure. After deduplication, pooled refunds are 19.50 per 1,000 treatment orders versus 18.08 control: +7.9%, within the overall guardrail. But the PRD requires compliance in every region, and Scotland fails badly. Strong acceptance cannot compensate for dietary breaches.

Analytics’ £4 basket concern is real in the pooled arithmetic, but not within regions. Pooled baskets are £61.09 treatment versus £65.15 control (−£4.06). Treatment disproportionately includes lower-basket Midlands and Scotland orders, while control disproportionately includes higher-basket London orders. Every within-region comparison is positive. This is a composition effect, not evidence here that AI reduces basket value. These are point estimates; the daily averages alone do not establish statistical certainty.

Scotland is not settling down. Following the supplier switch on 15 October:

  • Acceptance falls from 86.3% on 1–14 October to 79.1% on 15–28 October, below the launch threshold.
  • Treatment refunds rise from 17.25 to 43.25 per 1,000 orders. Post-switch control is 17.79, making treatment 143.1% worse.
  • The final week still has 79.2% acceptance and 43.64 refunds per 1,000.

The open pack-size mapping incident is consistent with this sustained deterioration. The full-period acceptance pass masks the current operating state.

Conditions and release plan

For London, South East, Midlands and North West, require:

  1. Deployed, validated safety protection: successful regression tests across gluten-free, vegan, vegetarian, nut-free, halal and kosher attributes, including missing/incorrect metadata and the three observed failures. “Merged” is not “protected.”
  2. Fresh post-fix evidence: a controlled validation window with zero dietary mismatches and all regional numerical criteria met, with explicit launch sign-off. Refunds and apologies do not erase historical criterion failures.
  3. Staged regional expansion: monitor acceptance, refunds and dietary incidents daily; stop immediately for any dietary mismatch and roll back for guardrail breaches. Reconcile the North West export before reporting.

For Scotland, also require a completed catalogue re-map, pack-size validation and fresh post-remediation evidence. With no completion date and continuing deterioration, make no Christmas rollout commitment.

Meeting decision: reject the all-region 2 November launch; approve a gated path for four regions, while keeping Scotland blocked.

Check by check

Got wrong · 2
  • Catches the duplicate rowsFinds and removes the duplicate rows but never states whether they change the results.
  • Limits the downside of being wrongGives monitoring and rollback thresholds but does not flag irreversible consequences such as customer-facing dietary substitutions or communications.
Mixed · 1
  • Gets the base of every number rightAll percentages, per-1,000 rates, and basket differences use correct summed or order-weighted bases.The two graders disagreed on this one.
Got right · 10
  • Uses the supplied evidence correctlyAll current-state numbers and facts trace to the brief, incidents.csv, or arithmetic from daily_metrics.csv.
  • Addresses the actual decisionCommits to a clear region-by-region go/no-go with dates and conditions for the launch meeting.
  • Respects explicit constraintsIncludes the required region-by-criterion table and keeps prose within the length while respecting the no-early-launch safety constraint.
  • Identifies material uncertaintyNames filter deployment/validation and Scotland catalogue remap as unresolved and ties go decisions to fresh evidence.
  • Avoids unsupported claimsCausal claims are hedged ('consistent with') and interpretations are labelled as composition effects or point estimates.
  • Produces the required deliverableProvides a usable recommendation, table, and release conditions that the reader could act on with light edits.
  • Checks every criterion, region by regionRegion-by-criterion table covers acceptance, refunds vs control, dietary incidents, and basket values with computed figures.
  • Finds the Scotland breachFinds Scotland's refund breach, ties it to the supplier change with post-15 Oct numbers, and holds Scotland.
  • Explains the basket gapExplains the pooled basket gap as a traffic-mix effect and shows all within-region treatment baskets positive.
  • Makes a clear callMakes a clear per-region call with dates and conditions, and Scotland is blocked.

Grades and run details

Decision model 85 · LLM judge 12 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly26%
  • passAddresses the actual decision81%
  • passRespects explicit constraints41%
  • passIdentifies material uncertainty46%
  • passAvoids unsupported claims42%
  • passProduces the required deliverable79%
  • passChecks every criterion, region by region71%
  • passFinds the Scotland breach100%
  • passExplains the basket gap99%
  • partialCatches the duplicate rows49%
  • passMakes a clear call34%
  • partialLimits the downside of being wrong43%
  • failGets the base of every number right22%
Run
Run
#1
API response time
1.9 min
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI93.788.32None
2GPT-6 AstrawithChatGPT89.888.32None
3GPT-6.1 SolwithAPI87.388.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI86.077.62None
6Gemini 3.5 Flash-LitewithGemini53.142.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook