Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Addresses the actual decision95% pass
    The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong32% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows50% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Uses the supplied evidence correctly54% pass
    The output states that pilot agents were hands-on, engaged and motivated, which is not supported by the supplied context (only that 7/8 wanted to keep it).
    Opus 5.5 · Claude · Go/no-go for AI-drafted support replies

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Staff PM for fulfilment at Basketful. On 30 October the launch meeting decides whether AI auto-substitutions (the model picks a replacement when an item is out of stock, instead of the picker) roll out to every region on Monday 2 November, ahead of the Christmas delivery-slot booking window that opens on 9 November. The rollout workbook is attached as two files: daily_metrics.csv (28 days of the regional test, by region and arm) and incidents.csv. Work from the data, not the summaries people have given you. Write your recommendation for the meeting: go, no-go or go with conditions, region by region, and why. Include a table that checks each launch criterion for each region with the figures. Keep the prose under 900 words.

What the model was given6 items: The test, Launch criteria (agreed in the PRD), What people have said, Engineering, daily_metrics.csv, incidents.csv
The test1–28 October, five regions. In each region orders were split between control (the picker chooses substitutes, as today) and treatment (AI auto-substitutions). The split was not 50/50 everywhere: treatment got 30% of orders in London, 50% in the South East and North West, and 70% in the Midlands and Scotland.
Launch criteria (agreed in the PRD)1. Substitution acceptance (subs_accepted ÷ subs_offered) of at least 80% in treatment, in every region. 2. Refund requests per 1,000 orders in treatment no more than 10% above control, in every region. 3. Zero dietary or allergen mismatches (a substitute that breaks a gluten-free, vegan, vegetarian, nut-free, halal or kosher attribute). 4. Average basket value not lower in treatment than control.
What people have saidHead of Commercial: “Acceptance is up 15 points and refunds are within the guardrail overall. Every day we wait costs us Christmas.” Head of Analytics: “Treatment baskets are £4 smaller across the test. That worries me.” Ops director (Scotland): “The supplier change caused some noise, but it's settling down.”
EngineeringThe dietary-attribute hard filter (never substitute across those attributes) is merged and scheduled to deploy on 4 November after regression tests. Scotland's catalogue re-map for the new dairy supplier is in progress with no date. Rollout can be switched on region by region.
daily_metrics.csv283 rows · Download
date,region,arm,orders,items_ordered,items_out_of_stock,subs_offered,subs_accepted,refund_requests,refund_value_gbp,complaints,avg_basket_gbp
2026-10-01,London,control,1757,52273,2351,1953,1387,32,160.71,6,72.89
2026-10-01,London,treatment,692,20455,871,820,708,12,62.05,2,72.42
2026-10-01,South East,control,853,22846,996,833,593,16,74.49,3,65.62
2026-10-01,South East,treatment,928,25223,1166,1106,952,15,75.93,3,66.59
2026-10-01,Midlands,control,467,10468,465,383,268,8,34.57,1,54.92
2026-10-01,Midlands,treatment,1113,25136,1166,1087,947,21,106.35,3,55.33
…
incidents.csv13 rows · Download
date,region,arm,type,severity,summary,status
2026-10-02,London,treatment,wrong_size,low,"4-pint milk substituted with 1-pint, customer accepted at door",closed
2026-10-04,Midlands,treatment,dietary_mismatch,high,"Gluten-free sliced loaf substituted with standard white loaf; customer coeliac, noticed at home",closed: refund and apology
2026-10-06,South East,control,wrong_item,low,"Picker substituted oat milk with soya milk; customer rejected at door",closed
2026-10-07,North West,treatment,data_export,low,"Daily export job re-ran after a timeout on 8, 9 and 10 Oct; analytics team flagged possible duplicate rows",open
2026-10-09,London,treatment,price,medium,"Substitute priced higher than original; customer charged the difference against policy",closed: policy fix deployed 11 Oct
2026-10-12,Scotland,control,late_delivery,low,"Van breakdown, 14 orders late",closed
…
What a strong answer doesThe answer key the graders mark against

No-go for a national launch on 2 November, because criterion 3 fails everywhere until the dietary hard filter ships: incidents.csv has three treatment dietary mismatches (4 Oct Midlands, 19 Oct South East, 23 Oct London, the last involving nuts and a child with an allergy). Go with conditions, region by region, once the filter is deployed and verified (planned 4 November, still before the 9 November booking window): London, South East, Midlands and North West pass the other criteria (acceptance about 86% against 71% control; refunds 2–5% below control). Hold Scotland: over the month its treatment refunds are about 66% above control, and from the 15 October supplier change about 143% above (43 against 18 per 1,000 orders), with acceptance falling to about 79%, below the 80% bar; it waits for the catalogue re-map and a clean re-test. The pooled refund rate (about 8% above control) hides this, and the criterion is per region. The pooled basket gap (about £65 against £61) is an artefact of the uneven split, since treatment baskets are equal or higher in every region (Simpson's paradox), so criterion 4 passes. Notices that North West treatment rows for 8–10 October are duplicated (an extra 2,194 orders), removes them, and says it doesn't change the result. States what would change the call: the filter slipping past 9 November, or Scotland's refunds not recovering.

Critical failures (cap the score at 40)
  • Recommends any rollout before the dietary hard filter is deployed
  • Recommends rolling out to Scotland while its refund guardrail is breached
Case

Staff-level · v1.8 · synthetic · e-commerce, AI product, data analysis, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: Recommends any rollout before the dietary hard filter is deployed. Score capped at 40.

Needs more than a quick edit5 of 13 checks passed

A PM could not use it without rechecking evidence because the North West GO and duplicate-row figure are wrong. It also needs added rollback triggers and change conditions.

Output

Recommendation Summary: Go With Conditions

Do not execute a blanket rollout on Monday, 2 November.

  • North West: GO (2 November). Meets all four PRD launch criteria cleanly.
  • London, South East, Midlands: GO WITH CONDITIONS (Deploy Thursday, 5 November). Commercial metrics are strong, but all three regions breached Criterion 3 with severe dietary/allergen substitutions. Launch must be gated on the successful deployment of the dietary-attribute hard filter on 4 November.
  • Scotland: NO-GO. Fails Criterion 2 (+67.4% refund spike vs. control) and Criterion 1 post-supplier transition (76.6% acceptance). Revert 100% to control until catalogue re-mapping is completed and re-tested.

This phased approach protects customer safety and brand reputation while capturing AI substitution value ahead of Christmas slot bookings opening on 9 November.

---

PRD Launch Criteria Assessment

Note: North West treatment metrics reflect deduplication of export retries on 8–10 October (3,194 duplicate rows removed).

RegionArmOrders1. Subs Acceptance (Target: $\ge$80%)2. Refunds / 1k Orders (Target: $\le$+10% vs Ctrl)3. Allergen / Dietary Mismatches (Target: 0)4. Avg Basket Value (Target: $\ge$ Ctrl)Status
:---:---:---:---:---:---:---:---
LondonControl<br>Treatment49,159<br>21,05471.5%<br>87.4%18.06<br>17.53 (-2.9%)0<br>1 (FAIL)£72.52<br>£72.68 (+£0.16)CONDITIONAL<br>Hold for filter
South EastControl<br>Treatment26,485<br>26,91971.7%<br>86.5%17.86<br>17.35 (-2.9%)0<br>1 (FAIL)£66.27<br>£66.57 (+£0.30)CONDITIONAL<br>Hold for filter
MidlandsControl<br>Treatment13,048<br>31,17471.3%<br>85.2%17.86<br>17.58 (-1.6%)0<br>1 (FAIL)£55.37<br>£55.62 (+£0.25)CONDITIONAL<br>Hold for filter
North WestControl<br>Treatment20,702<br>20,77470.8%<br>85.8%18.26<br>17.47 (-4.3%)0<br>0 (PASS)£57.29<br>£57.53 (+£0.24)GO<br>Launch 2 Nov
ScotlandControl<br>Treatment7,853<br>18,37171.3%<br>81.5% (76.6% post-15 Oct)18.21<br>30.48 (+67.4% - FAIL)0<br>0 (PASS)£52.62<br>£52.89 (+£0.27)NO-GO<br>Catalogue fault

---

Analysis & Stakeholder Debunking

#### 1. Head of Analytics: “Treatment baskets are £4 smaller across the test.” Finding: False alarm caused by Simpson’s Paradox. Across all pooled orders, treatment average basket appears lower (£61.06 treatment vs. £65.18 control, a £4.12 deficit). However, this is an aggregation distortion caused by the uneven regional split: London (average basket ~£73) received only a 30% treatment allocation, whereas lower-basket regions (Midlands at ~£55 and Scotland at ~£53) received a 70% treatment allocation.

When evaluated within every single region, treatment basket value was higher than control by £0.16 to £0.30. Criterion 4 is fully satisfied across all regions.

#### 2. Head of Commercial: “Refunds are within the guardrail overall. Every day we wait costs Christmas.” Finding: Misleading pooled metric masking a severe regional failure. While aggregate refunds appear within limits (+8.1% overall), the PRD explicitly requires compliance in every region. Scotland completely breaches this guardrail (+67.4% above control). Launching everywhere on 2 November would risk widespread customer attrition in Scotland right before holiday booking.

#### 3. Ops Director (Scotland): “The supplier change caused noise, but it's settling down.” Finding: Factually incorrect; Scotland is failing and persistent. Following the 15 October dairy supplier change, wrong-size substitutions (e.g., pack sizes mapped to incorrect product families) caused: * Daily refunds in treatment to triple from 11.6/day (Oct 1–14) to 28.4/day (Oct 15–28), reaching 43.7 refunds per 1,000 orders (+140% vs. control). * Treatment acceptance dropped from 86.3% down to 76.6% (failing Criterion 1). * Daily refund value increased 4x to £180–£247/day. Performance has not stabilized. Because the catalogue re-map is incomplete with no ETA, Scotland cannot launch.

#### 4. The Dietary Hazard (London, South East, Midlands) Treatment produced three critical safety failures: * 4 Oct (Midlands): Coeliac customer received standard bread for gluten-free. * 19 Oct (South East): Vegan customer received dairy cheddar. * 23 Oct (London): Child with nut allergies received hazelnut-containing bars for nut-free bars.

Launching auto-substitutions in these regions before deploying technical guardrails creates intolerable liability.

---

Implementation Plan

1. Monday, 2 November: * Roll out North West to 100% treatment. (Clean test data, no allergen incidents, +15pp acceptance lift). * Revert Scotland treatment to 0% (100% control). Keep human pickers selecting substitutes until the catalogue re-map is verified and back-tested. * Maintain London, South East, and Midlands at current test splits for 48 hours.

2. Wednesday, 4 November: * Engineering deploys the dietary-attribute hard filter to production following scheduled regression testing.

3. Thursday, 5 November: * Verify filter telemetry on live orders. * Roll out London, South East, and Midlands to 100% treatment.

4. Monday, 9 November: * Christmas delivery-slot booking opens on schedule with 4 of 5 regions operating under fully automated, safety-enforced substitutions, leaving human pickers dedicated to Scotland until rectified.

Claims the judge couldn’t find in the brief

  • North West: GO (2 November).
  • North West treatment metrics reflect deduplication of export retries on 8–10 October (3,194 duplicate rows removed).

Grades and run details

Decision model 54 · LLM judge 5 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly75%
  • passAddresses the actual decision51%
  • partialRespects explicit constraints23%
  • partialIdentifies material uncertainty39%
  • failAvoids unsupported claims43%
  • passProduces the required deliverable53%
  • partialChecks every criterion, region by region50%
  • passFinds the Scotland breach100%
  • passExplains the basket gap99%
  • partialCatches the duplicate rows84%
  • passMakes a clear call42%
  • failLimits the downside of being wrong30%
  • failGets the base of every number right69%
Run
Run
#1
API response time
85 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 13 checks passed

The output is fully reasoned, includes the required table and conditions, and could be used with a quick edit to remove the one unsupported sentence.

Output

Recommendation for 30 October: no-go in every region on 2 November. Conditional phased rollout thereafter in London, South East, Midlands and North West. Scotland remains blocked pending catalogue repair and fresh evidence.

The dietary filter is scheduled for 4 November, two days after the proposed launch. Three treatment substitutions breached dietary requirements, including gluten-free bread replaced with standard bread and nut-free bars replaced with hazelnut bars. The shared system needs this control before expansion, including in regions with no recorded mismatch.

I would pause existing auto-substitution exposure and return to picker selection until the filter is deployed and verified.

Regional checks, 1–28 October

I removed three exact duplicate North West treatment rows dated 8–10 October4, consistent with the export incident, leaving 280 unique region–arm–day records. Acceptance and refund rates use summed numerators and denominators. Basket values are weighted by orders.

T = treatment; C = control. Refund changes are relative changes in requests per 1,000 orders.

RegionAcceptance T ≥80%Refunds/1,000: T vs C; change ≤+10%Dietary mismatches: zeroBasket T vs C; difference ≥£0Decision for 2 Nov
London86.68%: pass17.44 vs 18.04; −3.33%: pass1: fail£72.70 vs £72.49; +£0.21: passNo-go
South East86.40%: pass17.54 vs 17.89; −1.93%: pass1: fail£66.64 vs £66.20; +£0.44: passNo-go
Midlands86.35%: pass17.48 vs 17.89; −2.25%: pass1: fail£55.66 vs £55.35; +£0.31: passNo-go
North West86.28%: pass17.58 vs 18.47; −4.80%: pass0 recorded: observed pass£57.51 vs £57.29; +£0.22: passNo-go: shared safety gap
Scotland82.76%: pass over full test230.16 vs 18.20; +65.70%: fail0 recorded: observed pass£52.98 vs £52.62; +£0.36: passNo-go: safety and catalogue

Zero recorded incidents does not establish that an unprotected region is safe. Refunding a customer closes the complaint, not the underlying failure.

The aggregate claims conceal different problems.

Commercial’s acceptance claim is broadly correct: pooled acceptance rose from 71.20% to 85.93%, a 14.74 percentage-point improvement. Pooled refunds increased 7.86%, inside the overall guardrail. But the agreed criterion applies to every region. Scotland fails substantially.

Analytics’ £4 concern is also numerically correct but misleading as a treatment comparison: pooled baskets are £61.09 in treatment versus £65.15 in control. Treatment contains proportionally fewer high-value London orders and more lower-value Midlands and Scotland orders because allocation varies by region. Within every region, treatment baskets are higher across the full test. Applying the same combined regional order mix to both arms gives £63.26 versus £62.96, approximately £0.30 higher in treatment. This resolves the composition distortion; it does not establish a statistically reliable basket uplift.

Scotland is not settling down.

After the supplier switch on 15 October, treatment acceptance fell from 86.34% to 79.08%, below the launch threshold. Refunds rose from 17.25 to 43.25 per 1,000 orders, versus 17.79 in concurrent control, approximately 143% higher.

The final seven days still show 79.17% acceptance and 43.64 refunds per 1,000. The full-period acceptance pass masks a persistent failure in the current catalogue. The open pack-size mapping incident provides a plausible mechanism, and the repair has no date.

Conditions for reconsideration

  • Engineering and QA, by 4 November: deploy and verify the dietary filter across all six protected attributes. Replay the three incident cases, test missing or conflicting attributes and confirm unsafe or unverified candidates cannot be automatically selected. A merged change is insufficient.
  • PM, Analytics and regional Ops, review on 6 November:3 assess a limited post-fix pilot in London, South East, Midlands and North West, retaining randomised controls. Recheck every regional criterion, audit dietary compliance and verify London’s pricing fix. Expand only where the evidence supports it; insufficient evidence means continued limited exposure.1
  • Scotland Ops and catalogue Engineering, owner and repair date agreed on 30 October: complete and validate the re-map before restarting a controlled pilot. Judge readiness on post-repair results, not the earlier healthy fortnight.
  • During any restart: immediately disable auto-substitutions for a dietary mismatch and pause affected regions for guardrail breaches, with daily regional monitoring.

The 9 November booking window is a commercial deadline. It supplies no evidence that either unresolved defect will be fixed in time.

What a PM had to fix

  1. 1Decision deferredMake the callTargeted repair

    What we’d changeGive Commercial something to plan around: say whether the four clean regions can go by 9 November, set a threshold for pausing a region, and give Scotland a Christmas fallback.

  2. 2Contradiction missedSurface the contradictionQuick edit

    What we’d changeShow Scotland failing acceptance since the supplier switch in the table, as the text does, rather than passing.

  3. 3Test or gate too weakTighten the testQuick edit

    What we’d changeA 6 November review would have only a day or two of post-filter data. Allow enough to judge.

  4. 4OtherFix itQuick edit

    What we’d changeSay whether removing them changes any result. The brief asks for it.

Claims the judge couldn’t find in the brief

  • Refunding a customer closes the complaint, not the underlying failure.

Grades and run details

Decision model 77 · LLM judge 12 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly13%
  • passAddresses the actual decision94%
  • passRespects explicit constraints28%
  • passIdentifies material uncertainty85%
  • passAvoids unsupported claims32%
  • passProduces the required deliverable81%
  • passChecks every criterion, region by region58%
  • passFinds the Scotland breach100%
  • passExplains the basket gap98%
  • partialCatches the duplicate rows63%
  • passMakes a clear call77%
  • partialLimits the downside of being wrong63%
  • failGets the base of every number right30%
Run
Run
#1
Time to output
71 s
Submitted
25 Sept 2026

Usable with a quick edit10 of 13 checks passed

The recommendation, table, and calculations are solid; only minor wording additions to state duplicate impact and flag irreversible dietary harm are needed.

Output

Recommendation for 30 October: go with conditions in four regions; no-go in Scotland

Do not expand any region on Monday 2 November. London, South East, Midlands and North West are candidates for a gated rollout after the dietary hard filter is deployed and validated—not before its scheduled 4 November deployment. Scotland must additionally resolve and re-test its catalogue mapping.

Christmas urgency does not justify launching ahead of a known safety fix. The 9 November booking window should guide sequencing, not override the PRD.

Method and launch-criterion checks

I removed three exact duplicate North West treatment rows, dated 8–10 October, consistent with the export incident. This leaves 280 unique date–region–arm records.

Acceptance and refunds below use ratios of summed counts, not averages of daily rates. Basket values are order-weighted. T = treatment; C = control. Dietary counts are recorded treatment incidents; zero recorded incidents is not proof that none occurred.

RegionTreatment acceptance ≥80%Refunds/1,000 orders: T vs C; relative change ≤+10%Zero dietary/allergen mismatchesAverage basket: T vs C; difference ≥£0Recommendation
LondonPass: 86.7% (23,607/27,234)Pass: 17.44 vs 18.04; −3.3%Fail: 1, nut-free → hazelnuts, 23 OctPass: £72.70 vs £72.49; +£0.21Go with conditions; not 2 Nov
South EastPass: 86.4% (26,586/30,770)Pass: 17.54 vs 17.89; −1.9%Fail: 1, vegan → dairy, 19 OctPass: £66.64 vs £66.20; +£0.44Go with conditions; not 2 Nov
MidlandsPass: 86.3% (26,732/30,958)Pass: 17.48 vs 17.89; −2.3%Fail: 1, gluten-free → standard bread, 4 OctPass: £55.66 vs £55.35; +£0.31Go with conditions; not 2 Nov
North WestPass: 86.3% (18,153/21,039)Pass: 17.58 vs 18.47; −4.8%Pass on incident log: 0 recordedPass: £57.51 vs £57.29; +£0.22Go with conditions; not 2 Nov
ScotlandPass over 28 days: 82.8% (14,333/17,319)Fail: 30.16 vs 18.20; +65.7%Pass on incident log: 0 recordedPass: £52.98 vs £52.62; +£0.36No-go

What the data actually say

Commercial’s overall refund claim hides a regional failure. After deduplication, pooled refunds are 19.50 per 1,000 treatment orders versus 18.08 control: +7.9%, within the overall guardrail. But the PRD requires compliance in every region, and Scotland fails badly. Strong acceptance cannot compensate for dietary breaches.

Analytics’ £4 basket concern is real in the pooled arithmetic, but not within regions. Pooled baskets are £61.09 treatment versus £65.15 control (−£4.06). Treatment disproportionately includes lower-basket Midlands and Scotland orders, while control disproportionately includes higher-basket London orders. Every within-region comparison is positive. This is a composition effect, not evidence here that AI reduces basket value. These are point estimates; the daily averages alone do not establish statistical certainty.

Scotland is not settling down. Following the supplier switch on 15 October:

  • Acceptance falls from 86.3% on 1–14 October to 79.1% on 15–28 October, below the launch threshold.
  • Treatment refunds rise from 17.25 to 43.25 per 1,000 orders. Post-switch control is 17.79, making treatment 143.1% worse.
  • The final week still has 79.2% acceptance and 43.64 refunds per 1,000.

The open pack-size mapping incident is consistent with this sustained deterioration. The full-period acceptance pass masks the current operating state.

Conditions and release plan

For London, South East, Midlands and North West, require:

  1. Deployed, validated safety protection: successful regression tests across gluten-free, vegan, vegetarian, nut-free, halal and kosher attributes, including missing/incorrect metadata and the three observed failures. “Merged” is not “protected.”
  2. Fresh post-fix evidence: a controlled validation window with zero dietary mismatches and all regional numerical criteria met, with explicit launch sign-off. Refunds and apologies do not erase historical criterion failures.
  3. Staged regional expansion: monitor acceptance, refunds and dietary incidents daily; stop immediately for any dietary mismatch and roll back for guardrail breaches. Reconcile the North West export before reporting.

For Scotland, also require a completed catalogue re-map, pack-size validation and fresh post-remediation evidence. With no completion date and continuing deterioration, make no Christmas rollout commitment.

Meeting decision: reject the all-region 2 November launch; approve a gated path for four regions, while keeping Scotland blocked.

Grades and run details

Decision model 85 · LLM judge 12 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly15%
  • passAddresses the actual decision78%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty45%
  • passAvoids unsupported claims28%
  • passProduces the required deliverable77%
  • passChecks every criterion, region by region62%
  • passFinds the Scotland breach100%
  • passExplains the basket gap99%
  • partialCatches the duplicate rows72%
  • passMakes a clear call37%
  • partialLimits the downside of being wrong40%
  • failGets the base of every number right26%
Run
Run
#1
API response time
1.9 min
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyWrongMixedRight
Gemini 3.8 Flash · API

It contains material unsupported or incorrect current-situation claims, including North West GO despite a dietary incident and a wrong duplicate-row figure.

GPT-6 Astra · ChatGPT

All factual statements about the current situation are supported by the brief or correct arithmetic; no invented facts.

GPT-6.1 Sol · API

All current-state numbers and facts trace to the brief, incidents.csv, or arithmetic from daily_metrics.csv.

Identifies material uncertaintyWrongRightRight
Gemini 3.8 Flash · API

It does not clearly name the key unknowns that would change the call, such as the hard filter slipping past 9 November or Scotland failing to recover after re-map.

GPT-6 Astra · ChatGPT

Names the unknowns (filter delivery, Scotland catalogue repair, post-fix evidence) and says what would change the call.

GPT-6.1 Sol · API

Names filter deployment/validation and Scotland catalogue remap as unresolved and ties go decisions to fresh evidence.

Avoids unsupported claimsWrongMixedRight
Gemini 3.8 Flash · API

It presents North West as clean and safe to launch despite an open dietary incident, and asserts duplicate-row removal with an incorrect figure.

GPT-6 Astra · ChatGPT

Presents 'Refunding a customer closes the complaint, not the underlying failure' as fact without support from the supplied context.

GPT-6.1 Sol · API

Causal claims are hedged ('consistent with') and interpretations are labelled as composition effects or point estimates.

Checks every criterion, region by regionWrongRightRight
Gemini 3.8 Flash · API

The table checks all four criteria by region, but the North West dietary result is wrong and the duplicate handling figure is wrong.

GPT-6 Astra · ChatGPT

Table checks every criterion for each region using figures computed from the attached files.

GPT-6.1 Sol · API

Region-by-criterion table covers acceptance, refunds vs control, dietary incidents, and basket values with computed figures.

Makes a clear callMixedRightRight
Gemini 3.8 Flash · API

It gives dates and conditions, but the North West call is wrong and it does not state what would change the overall call.

GPT-6 Astra · ChatGPT

One clear call per region, with conditions and dates, and says what would change it.

GPT-6.1 Sol · API

Makes a clear per-region call with dates and conditions, and Scotland is blocked.

Limits the downside of being wrongWrongRightWrong
Gemini 3.8 Flash · API

It lacks post-launch monitoring signals, pause/rollback thresholds, and irreversible-risk triggers.

GPT-6 Astra · ChatGPT

Proposes daily regional monitoring, a trigger (dietary mismatch or guardrail breach) and action (disable, pause).

GPT-6.1 Sol · API

Gives monitoring and rollback thresholds but does not flag irreversible consequences such as customer-facing dietary substitutions or communications.

Gets the base of every number rightWrongMixedMixed
Gemini 3.8 Flash · API

Several derived figures are correct, but the duplicate-row count and North West dietary status are wrong.

GPT-6 Astra · ChatGPT

All derived percentages and rates use correct numerators and denominators, and basket values are weighted by orders.

GPT-6.1 Sol · API

All percentages, per-1,000 rates, and basket differences use correct summed or order-weighted bases.

All got wrong 1

Catches the duplicate rowsWrongWrongWrong
Gemini 3.8 Flash · API

It notices the North West duplicates but misstates them as 3,194 duplicate rows rather than 2,194 duplicate orders across three rows.

GPT-6 Astra · ChatGPT

Finds and removes the duplicate North West rows but does not explicitly state whether they change the overall assessment.

GPT-6.1 Sol · API

Finds and removes the duplicate rows but never states whether they change the results.

All got right 5

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

It commits early to a clear region-by-region call and states conditions and dates.

GPT-6 Astra · ChatGPT

Commits to a clear no-go on 2 Nov, then conditional rollout by region, and states what would change the call.

GPT-6.1 Sol · API

Commits to a clear region-by-region go/no-go with dates and conditions for the launch meeting.

Respects explicit constraintsRightRightRight
Gemini 3.8 Flash · API

It is under 900 words, includes a region-by-criterion table, and respects the launch-criteria framing.

GPT-6 Astra · ChatGPT

Respects the requested form, reader, word limit, and uses the data rather than stakeholder summaries.

GPT-6.1 Sol · API

Includes the required region-by-criterion table and keeps prose within the length while respecting the no-early-launch safety constraint.

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

The requested recommendation and table are present and usable in form, though some figures need correction.

GPT-6 Astra · ChatGPT

A complete recommendation with a criteria table, under 900 words, usable for the launch meeting.

GPT-6.1 Sol · API

Provides a usable recommendation, table, and release conditions that the reader could act on with light edits.

Finds the Scotland breachRightRightRight
Gemini 3.8 Flash · API

It identifies Scotland's refund breach, ties it to the supplier change, notes acceptance dip, and holds Scotland.

GPT-6 Astra · ChatGPT

Identifies the Scotland refund breach, ties it to the 15 October supplier change, notes acceptance dip, and blocks Scotland.

GPT-6.1 Sol · API

Finds Scotland's refund breach, ties it to the supplier change with post-15 Oct numbers, and holds Scotland.

Explains the basket gapRightRightRight
Gemini 3.8 Flash · API

It compares within-region baskets and explains the pooled gap as a Simpson's Paradox mix effect.

GPT-6 Astra · ChatGPT

Shows treatment baskets are higher within every region and explains the pooled gap as a mix artefact.

GPT-6.1 Sol · API

Explains the pooled basket gap as a traffic-mix effect and shows all within-region treatment baskets positive.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.388.32None
2GPT-6 AstrawithChatGPT86.088.32None
3Sonnet 5.5withAPI85.488.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI82.177.62None
6Gemini 3.8 FlashwithAPI71.931.521 capped
7Gemini 3.5 Flash-LitewithGemini50.642.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook