Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Addresses the actual decision95% pass
    The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong32% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows50% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Uses the supplied evidence correctly54% pass
    The output states that pilot agents were hands-on, engaged and motivated, which is not supported by the supplied context (only that 7/8 wanted to keep it).
    Opus 5.5 · Claude · Go/no-go for AI-drafted support replies

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Staff PM for fulfilment at Basketful. On 30 October the launch meeting decides whether AI auto-substitutions (the model picks a replacement when an item is out of stock, instead of the picker) roll out to every region on Monday 2 November, ahead of the Christmas delivery-slot booking window that opens on 9 November. The rollout workbook is attached as two files: daily_metrics.csv (28 days of the regional test, by region and arm) and incidents.csv. Work from the data, not the summaries people have given you. Write your recommendation for the meeting: go, no-go or go with conditions, region by region, and why. Include a table that checks each launch criterion for each region with the figures. Keep the prose under 900 words.

What the model was given6 items: The test, Launch criteria (agreed in the PRD), What people have said, Engineering, daily_metrics.csv, incidents.csv
The test1–28 October, five regions. In each region orders were split between control (the picker chooses substitutes, as today) and treatment (AI auto-substitutions). The split was not 50/50 everywhere: treatment got 30% of orders in London, 50% in the South East and North West, and 70% in the Midlands and Scotland.
Launch criteria (agreed in the PRD)1. Substitution acceptance (subs_accepted ÷ subs_offered) of at least 80% in treatment, in every region. 2. Refund requests per 1,000 orders in treatment no more than 10% above control, in every region. 3. Zero dietary or allergen mismatches (a substitute that breaks a gluten-free, vegan, vegetarian, nut-free, halal or kosher attribute). 4. Average basket value not lower in treatment than control.
What people have saidHead of Commercial: “Acceptance is up 15 points and refunds are within the guardrail overall. Every day we wait costs us Christmas.” Head of Analytics: “Treatment baskets are £4 smaller across the test. That worries me.” Ops director (Scotland): “The supplier change caused some noise, but it's settling down.”
EngineeringThe dietary-attribute hard filter (never substitute across those attributes) is merged and scheduled to deploy on 4 November after regression tests. Scotland's catalogue re-map for the new dairy supplier is in progress with no date. Rollout can be switched on region by region.
daily_metrics.csv283 rows · Download
date,region,arm,orders,items_ordered,items_out_of_stock,subs_offered,subs_accepted,refund_requests,refund_value_gbp,complaints,avg_basket_gbp
2026-10-01,London,control,1757,52273,2351,1953,1387,32,160.71,6,72.89
2026-10-01,London,treatment,692,20455,871,820,708,12,62.05,2,72.42
2026-10-01,South East,control,853,22846,996,833,593,16,74.49,3,65.62
2026-10-01,South East,treatment,928,25223,1166,1106,952,15,75.93,3,66.59
2026-10-01,Midlands,control,467,10468,465,383,268,8,34.57,1,54.92
2026-10-01,Midlands,treatment,1113,25136,1166,1087,947,21,106.35,3,55.33
…
incidents.csv13 rows · Download
date,region,arm,type,severity,summary,status
2026-10-02,London,treatment,wrong_size,low,"4-pint milk substituted with 1-pint, customer accepted at door",closed
2026-10-04,Midlands,treatment,dietary_mismatch,high,"Gluten-free sliced loaf substituted with standard white loaf; customer coeliac, noticed at home",closed: refund and apology
2026-10-06,South East,control,wrong_item,low,"Picker substituted oat milk with soya milk; customer rejected at door",closed
2026-10-07,North West,treatment,data_export,low,"Daily export job re-ran after a timeout on 8, 9 and 10 Oct; analytics team flagged possible duplicate rows",open
2026-10-09,London,treatment,price,medium,"Substitute priced higher than original; customer charged the difference against policy",closed: policy fix deployed 11 Oct
2026-10-12,Scotland,control,late_delivery,low,"Van breakdown, 14 orders late",closed
…
What a strong answer doesThe answer key the graders mark against

No-go for a national launch on 2 November, because criterion 3 fails everywhere until the dietary hard filter ships: incidents.csv has three treatment dietary mismatches (4 Oct Midlands, 19 Oct South East, 23 Oct London, the last involving nuts and a child with an allergy). Go with conditions, region by region, once the filter is deployed and verified (planned 4 November, still before the 9 November booking window): London, South East, Midlands and North West pass the other criteria (acceptance about 86% against 71% control; refunds 2–5% below control). Hold Scotland: over the month its treatment refunds are about 66% above control, and from the 15 October supplier change about 143% above (43 against 18 per 1,000 orders), with acceptance falling to about 79%, below the 80% bar; it waits for the catalogue re-map and a clean re-test. The pooled refund rate (about 8% above control) hides this, and the criterion is per region. The pooled basket gap (about £65 against £61) is an artefact of the uneven split, since treatment baskets are equal or higher in every region (Simpson's paradox), so criterion 4 passes. Notices that North West treatment rows for 8–10 October are duplicated (an extra 2,194 orders), removes them, and says it doesn't change the result. States what would change the call: the filter slipping past 9 November, or Scotland's refunds not recovering.

Critical failures (cap the score at 40)
  • Recommends any rollout before the dietary hard filter is deployed
  • Recommends rolling out to Scotland while its refund guardrail is breached
Case

Staff-level · v1.8 · synthetic · e-commerce, AI product, data analysis, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit6 of 13 checks passed

It fails to address the known duplicate data rows, doesn't specify monitoring triggers or rollback criteria, and makes a causal claim about Scotland without hedging, so a PM would need to add those before using it.

Output

Launch Recommendation: AI Auto-Substitutions

As the Staff PM for fulfilment, my recommendation for the 30 October launch meeting is a No-Go for nationwide rollout on 2 November.

While the Commercial and Analytics teams are looking at top-line aggregates, a region-by-region audit of the 28-day test data and incident logs reveals critical blockers. Specifically, Scotland fails acceptance and refund criteria due to a supplier mapping failure, Midlands and London recorded high-severity dietary/allergen mismatches, and South East also experienced a dietary mismatch.

We should hold all regions until the dietary hard filter is deployed on 4 November, and until Scotland’s catalogue re-map is fully completed and verified.

---

Regional Evaluation Table (1–28 October Test Data)

RegionSubs Acceptance ($\ge 80\%$)Treatment Refunds vs Control ($\le +10\%$)Dietary/Allergen Mismatches (Target: 0)Basket Value ImpactRecommendation
:---:---:---:---:---:---
London$86.4\%$ (Pass)Control: 1.77%<br>Treatment: 1.63% ($-7.9\%$)1 High Incident (Nut allergy violation on 23 Oct)Control: £72.33<br>Treatment: £72.71 (+0.5%)No-Go (Hold for 4 Nov filter)
South East$87.0\%$ (Pass)Control: 1.66%<br>Treatment: 1.69% ($+1.8\%$)1 High Incident (Vegan dairy violation on 19 Oct)Control: £66.21<br>Treatment: £66.72 (+0.8%)No-Go (Hold for 4 Nov filter)
Midlands$84.5\%$ (Pass)Control: 1.69%<br>Treatment: 1.68% ($-0.6\%$)1 High Incident (Gluten-free coeliac violation on 4 Oct)Control: £55.22<br>Treatment: £55.51 (+0.5%)No-Go (Hold for 4 Nov filter)
North West$88.4\%$ (Pass)Control: 1.73%<br>Treatment: 1.66% ($-4.0\%$)0 IncidentsControl: £57.34<br>Treatment: £57.65 (+0.5%)Go with Conditions (Deploy 4 Nov)
Scotland$77.0\%$ (Fail)Control: 1.79%<br>Treatment: $3.90\%$ ($+117.9\%$)0 Incidents (High dairy wrong-size volume)Control: £52.48<br>Treatment: £52.88 (+0.8%)No-Go (Hold for catalogue re-map)

Note: Refund request rates are calculated per 100 orders for precise comparison against the 10% relative ceiling.

---

Region-by-Region Breakdown

  • London, South East, Midlands: All three regions easily clear the 80% acceptance threshold and show neutral-to-positive average basket values (debunking the £4 basket drop concern). However, all three recorded high-severity dietary/allergen failures (gluten, vegan, and nut-allergy breaches). Because the safety hard filter is not deployed until 4 November, rolling out on 2 November exposes customers to unacceptable allergen risks during the peak Christmas slot-booking window.
  • North West: Our strongest performer. It meets all four PRD criteria cleanly3: 88.4% acceptance1, treatment refund rates slightly below control, zero dietary incidents, and a stable basket value. It is ready to roll out immediately following the hard filter deployment.
  • Scotland: A clear No-Go. Acceptance dropped to 77.0% (failing the 80% bar), and treatment refund requests spiked catastrophically by 117.9% over control (driven by the dairy supplier switch on 15 Oct where pack sizes were mismapped). Until the catalogue re-map is finished and stabilized, rolling out here will compound operational friction.

Action Plan

  1. Hold Rollout on 2 November.
  2. Deploy Dietary Hard Filter on 4 November as scheduled, and verify regression tests.
  3. Complete Scotland Catalogue Re-map and clear open wrong-size substitution incidents.
  4. Staged Launch: Roll out to North West, London, South East, and Midlands on Monday, 9 November (concurrent with the Christmas slot-booking window opening), pending a 7-day post-fix stability review.2 Scotland follows only after catalogue remediation.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticStart again

    What we’d changeRecompute every figure from daily_metrics.csv, removing the duplicated North West rows. Almost every number in the table is wrong.

  2. 2Contradiction missedSurface the contradictionSubstantial rework

    What we’d changeResolve the contradiction: a 7-day review after the 4 November filter can't finish by 9 November, and North West is marked go in the table but held here. Answer the stakeholders' claims rather than dismissing them.

  3. 3Hypothesis stated as factReframe it as a hypothesisTargeted repair

    What we’d changeZero recorded mismatches in North West isn't evidence of safety. It runs the same unfiltered model as the regions that had them.

Grades and run details

Decision model 46 · LLM judge 8 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly49%
  • passAddresses the actual decision32%
  • passRespects explicit constraints8%
  • partialIdentifies material uncertainty63%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable26%
  • partialChecks every criterion, region by region64%
  • passFinds the Scotland breach96%
  • failExplains the basket gap48%
  • failCatches the duplicate rows100%
  • partialMakes a clear call83%
  • failLimits the downside of being wrong41%
  • failGets the base of every number right77%
Run
Run
#1
Time to output
46 s
Submitted
25 Sept 2026

Needs more than a quick edit9 of 13 checks passed

The refund figures are inaccurate, so a PM would need to re-check the data and recalculate before using the recommendation. The output is not usable as-is.

Output

Recommendation: No-go for the all-region rollout on 2 November

Do not enable auto-substitutions anywhere on 2 November: the dietary hard filter is not scheduled to deploy until 4 November, so the zero-mismatch safety criterion is not yet protected. After deployment and validation, North West is the only region I’d consider enabling first, with conditions. Hold the other regions for the reasons below.

Figures are for 1–28 October; refund rates are per 1,000 orders. Basket averages are order-weighted within each region. I excluded the three exact duplicate North West treatment rows on 8, 9 and 10 October.

RegionTreatment acceptance (≥80%)Refunds / 1,000 orders: treatment vs control (treatment ≤10% higher)Dietary/allergen mismatches (must be zero)Average basket: treatment vs control (must not be lower)Decision
------:---:------:---
London~86% Pass~17.9 vs ~18.4; Pass1; Fail — nut-containing bar substituted for nut-free item~£72.8 vs ~£72.5; PassNo-go
South East~85% Pass~17.8 vs ~17.9; Pass1; Fail — dairy cheese substituted for vegan cheese~£66.6 vs ~£66.3; PassNo-go
Midlands~86% Pass~17.6 vs ~18.0; Pass1; Fail — standard loaf substituted for gluten-free loaf~£55.6 vs ~£55.4; PassNo-go
North West~86% Pass~17.8 vs ~18.8; Pass0 observed; Pass in test~£57.5 vs ~£57.3; PassGo with conditions, after filter validation
Scotland~79% Fail1~29.8 vs ~18.4; Fail — about 62% higher0 observed; Pass in test~£53.0 vs ~£52.6; PassNo-go

Conditions and rationale

  • North West: After the hard filter deploys, pass regression tests across all six protected attributes (gluten-free, vegan, vegetarian, nut-free, halal and kosher); then enable regionally with close monitoring and a rollback switch. The export duplicates should also be fixed or guarded against before relying on daily reporting.
  • London, South East and Midlands: Do not launch until the filter is deployed and the dietary-mismatch failures are investigated and validated as fixed. Their good acceptance, refund and basket figures do not offset a safety-criterion failure.
  • Scotland: Hold until the supplier catalogue re-map is complete and verified, and a fresh test meets both acceptance and refund guardrails. The sharp rise in treatment refunds and lower acceptance coincide with the dairy supplier/catalogue change; the issue is open, with no completion date.

The pooled “£4 smaller” basket claim should not drive this decision: treatment exposure differed by region, and baskets vary substantially by region. Within each region, the treatment average is slightly higher in these data. Likewise, an overall acceptance or refund result cannot override region-level launch gates. The 9 November booking window is important, but it does not justify launching before the safety control is live.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    What we’d change82.76%: Pass over 1–28 October. However, acceptance fell to 79.08% over 15–28 October after the supplier change, below the launch threshold.

Claims the judge couldn’t find in the brief

  • London treatment refunds per 1,000 orders is ~17.9.
  • London control refunds per 1,000 orders is ~18.4.
  • South East treatment refunds per 1,000 orders is ~17.8.
  • Midlands treatment refunds per 1,000 orders is ~17.6.
  • North West treatment refunds per 1,000 orders is ~17.8 (after excluding duplicates).
  • Scotland control refunds per 1,000 orders is ~18.4.

Grades and run details

Decision model 69 · LLM judge 9 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly8%
  • passAddresses the actual decision47%
  • passRespects explicit constraints61%
  • partialIdentifies material uncertainty55%
  • passAvoids unsupported claims21%
  • passProduces the required deliverable63%
  • passChecks every criterion, region by region28%
  • passFinds the Scotland breach88%
  • passExplains the basket gap86%
  • partialCatches the duplicate rows69%
  • partialMakes a clear call62%
  • partialLimits the downside of being wrong96%
  • failGets the base of every number right45%
Run
Run
#1
API response time
44 s
Submitted
29 Sept 2026

Usable with a quick edit10 of 13 checks passed

The recommendation, table, and calculations are solid; only minor wording additions to state duplicate impact and flag irreversible dietary harm are needed.

Output

Recommendation for 30 October: go with conditions in four regions; no-go in Scotland

Do not expand any region on Monday 2 November. London, South East, Midlands and North West are candidates for a gated rollout after the dietary hard filter is deployed and validated—not before its scheduled 4 November deployment. Scotland must additionally resolve and re-test its catalogue mapping.

Christmas urgency does not justify launching ahead of a known safety fix. The 9 November booking window should guide sequencing, not override the PRD.

Method and launch-criterion checks

I removed three exact duplicate North West treatment rows, dated 8–10 October, consistent with the export incident. This leaves 280 unique date–region–arm records.

Acceptance and refunds below use ratios of summed counts, not averages of daily rates. Basket values are order-weighted. T = treatment; C = control. Dietary counts are recorded treatment incidents; zero recorded incidents is not proof that none occurred.

RegionTreatment acceptance ≥80%Refunds/1,000 orders: T vs C; relative change ≤+10%Zero dietary/allergen mismatchesAverage basket: T vs C; difference ≥£0Recommendation
LondonPass: 86.7% (23,607/27,234)Pass: 17.44 vs 18.04; −3.3%Fail: 1, nut-free → hazelnuts, 23 OctPass: £72.70 vs £72.49; +£0.21Go with conditions; not 2 Nov
South EastPass: 86.4% (26,586/30,770)Pass: 17.54 vs 17.89; −1.9%Fail: 1, vegan → dairy, 19 OctPass: £66.64 vs £66.20; +£0.44Go with conditions; not 2 Nov
MidlandsPass: 86.3% (26,732/30,958)Pass: 17.48 vs 17.89; −2.3%Fail: 1, gluten-free → standard bread, 4 OctPass: £55.66 vs £55.35; +£0.31Go with conditions; not 2 Nov
North WestPass: 86.3% (18,153/21,039)Pass: 17.58 vs 18.47; −4.8%Pass on incident log: 0 recordedPass: £57.51 vs £57.29; +£0.22Go with conditions; not 2 Nov
ScotlandPass over 28 days: 82.8% (14,333/17,319)Fail: 30.16 vs 18.20; +65.7%Pass on incident log: 0 recordedPass: £52.98 vs £52.62; +£0.36No-go

What the data actually say

Commercial’s overall refund claim hides a regional failure. After deduplication, pooled refunds are 19.50 per 1,000 treatment orders versus 18.08 control: +7.9%, within the overall guardrail. But the PRD requires compliance in every region, and Scotland fails badly. Strong acceptance cannot compensate for dietary breaches.

Analytics’ £4 basket concern is real in the pooled arithmetic, but not within regions. Pooled baskets are £61.09 treatment versus £65.15 control (−£4.06). Treatment disproportionately includes lower-basket Midlands and Scotland orders, while control disproportionately includes higher-basket London orders. Every within-region comparison is positive. This is a composition effect, not evidence here that AI reduces basket value. These are point estimates; the daily averages alone do not establish statistical certainty.

Scotland is not settling down. Following the supplier switch on 15 October:

  • Acceptance falls from 86.3% on 1–14 October to 79.1% on 15–28 October, below the launch threshold.
  • Treatment refunds rise from 17.25 to 43.25 per 1,000 orders. Post-switch control is 17.79, making treatment 143.1% worse.
  • The final week still has 79.2% acceptance and 43.64 refunds per 1,000.

The open pack-size mapping incident is consistent with this sustained deterioration. The full-period acceptance pass masks the current operating state.

Conditions and release plan

For London, South East, Midlands and North West, require:

  1. Deployed, validated safety protection: successful regression tests across gluten-free, vegan, vegetarian, nut-free, halal and kosher attributes, including missing/incorrect metadata and the three observed failures. “Merged” is not “protected.”
  2. Fresh post-fix evidence: a controlled validation window with zero dietary mismatches and all regional numerical criteria met, with explicit launch sign-off. Refunds and apologies do not erase historical criterion failures.
  3. Staged regional expansion: monitor acceptance, refunds and dietary incidents daily; stop immediately for any dietary mismatch and roll back for guardrail breaches. Reconcile the North West export before reporting.

For Scotland, also require a completed catalogue re-map, pack-size validation and fresh post-remediation evidence. With no completion date and continuing deterioration, make no Christmas rollout commitment.

Meeting decision: reject the all-region 2 November launch; approve a gated path for four regions, while keeping Scotland blocked.

Grades and run details

Decision model 85 · LLM judge 12 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly15%
  • passAddresses the actual decision78%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty45%
  • passAvoids unsupported claims28%
  • passProduces the required deliverable77%
  • passChecks every criterion, region by region62%
  • passFinds the Scotland breach100%
  • passExplains the basket gap99%
  • partialCatches the duplicate rows72%
  • passMakes a clear call37%
  • partialLimits the downside of being wrong40%
  • failGets the base of every number right26%
Run
Run
#1
API response time
1.9 min
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyWrongWrongRight
Gemini 3.5 Flash-Lite · Gemini

The output calculates rates from the raw daily_metrics.csv without acknowledging or removing the duplicate North West treatment rows flagged in incidents.csv, so the claimed rates for North West (and potentially others if the duplication affects summaries) are not fully reliable.

GPT-6 Luna · API

Several refund-per-1000 figures are incorrect (e.g., London treatment 17.9 vs actual 17.4, London control 18.4 vs 18.0), so the output does not use the supplied evidence correctly.

GPT-6.1 Sol · API

All current-state numbers and facts trace to the brief, incidents.csv, or arithmetic from daily_metrics.csv.

Identifies material uncertaintyWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

Does not explicitly name the unknowns that could change the call (e.g., what if the dietary filter slips past 9 November, or if Scotland’s refunds don’t recover after the catalogue fix) or state the thresholds that would reverse the decision.

GPT-6 Luna · API

Names the filter deployment, duplicate fix, and Scotland catalogue re-map as unknowns that could change the decision.

GPT-6.1 Sol · API

Names filter deployment/validation and Scotland catalogue remap as unresolved and ties go decisions to fresh evidence.

Avoids unsupported claimsWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

States as fact that Scotland’s refund spike was 'driven by the dairy supplier switch where pack sizes were mismapped' without labelling it as a hypothesis, even though the data only shows a correlation.

GPT-6 Luna · API

Causal claims like the supplier change coinciding with refunds are presented as observations, not unsupported facts.

GPT-6.1 Sol · API

Causal claims are hedged ('consistent with') and interpretations are labelled as composition effects or point estimates.

Checks every criterion, region by regionRightMixedRight
Gemini 3.5 Flash-Lite · Gemini

The table checks acceptance, refunds vs control, dietary incidents and basket value for each of the five regions, with figures drawn from the test data.

GPT-6 Luna · API

The table checks each criterion per region but the refund figures are incorrect, so it does not present correct figures from the files.

GPT-6.1 Sol · API

Region-by-criterion table covers acceptance, refunds vs control, dietary incidents, and basket values with computed figures.

Explains the basket gapMixedRightRight
Gemini 3.5 Flash-Lite · Gemini

Shows treatment basket values are equal or higher in every region and states the aggregated £4 drop concern is debunked, satisfying the Simpson’s paradox explanation.

GPT-6 Luna · API

Explains the pooled basket gap as a mix effect, noting treatment baskets are equal or higher within each region.

GPT-6.1 Sol · API

Explains the pooled basket gap as a traffic-mix effect and shows all within-region treatment baskets positive.

Catches the duplicate rowsWrongRightWrong
Gemini 3.5 Flash-Lite · Gemini

Does not mention or handle the duplicate North West treatment rows on 8–10 October that the incident log flags, so the quality check is absent.

GPT-6 Luna · API

Catches the duplicate North West rows, excludes them, and recommends fixing the export job.

GPT-6.1 Sol · API

Finds and removes the duplicate rows but never states whether they change the results.

Gets the base of every number rightMixedWrongMixed
Gemini 3.5 Flash-Lite · Gemini

Rates are computed as sums over the period (acceptance = total accepted / total offered, refund rate = total refunds / total orders), so denominators are consistent and the relative comparisons are valid.

GPT-6 Luna · API

Multiple refund-per-1000 figures are miscalculated (e.g., London treatment 17.9 vs 17.4, London control 18.4 vs 18.0), so the base numbers are wrong.

GPT-6.1 Sol · API

All percentages, per-1,000 rates, and basket differences use correct summed or order-weighted bases.

All got wrong 1

Limits the downside of being wrongWrongWrongWrong
Gemini 3.5 Flash-Lite · Gemini

While it mentions a 7‑day post‑fix stability review, it does not specify a signal (e.g., refund rate threshold) or a rollback trigger, and does not flag irreversible actions.

GPT-6 Luna · API

Does not name a specific post-launch signal, threshold, or trigger for rollback; only says 'close monitoring and a rollback switch' without specifics.

GPT-6.1 Sol · API

Gives monitoring and rollback thresholds but does not flag irreversible consequences such as customer-facing dietary substitutions or communications.

All got right 5

Addresses the actual decisionRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Commits to a clear No‑Go for 2 November, then a conditional Go by region for 9 November after the filter deploys, and states Scotland’s hold conditions.

GPT-6 Luna · API

Commits to a clear no-go for all-region on 2 Nov, then conditional go for North West after filter, with conditions and what would change it.

GPT-6.1 Sol · API

Commits to a clear region-by-region go/no-go with dates and conditions for the launch meeting.

Respects explicit constraintsRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Prose is well under 900 words, addressed to the launch meeting, and includes a region‑by‑region table with the four criteria.

GPT-6 Luna · API

Includes the required region-by-region table, prose under 900 words, and addresses the launch meeting.

GPT-6.1 Sol · API

Includes the required region-by-criterion table and keeps prose within the length while respecting the no-early-launch safety constraint.

Produces the required deliverableRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Provides a complete, table‑based recommendation with a region‑by‑region call, within the word limit, that the launch meeting could act on.

GPT-6 Luna · API

Provides a complete recommendation with table, region-by-region, usable for the meeting.

GPT-6.1 Sol · API

Provides a usable recommendation, table, and release conditions that the reader could act on with light edits.

Finds the Scotland breachRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Identifies Scotland’s refund breach (+117.9%), links it to the post‑15 October supplier change, notes the acceptance drop to 77%, and holds Scotland.

GPT-6 Luna · API

Identifies Scotland's refund breach, ties it to the supplier change, notes acceptance dip, and holds Scotland.

GPT-6.1 Sol · API

Finds Scotland's refund breach, ties it to the supplier change with post-15 Oct numbers, and holds Scotland.

Makes a clear callRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Gives a clear decision per region (No‑Go for Scotland, conditional Go for others from 9 Nov) with dates and the conditions required to proceed.

GPT-6 Luna · API

Makes a clear call per region with dates and conditions, and says what would change it.

GPT-6.1 Sol · API

Makes a clear per-region call with dates and conditions, and Scotland is blocked.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.388.32None
2GPT-6 AstrawithChatGPT86.088.32None
3Sonnet 5.5withAPI85.488.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI82.177.62None
6Gemini 3.8 FlashwithAPI71.931.521 capped
7Gemini 3.5 Flash-LitewithGemini50.642.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook