Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Addresses the actual decision95% pass
    The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong32% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows50% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Uses the supplied evidence correctly54% pass
    The output states that pilot agents were hands-on, engaged and motivated, which is not supported by the supplied context (only that 7/8 wanted to keep it).
    Opus 5.5 · Claude · Go/no-go for AI-drafted support replies

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Staff PM for fulfilment at Basketful. On 30 October the launch meeting decides whether AI auto-substitutions (the model picks a replacement when an item is out of stock, instead of the picker) roll out to every region on Monday 2 November, ahead of the Christmas delivery-slot booking window that opens on 9 November. The rollout workbook is attached as two files: daily_metrics.csv (28 days of the regional test, by region and arm) and incidents.csv. Work from the data, not the summaries people have given you. Write your recommendation for the meeting: go, no-go or go with conditions, region by region, and why. Include a table that checks each launch criterion for each region with the figures. Keep the prose under 900 words.

What the model was given6 items: The test, Launch criteria (agreed in the PRD), What people have said, Engineering, daily_metrics.csv, incidents.csv
The test1–28 October, five regions. In each region orders were split between control (the picker chooses substitutes, as today) and treatment (AI auto-substitutions). The split was not 50/50 everywhere: treatment got 30% of orders in London, 50% in the South East and North West, and 70% in the Midlands and Scotland.
Launch criteria (agreed in the PRD)1. Substitution acceptance (subs_accepted ÷ subs_offered) of at least 80% in treatment, in every region. 2. Refund requests per 1,000 orders in treatment no more than 10% above control, in every region. 3. Zero dietary or allergen mismatches (a substitute that breaks a gluten-free, vegan, vegetarian, nut-free, halal or kosher attribute). 4. Average basket value not lower in treatment than control.
What people have saidHead of Commercial: “Acceptance is up 15 points and refunds are within the guardrail overall. Every day we wait costs us Christmas.” Head of Analytics: “Treatment baskets are £4 smaller across the test. That worries me.” Ops director (Scotland): “The supplier change caused some noise, but it's settling down.”
EngineeringThe dietary-attribute hard filter (never substitute across those attributes) is merged and scheduled to deploy on 4 November after regression tests. Scotland's catalogue re-map for the new dairy supplier is in progress with no date. Rollout can be switched on region by region.
daily_metrics.csv283 rows · Download
date,region,arm,orders,items_ordered,items_out_of_stock,subs_offered,subs_accepted,refund_requests,refund_value_gbp,complaints,avg_basket_gbp
2026-10-01,London,control,1757,52273,2351,1953,1387,32,160.71,6,72.89
2026-10-01,London,treatment,692,20455,871,820,708,12,62.05,2,72.42
2026-10-01,South East,control,853,22846,996,833,593,16,74.49,3,65.62
2026-10-01,South East,treatment,928,25223,1166,1106,952,15,75.93,3,66.59
2026-10-01,Midlands,control,467,10468,465,383,268,8,34.57,1,54.92
2026-10-01,Midlands,treatment,1113,25136,1166,1087,947,21,106.35,3,55.33
…
incidents.csv13 rows · Download
date,region,arm,type,severity,summary,status
2026-10-02,London,treatment,wrong_size,low,"4-pint milk substituted with 1-pint, customer accepted at door",closed
2026-10-04,Midlands,treatment,dietary_mismatch,high,"Gluten-free sliced loaf substituted with standard white loaf; customer coeliac, noticed at home",closed: refund and apology
2026-10-06,South East,control,wrong_item,low,"Picker substituted oat milk with soya milk; customer rejected at door",closed
2026-10-07,North West,treatment,data_export,low,"Daily export job re-ran after a timeout on 8, 9 and 10 Oct; analytics team flagged possible duplicate rows",open
2026-10-09,London,treatment,price,medium,"Substitute priced higher than original; customer charged the difference against policy",closed: policy fix deployed 11 Oct
2026-10-12,Scotland,control,late_delivery,low,"Van breakdown, 14 orders late",closed
…
What a strong answer doesThe answer key the graders mark against

No-go for a national launch on 2 November, because criterion 3 fails everywhere until the dietary hard filter ships: incidents.csv has three treatment dietary mismatches (4 Oct Midlands, 19 Oct South East, 23 Oct London, the last involving nuts and a child with an allergy). Go with conditions, region by region, once the filter is deployed and verified (planned 4 November, still before the 9 November booking window): London, South East, Midlands and North West pass the other criteria (acceptance about 86% against 71% control; refunds 2–5% below control). Hold Scotland: over the month its treatment refunds are about 66% above control, and from the 15 October supplier change about 143% above (43 against 18 per 1,000 orders), with acceptance falling to about 79%, below the 80% bar; it waits for the catalogue re-map and a clean re-test. The pooled refund rate (about 8% above control) hides this, and the criterion is per region. The pooled basket gap (about £65 against £61) is an artefact of the uneven split, since treatment baskets are equal or higher in every region (Simpson's paradox), so criterion 4 passes. Notices that North West treatment rows for 8–10 October are duplicated (an extra 2,194 orders), removes them, and says it doesn't change the result. States what would change the call: the filter slipping past 9 November, or Scotland's refunds not recovering.

Critical failures (cap the score at 40)
  • Recommends any rollout before the dietary hard filter is deployed
  • Recommends rolling out to Scotland while its refund guardrail is breached
Case

Staff-level · v1.8 · synthetic · e-commerce, AI product, data analysis, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 13 checks passed

The output is fully reasoned, includes the required table and conditions, and could be used with a quick edit to remove the one unsupported sentence.

Output

Recommendation for 30 October: no-go in every region on 2 November. Conditional phased rollout thereafter in London, South East, Midlands and North West. Scotland remains blocked pending catalogue repair and fresh evidence.

The dietary filter is scheduled for 4 November, two days after the proposed launch. Three treatment substitutions breached dietary requirements, including gluten-free bread replaced with standard bread and nut-free bars replaced with hazelnut bars. The shared system needs this control before expansion, including in regions with no recorded mismatch.

I would pause existing auto-substitution exposure and return to picker selection until the filter is deployed and verified.

Regional checks, 1–28 October

I removed three exact duplicate North West treatment rows dated 8–10 October4, consistent with the export incident, leaving 280 unique region–arm–day records. Acceptance and refund rates use summed numerators and denominators. Basket values are weighted by orders.

T = treatment; C = control. Refund changes are relative changes in requests per 1,000 orders.

RegionAcceptance T ≥80%Refunds/1,000: T vs C; change ≤+10%Dietary mismatches: zeroBasket T vs C; difference ≥£0Decision for 2 Nov
London86.68%: pass17.44 vs 18.04; −3.33%: pass1: fail£72.70 vs £72.49; +£0.21: passNo-go
South East86.40%: pass17.54 vs 17.89; −1.93%: pass1: fail£66.64 vs £66.20; +£0.44: passNo-go
Midlands86.35%: pass17.48 vs 17.89; −2.25%: pass1: fail£55.66 vs £55.35; +£0.31: passNo-go
North West86.28%: pass17.58 vs 18.47; −4.80%: pass0 recorded: observed pass£57.51 vs £57.29; +£0.22: passNo-go: shared safety gap
Scotland82.76%: pass over full test230.16 vs 18.20; +65.70%: fail0 recorded: observed pass£52.98 vs £52.62; +£0.36: passNo-go: safety and catalogue

Zero recorded incidents does not establish that an unprotected region is safe. Refunding a customer closes the complaint, not the underlying failure.

The aggregate claims conceal different problems.

Commercial’s acceptance claim is broadly correct: pooled acceptance rose from 71.20% to 85.93%, a 14.74 percentage-point improvement. Pooled refunds increased 7.86%, inside the overall guardrail. But the agreed criterion applies to every region. Scotland fails substantially.

Analytics’ £4 concern is also numerically correct but misleading as a treatment comparison: pooled baskets are £61.09 in treatment versus £65.15 in control. Treatment contains proportionally fewer high-value London orders and more lower-value Midlands and Scotland orders because allocation varies by region. Within every region, treatment baskets are higher across the full test. Applying the same combined regional order mix to both arms gives £63.26 versus £62.96, approximately £0.30 higher in treatment. This resolves the composition distortion; it does not establish a statistically reliable basket uplift.

Scotland is not settling down.

After the supplier switch on 15 October, treatment acceptance fell from 86.34% to 79.08%, below the launch threshold. Refunds rose from 17.25 to 43.25 per 1,000 orders, versus 17.79 in concurrent control, approximately 143% higher.

The final seven days still show 79.17% acceptance and 43.64 refunds per 1,000. The full-period acceptance pass masks a persistent failure in the current catalogue. The open pack-size mapping incident provides a plausible mechanism, and the repair has no date.

Conditions for reconsideration

  • Engineering and QA, by 4 November: deploy and verify the dietary filter across all six protected attributes. Replay the three incident cases, test missing or conflicting attributes and confirm unsafe or unverified candidates cannot be automatically selected. A merged change is insufficient.
  • PM, Analytics and regional Ops, review on 6 November:3 assess a limited post-fix pilot in London, South East, Midlands and North West, retaining randomised controls. Recheck every regional criterion, audit dietary compliance and verify London’s pricing fix. Expand only where the evidence supports it; insufficient evidence means continued limited exposure.1
  • Scotland Ops and catalogue Engineering, owner and repair date agreed on 30 October: complete and validate the re-map before restarting a controlled pilot. Judge readiness on post-repair results, not the earlier healthy fortnight.
  • During any restart: immediately disable auto-substitutions for a dietary mismatch and pause affected regions for guardrail breaches, with daily regional monitoring.

The 9 November booking window is a commercial deadline. It supplies no evidence that either unresolved defect will be fixed in time.

What a PM had to fix

  1. 1Decision deferredMake the callTargeted repair

    What we’d changeGive Commercial something to plan around: say whether the four clean regions can go by 9 November, set a threshold for pausing a region, and give Scotland a Christmas fallback.

  2. 2Contradiction missedSurface the contradictionQuick edit

    What we’d changeShow Scotland failing acceptance since the supplier switch in the table, as the text does, rather than passing.

  3. 3Test or gate too weakTighten the testQuick edit

    What we’d changeA 6 November review would have only a day or two of post-filter data. Allow enough to judge.

  4. 4OtherFix itQuick edit

    What we’d changeSay whether removing them changes any result. The brief asks for it.

Claims the judge couldn’t find in the brief

  • Refunding a customer closes the complaint, not the underlying failure.

Grades and run details

Decision model 77 · LLM judge 12 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly13%
  • passAddresses the actual decision94%
  • passRespects explicit constraints28%
  • passIdentifies material uncertainty85%
  • passAvoids unsupported claims32%
  • passProduces the required deliverable81%
  • passChecks every criterion, region by region58%
  • passFinds the Scotland breach100%
  • passExplains the basket gap98%
  • partialCatches the duplicate rows63%
  • passMakes a clear call77%
  • partialLimits the downside of being wrong63%
  • failGets the base of every number right30%
Run
Run
#1
Time to output
71 s
Submitted
25 Sept 2026

Usable with a quick edit10 of 13 checks passed

A PM could use it with light edits, mainly correcting minor arithmetic and removing the unsupported incident-reporting inference. The core decision, table, and conditions are actionable.

Output

Recommendation: AI auto-substitutions, launch meeting 30 October

Decision: no-go for Monday 2 November. Go with conditions from 4 November for London, South East, Midlands and North West. Scotland stays on picker substitutions until it re-passes.

Criteria by region

Figures are recomputed from the daily rows. Treatment (T) and control (C) are pooled over the 28 days.

Region1. Acceptance (T) ≥80%2. Refunds per 1,000 orders, T vs C (≤ +10%)3. Dietary mismatches (0)4. Basket T vs C (not lower)Verdict
London86.7% ✅17.4 vs 18.0 (−3%) ✅1 (nut allergen, 23 Oct, open) ❌£72.72 vs £72.48 (+£0.24) ✅Conditional
South East86.4% ✅17.5 vs 17.9 (−2%) ✅1 (vegan cheese, 19 Oct) ❌£66.66 vs £66.20 (+£0.46) ✅Conditional
Midlands86.3% ✅17.5 vs 17.9 (−2%) ✅1 (gluten-free, 4 Oct) ❌£55.66 vs £55.35 (+£0.31) ✅Conditional
North West86.3% ✅17.6 vs 18.5 (−5%) ✅0 logged ✅£57.52 vs £57.28 (+£0.24) ✅Conditional
Scotland82.8% overall, 79.1% since 15 Oct ❌30.2 vs 18.2 (+66%); since 15 Oct 43.3 vs 17.8 (+143%) ❌0 logged ✅£52.98 vs £52.62 (+£0.36) ✅No-go

Baskets are means of daily averages. The differences are small and I have not tested them for significance, but the point estimate is positive everywhere.

Checking what people have said

  • Head of Commercial. The pooled numbers are right: acceptance is 85.9% against 71.2% (+14.7 points), and refunds are 19.5 against 18.1 per 1,000 (+8%). But the pooled refund figure hides Scotland, which is failing badly. Four regions are 2–5% better than control on refunds.
  • Head of Analytics. The £4 gap is real in the pooled data (£61.10 vs £65.15) but it is a mix effect, not a treatment effect. London, with the highest baskets (about £72), was 42% of control orders but only 18% of treatment orders. The Midlands and Scotland (about £53–56) were 18% of control but 42% of treatment. Within every region, treatment baskets are £0.24–£0.46 higher. Criterion 4 passes. The uneven split did not bias the regional comparisons.
  • Ops director (Scotland). The data does not support "settling down". Before the supplier switch on 15 October, Scotland treatment was healthy (86.3% acceptance, 17.3 refunds per 1,000). From 15 October to 28 October, acceptance has been flat at 78–81%, and daily refunds have been 25–33 against a pre-switch level of 10–14. The cause is the open incident: the new supplier's pack sizes are mapped to the wrong product family, so the model makes wrong-size dairy substitutions. That is a catalogue defect with no fix date, not noise. Scotland control is unaffected.

Why the date moves

Criterion 3 fails in three regions. The three mismatches include a nut-allergy case and a coeliac case. The hard filter that prevents them deploys on 4 November. Until it is live, every region has the same exposure. North West and Scotland show zero only because none were reported, and the incidents come from customer complaints, so the true rate is unknown. I would not switch on a known allergen risk two days before the fix.

Waiting is cheap. The booking window opens on 9 November, not 2 November. Deploying the filter on 4 November and switching regions on 4–5 November leaves several days before the window. Only Scotland has no route to launch before then.

Conditions for the four regions

  1. The filter deploys on 4 November and passes regression tests. Before any region goes live, it is tested in production against the three incident SKUs (gluten-free loaf, vegan cheese, nut-free bar) and the other attributes: vegetarian, halal and kosher. The London incident is still open with the filter "in review", so confirm the filter would have blocked it.
  2. Regions are switched on region by region, North West and South East first, then London and the Midlands. This gives a day of live data before the larger volumes go on.
  3. Any dietary or allergen mismatch after launch switches that region back to picker substitution immediately, pending review. That includes any mismatch in the first week.
  4. Daily monitoring of acceptance (80% floor) and refunds (10% guardrail) per region against a small held-back control. Delegate the switch-on decision so it does not need another meeting.

Scotland

Keep Scotland on picker substitutions. Retest only after the catalogue re-map ships, on a small treatment share. It should show at least three consecutive days at 80% or better acceptance, and refunds within 10% of control, before it goes to full volume. Engineering has no date, so Scotland should not be in the Christmas plan and should be treated as upside.

Data notes

  • North West treatment has three duplicated rows (8, 9 and 10 October), matching the open export-job incident. I removed them. This restores the intended 50% split, and the North West rates barely move either way.
  • Incident logs are a floor on dietary failures, not a measure of them.
  • This is 28 days of data in a low-volume region (Scotland control is about 265 orders a day). Even so, the Scotland gap is far outside noise.

Grades and run details

Decision model 81 · LLM judge 12 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly9%
  • passAddresses the actual decision98%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty92%
  • partialAvoids unsupported claims36%
  • passProduces the required deliverable79%
  • passChecks every criterion, region by region56%
  • passFinds the Scotland breach100%
  • passExplains the basket gap99%
  • passCatches the duplicate rows100%
  • passMakes a clear call85%
  • passLimits the downside of being wrong76%
  • failGets the base of every number right29%
Run
Run
#1
API response time
2.3 min
Submitted
29 Sept 2026

Needs more than a quick edit9 of 13 checks passed

The refund figures are inaccurate, so a PM would need to re-check the data and recalculate before using the recommendation. The output is not usable as-is.

Output

Recommendation: No-go for the all-region rollout on 2 November

Do not enable auto-substitutions anywhere on 2 November: the dietary hard filter is not scheduled to deploy until 4 November, so the zero-mismatch safety criterion is not yet protected. After deployment and validation, North West is the only region I’d consider enabling first, with conditions. Hold the other regions for the reasons below.

Figures are for 1–28 October; refund rates are per 1,000 orders. Basket averages are order-weighted within each region. I excluded the three exact duplicate North West treatment rows on 8, 9 and 10 October.

RegionTreatment acceptance (≥80%)Refunds / 1,000 orders: treatment vs control (treatment ≤10% higher)Dietary/allergen mismatches (must be zero)Average basket: treatment vs control (must not be lower)Decision
------:---:------:---
London~86% Pass~17.9 vs ~18.4; Pass1; Fail — nut-containing bar substituted for nut-free item~£72.8 vs ~£72.5; PassNo-go
South East~85% Pass~17.8 vs ~17.9; Pass1; Fail — dairy cheese substituted for vegan cheese~£66.6 vs ~£66.3; PassNo-go
Midlands~86% Pass~17.6 vs ~18.0; Pass1; Fail — standard loaf substituted for gluten-free loaf~£55.6 vs ~£55.4; PassNo-go
North West~86% Pass~17.8 vs ~18.8; Pass0 observed; Pass in test~£57.5 vs ~£57.3; PassGo with conditions, after filter validation
Scotland~79% Fail1~29.8 vs ~18.4; Fail — about 62% higher0 observed; Pass in test~£53.0 vs ~£52.6; PassNo-go

Conditions and rationale

  • North West: After the hard filter deploys, pass regression tests across all six protected attributes (gluten-free, vegan, vegetarian, nut-free, halal and kosher); then enable regionally with close monitoring and a rollback switch. The export duplicates should also be fixed or guarded against before relying on daily reporting.
  • London, South East and Midlands: Do not launch until the filter is deployed and the dietary-mismatch failures are investigated and validated as fixed. Their good acceptance, refund and basket figures do not offset a safety-criterion failure.
  • Scotland: Hold until the supplier catalogue re-map is complete and verified, and a fresh test meets both acceptance and refund guardrails. The sharp rise in treatment refunds and lower acceptance coincide with the dairy supplier/catalogue change; the issue is open, with no completion date.

The pooled “£4 smaller” basket claim should not drive this decision: treatment exposure differed by region, and baskets vary substantially by region. Within each region, the treatment average is slightly higher in these data. Likewise, an overall acceptance or refund result cannot override region-level launch gates. The 9 November booking window is important, but it does not justify launching before the safety control is live.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    What we’d change82.76%: Pass over 1–28 October. However, acceptance fell to 79.08% over 15–28 October after the supplier change, below the launch threshold.

Claims the judge couldn’t find in the brief

  • London treatment refunds per 1,000 orders is ~17.9.
  • London control refunds per 1,000 orders is ~18.4.
  • South East treatment refunds per 1,000 orders is ~17.8.
  • Midlands treatment refunds per 1,000 orders is ~17.6.
  • North West treatment refunds per 1,000 orders is ~17.8 (after excluding duplicates).
  • Scotland control refunds per 1,000 orders is ~18.4.

Grades and run details

Decision model 69 · LLM judge 9 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly8%
  • passAddresses the actual decision47%
  • passRespects explicit constraints61%
  • partialIdentifies material uncertainty55%
  • passAvoids unsupported claims21%
  • passProduces the required deliverable63%
  • passChecks every criterion, region by region28%
  • passFinds the Scotland breach88%
  • passExplains the basket gap86%
  • partialCatches the duplicate rows69%
  • partialMakes a clear call62%
  • partialLimits the downside of being wrong96%
  • failGets the base of every number right45%
Run
Run
#1
API response time
44 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyMixedMixedWrong
GPT-6 Astra · ChatGPT

All factual statements about the current situation are supported by the brief or correct arithmetic; no invented facts.

Sonnet 5.5 · API

The output's factual claims are drawn from the brief and supplied CSVs, with only minor rounding and one unsupported inference about incident reporting.

GPT-6 Luna · API

Several refund-per-1000 figures are incorrect (e.g., London treatment 17.9 vs actual 17.4, London control 18.4 vs 18.0), so the output does not use the supplied evidence correctly.

Avoids unsupported claimsMixedWrongRight
GPT-6 Astra · ChatGPT

Presents 'Refunding a customer closes the complaint, not the underlying failure' as fact without support from the supplied context.

Sonnet 5.5 · API

It asserts that zero logged dietary incidents in North West and Scotland are only because none were reported and that incidents come from customer complaints, which is not established by the supplied evidence.

GPT-6 Luna · API

Causal claims like the supplier change coinciding with refunds are presented as observations, not unsupported facts.

Checks every criterion, region by regionRightRightMixed
GPT-6 Astra · ChatGPT

Table checks every criterion for each region using figures computed from the attached files.

Sonnet 5.5 · API

It checks all four launch criteria for each region with figures computed from the files, including the duplicate correction.

GPT-6 Luna · API

The table checks each criterion per region but the refund figures are incorrect, so it does not present correct figures from the files.

Catches the duplicate rowsWrongRightRight
GPT-6 Astra · ChatGPT

Finds and removes the duplicate North West rows but does not explicitly state whether they change the overall assessment.

Sonnet 5.5 · API

It finds the duplicated North West treatment rows, removes them, and says the result barely changes.

GPT-6 Luna · API

Catches the duplicate North West rows, excludes them, and recommends fixing the export job.

Limits the downside of being wrongRightRightWrong
GPT-6 Astra · ChatGPT

Proposes daily regional monitoring, a trigger (dietary mismatch or guardrail breach) and action (disable, pause).

Sonnet 5.5 · API

It limits downside with staged switch-on, daily monitoring, immediate rollback on any dietary mismatch, and Scotland held out.

GPT-6 Luna · API

Does not name a specific post-launch signal, threshold, or trigger for rollback; only says 'close monitoring and a rollback switch' without specifics.

Gets the base of every number rightMixedWrongWrong
GPT-6 Astra · ChatGPT

All derived percentages and rates use correct numerators and denominators, and basket values are weighted by orders.

Sonnet 5.5 · API

Some derived figures are off, including the pooled basket gap and the control-order share for London, and the basket base is stated as means of daily averages rather than order-weighted averages.

GPT-6 Luna · API

Multiple refund-per-1000 figures are miscalculated (e.g., London treatment 17.9 vs 17.4, London control 18.4 vs 18.0), so the base numbers are wrong.

All got right 7

Addresses the actual decisionRightRightRight
GPT-6 Astra · ChatGPT

Commits to a clear no-go on 2 Nov, then conditional rollout by region, and states what would change the call.

Sonnet 5.5 · API

It commits early to no-go on 2 November, conditional go from 4 November for four regions, hold Scotland, and states what would change the call.

GPT-6 Luna · API

Commits to a clear no-go for all-region on 2 Nov, then conditional go for North West after filter, with conditions and what would change it.

Respects explicit constraintsRightRightRight
GPT-6 Astra · ChatGPT

Respects the requested form, reader, word limit, and uses the data rather than stakeholder summaries.

Sonnet 5.5 · API

It provides a region-by-region recommendation with a launch-criteria table and keeps prose under 900 words.

GPT-6 Luna · API

Includes the required region-by-region table, prose under 900 words, and addresses the launch meeting.

Identifies material uncertaintyRightRightRight
GPT-6 Astra · ChatGPT

Names the unknowns (filter delivery, Scotland catalogue repair, post-fix evidence) and says what would change the call.

Sonnet 5.5 · API

It names material unknowns such as filter deployment, Scotland re-map timing, duplicate handling, and incident under-reporting, and gives monitoring/rollback conditions.

GPT-6 Luna · API

Names the filter deployment, duplicate fix, and Scotland catalogue re-map as unknowns that could change the decision.

Produces the required deliverableRightRightRight
GPT-6 Astra · ChatGPT

A complete recommendation with a criteria table, under 900 words, usable for the launch meeting.

Sonnet 5.5 · API

The requested recommendation and table are present, usable for the launch meeting, and within the length limit.

GPT-6 Luna · API

Provides a complete recommendation with table, region-by-region, usable for the meeting.

Finds the Scotland breachRightRightRight
GPT-6 Astra · ChatGPT

Identifies the Scotland refund breach, ties it to the 15 October supplier change, notes acceptance dip, and blocks Scotland.

Sonnet 5.5 · API

It identifies Scotland's refund breach, ties it to the 15 October supplier change, notes the acceptance dip, and holds Scotland.

GPT-6 Luna · API

Identifies Scotland's refund breach, ties it to the supplier change, notes acceptance dip, and holds Scotland.

Explains the basket gapRightRightRight
GPT-6 Astra · ChatGPT

Shows treatment baskets are higher within every region and explains the pooled gap as a mix artefact.

Sonnet 5.5 · API

It compares baskets within each region and explains the pooled gap as an artefact of the uneven regional split.

GPT-6 Luna · API

Explains the pooled basket gap as a mix effect, noting treatment baskets are equal or higher within each region.

Makes a clear callRightRightRight
GPT-6 Astra · ChatGPT

One clear call per region, with conditions and dates, and says what would change it.

Sonnet 5.5 · API

It gives one clear call per region with dates, conditions, and triggers for change.

GPT-6 Luna · API

Makes a clear call per region with dates and conditions, and says what would change it.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.388.32None
2GPT-6 AstrawithChatGPT86.088.32None
3Sonnet 5.5withAPI85.488.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI82.177.62None
6Gemini 3.8 FlashwithAPI71.931.521 capped
7Gemini 3.5 Flash-LitewithGemini50.642.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook