Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.0 · 2 casesDifficulty

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model, LLM judge and blind PM review

Results

Every setup we’ve tested on this task, across all cases and repeats.

#Model · HarnessTask scoreDecision modelLLM judgePM reviewRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.

The brief

The new checkout for our grocery app is due to go to 100% of customers on Thursday. You're the PM. Using the material below, write your go/no-go recommendation for the release meeting. Keep it under 400 words.

ScenarioBasketful is an online grocery service with 220,000 orders a week. The new checkout replaces a four-step flow with a single page and adds Apple Pay and Google Pay.
Rollout so farLive for 10% of customers for 14 days (about 31,000 orders).
Results against the old checkoutCheckout completion 71.4% vs 68.9% (+2.5pp, 95% CI +1.9 to +3.1). Average order value £58.20 vs £58.90 (difference −£0.70, 95% CI −£1.60 to +£0.20). Payment failures 1.9% vs 1.2%. Most of the extra failures come from one card issuer, Northbank, whose cards decline the new flow's security check; Northbank cards are used by 9% of customers.
SupportCheckout complaints: 0.4% of orders vs 0.3% on the old checkout.
EngineeringA fix for the Northbank declines is written and in review; it could ship next Tuesday. The rollout can be targeted by card issuer. The old checkout is scheduled for deletion in the release after this one.
CommercialThursday is the start of the Christmas delivery-slot booking window; marketing emails go to 1.2 million customers that morning.
What a strong answer does

Go with conditions: take the new checkout to everyone except Northbank cardholders, who stay on the old flow until Tuesday's fix is verified; keep the old checkout until then; treat the order-value difference as noise, not a blocker. Weighs the Christmas-window traffic both ways and says what would change the call (the fix slipping, failure rates rising at scale).

Critical failures (cap the score)
  • Recommends a full rollout that includes Northbank cardholders before the fix
Case

v1.1 · synthetic · e-commerce, launch, peak season