Make the launch call
Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?
The PM job
Deciding whether a feature ships on the planned date, and on what conditions.
Why it matters
Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.
What good looks like
- Checks each agreed criterion against the evidence
- Makes one clear call, with conditions if needed
- Separates risks that block launch from risks that can be managed
- Says what would change the call
Deliberately not measured
- Rollout engineering detail
- Project-plan formatting
Deciding against agreed launch criteria
Rubber-stamps a launch that misses an agreed bar, or blocks it on noise
Decision model, LLM judge and blind PM review
Results
Every setup we’ve tested on this task, across all cases and repeats.
| # | Model · Harness | Task score | Decision model | LLM judge | PM review | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.
The new checkout for our grocery app is due to go to 100% of customers on Thursday. You're the PM. Using the material below, write your go/no-go recommendation for the release meeting. Keep it under 400 words.
Go with conditions: take the new checkout to everyone except Northbank cardholders, who stay on the old flow until Tuesday's fix is verified; keep the old checkout until then; treat the order-value difference as noise, not a blocker. Weighs the Christmas-window traffic both ways and says what would change the call (the fix slipping, failure rates rising at scale).
- Recommends a full rollout that includes Northbank cardholders before the fix
v1.1 · synthetic · e-commerce, launch, peak season