Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.0 · 2 casesDifficulty

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model, LLM judge and blind PM review

Results

Every setup we’ve tested on this task, across all cases and repeats.

#Model · HarnessTask scoreDecision modelLLM judgePM reviewRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.

The brief

Tomorrow's launch meeting decides whether AI-drafted support replies go live to all 42 support agents on Monday. You're the PM. Using the material below, write your launch recommendation for the meeting: go, no-go, or go with conditions, and why. Keep it under 400 words.

ScenarioLedgerly is accounting software for small businesses. The feature drafts a first reply to each support ticket; an agent reviews and edits every draft before it is sent.
Launch criteria agreed in the PRD1. At least 60% of drafts accepted as-is or with minor edits. 2. Wrong or harmful drafts (wrong account data, wrong policy, invented features) at most 1% on the 500-ticket eval set. 3. No draft may commit to a refund. 4. Median time to first response down at least 30% in the pilot.
Pilot (8 agents, 3 weeks, 2,940 tickets)Accepted as-is or with minor edits: 64%. Median time to first response: 7.1 hours before, 4.2 hours during (−41%). Seven of the eight agents want to keep it.
Eval set (500 tickets)Wrong or harmful drafts: 9 of 500 (1.8%). Six of the nine are billing tickets quoting the old prices, which changed on 1 September; the model's reference data stops in August. Two drafts said “we'll refund this charge”; the pilot agents caught and removed both.
Open bugsReplies to Welsh-language tickets come back in English (about 0.3% of tickets). Tables in drafts are badly formatted.
Support operationsTraining for all 42 agents is booked for Friday. Billing is one of four queues and handles 38% of tickets.
What a strong answer does

Not a full go: criterion 2 fails (1.8% against 1%) and criterion 3 is breached even though agents caught it. Strong answer: go with conditions, launching to the non-billing queues while billing waits for updated pricing data and a hard block on refund language, then re-running the eval to confirm it meets the bar; names exactly what would turn it into a full go.

Critical failures (cap the score)
  • Recommends a full launch while the error rate misses the agreed 1% bar
  • Treats the refund commitments as acceptable
Case

v1.1 · synthetic · AI product, launch