Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Respects explicit constraints98% pass
    The output is a launch recommendation, addresses the meeting, and stays within the 400-word limit (345 words).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong35% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows54% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Gets the base of every number right67% pass
    Some derived figures are off, including the pooled basket gap and the control-order share for London, and the basket base is stated as means of daily averages rather than order-weighted averages.
    Sonnet 5.5 · API · Go/no-go for AI substitutions, from the rollout data

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Tomorrow's launch meeting decides whether AI-drafted support replies go live to all 42 support agents on Monday. You're the PM. Using the material below, write your launch recommendation for the meeting: go, no-go, or go with conditions, and why. Keep it under 400 words.

ScenarioLedgerly is accounting software for small businesses. The feature drafts a first reply to each support ticket; an agent reviews and edits every draft before it is sent.
Launch criteria agreed in the PRD1. At least 60% of drafts accepted as-is or with minor edits. 2. Wrong or harmful drafts (wrong account data, wrong policy, invented features) at most 1% on the 500-ticket eval set. 3. No draft may commit to a refund. 4. Median time to first response down at least 30% in the pilot.
Pilot (8 agents, 3 weeks, 2,940 tickets)Accepted as-is or with minor edits: 64%. Median time to first response: 7.1 hours before, 4.2 hours during (−41%). Seven of the eight agents want to keep it.
Eval set (500 tickets)Wrong or harmful drafts: 9 of 500 (1.8%). Six of the nine are billing tickets quoting the old prices, which changed on 1 September; the model's reference data stops in August. Two drafts said “we'll refund this charge”; the pilot agents caught and removed both.
Open bugsReplies to Welsh-language tickets come back in English (about 0.3% of tickets). Tables in drafts are badly formatted.
Support operationsTraining for all 42 agents is booked for Friday. Billing is one of four queues and handles 38% of tickets.
What a strong answer does

Not a full go: criterion 2 fails (1.8% against 1%) and criterion 3 is breached even though agents caught it. Strong answer: go with conditions, launching to the non-billing queues while billing waits for updated pricing data and a hard block on refund language, then re-running the eval to confirm it meets the bar; names exactly what would turn it into a full go.

Critical failures (cap the score)
  • Recommends a full launch while the error rate misses the agreed 1% bar
  • Treats the refund commitments as acceptable
Case

v1.7 · synthetic · AI product, launch

Usable with a quick edit7 of 10 checks passed

The unsupported claim about pilot engagement is minor; remove it and the recommendation remains sound and usable with a quick edit.

Output

Launch recommendation: AI-drafted support replies

Recommendation: no-go for all 42 agents this Monday. Go the following Monday (5 Oct) if three conditions are met.

Scorecard against PRD criteria

CriterionTargetResultStatus
Accepted as-is / minor edits≥60%64%✅ Pass
Wrong or harmful drafts≤1%1.8% (9/500)❌ Fail
No refund commitments02 drafts❌ Fail
Median first response−30%−41% (7.1h → 4.2h)✅ Pass

Two of four criteria fail. The value is proven, but the safety bar isn't met yet.

Why not Monday

  • Refunds are a hard line, and we crossed it. Two drafts promised refunds. Agents caught both, but the criterion says no draft may do this. Human review is the backstop, not the control. The pilot used 8 motivated agents. With 42 agents, some of them tired and under queue pressure, a promise will eventually get through.2
  • Billing errors land on our biggest queue. Six of the nine failures quote pre-1 September prices. Billing is 38% of volume, so this is a systematic error, not an edge case.
  • Training comes after launch. Training is booked for next Friday, so a Monday launch puts the tool in front of untrained agents.1 The pilot agents were hands-on and engaged, and we can't assume the rest will review as carefully.

Conditions for a 5 October go

  1. Refund guardrail. Add a hard output check that blocks or flags any draft committing to a refund. Test it against the eval set and adversarial cases, with zero passes allowed.
  2. Refresh pricing data and re-run the full 500-ticket eval. Launch only if wrong or harmful drafts are at most 1%. Removing the six pricing errors would leave 0.6% on paper. We should prove that with a re-run, not assume it.
  3. Train before rollout. Keep Friday's training and include the refund and pricing failure modes explicitly, so agents know what to look for.

Not blocking

  • Welsh replies coming back in English (0.3%). Turn drafting off for Welsh-language tickets until this is fixed. It costs little and avoids a visible failure.
  • Table formatting. This is cosmetic, and agents edit every draft anyway. Fix it in the normal backlog.

Meanwhile

Keep the 8 pilot agents on the tool so we don't lose momentum.3 In the first two weeks after launch, spot-check a sample of sent billing replies.

The ask

Agree to the three conditions. Name owners for the guardrail and the eval re-run, with results due Thursday 1 October. If both pass, we confirm go on 2 October after training.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    Source: Support operationsTraining for all 42 agents is booked for Friday.

    What we’d changeThe brief says training is on Friday, not 'next Friday'. Don't turn an ambiguous date into a blocker, or build an October schedule around it.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeSay the risk grows with 42 agents rather than stating it as certain, and don't assume the pilot agents were unusually motivated.

  3. 3Constraint missedRestore the constraintQuick edit

    What we’d changeMomentum isn't a reason to keep a tool that fails two criteria in use. Add the refund check first, or pause it.

Check by check

Got wrong · 3
  • Uses the supplied evidence correctlyThe output states that pilot agents were hands-on, engaged and motivated, which is not supported by the supplied context (only that 7/8 wanted to keep it).
  • Avoids unsupported claimsPresents pilot agents' hands-on engagement and motivation as established fact when the evidence only shows retention preference, not engagement level.
  • Limits the downside of being wrongPost-launch spot-checks are mentioned but no threshold or rollback action is specified, and irreversible risks are not flagged.
Got right · 7
  • Addresses the actual decisionCommits to no-go for Monday, with a conditional go the following Monday, and states what would change the call.
  • Respects explicit constraintsRespects the length limit, addresses the meeting, and proposes conditions that would enforce the launch criteria.
  • Identifies material uncertaintyNames the remaining work (guardrail, pricing refresh, re-run) and says the call depends on their results.
  • Produces the required deliverableDelivers a complete launch recommendation under 400 words that the meeting could act on.
  • Checks every agreed criterionEvaluates all four PRD criteria with a status table and references to the evidence.
  • Makes a clear callClear no-go for Monday, go on 5 October if three named conditions are met.
  • Gets the base of every number rightAll percentages and differences are computed correctly from the supplied numbers.

Claims the judge couldn’t find in the brief

  • Pilot agents were hands-on, engaged and motivated

Grades and run details

Decision model 80 · LLM judge 8 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly16%
  • passAddresses the actual decision95%
  • passRespects explicit constraints38%
  • passIdentifies material uncertainty98%
  • partialAvoids unsupported claims29%
  • passProduces the required deliverable64%
  • passChecks every agreed criterion99%
  • passMakes a clear call100%
  • partialLimits the downside of being wrong84%
  • passGets the base of every number right85%
Run
Run
#1
Time to output
23 s
Submitted
25 Sept 2026

Usable with a quick edit9 of 10 checks passed

The recommendation is complete, evidence-based, and actionable; the only meaningful gap is a missing post-launch pause threshold, which is a quick edit.

Output

Recommendation: Go with conditions

Not a clean go. Two of the four agreed launch criteria are currently failed, and I don't think we should waive criteria we set ourselves.

CriterionResultStatus
≥60% accepted64%Met
≤1% wrong/harmful1.8% (9/500)Failed
No refund commitments2 drafts promised refundsFailed
≥30% faster first response−41% (7.1h → 4.2h)Met

The value is real: faster responses, good acceptance, and 7 of 8 pilot agents want to keep it. Both failures also have identifiable causes.

  • Wrong drafts: Six of the nine errors come from stale pricing, because the reference data stops in August. Fixing that should bring us to about 3/500 (0.6%), but that is a projection until we re-run the eval.
  • Refunds: Agents caught both drafts, but the criterion says no draft, and with 42 agents we shouldn't rely on catching them by habit.

Conditions for Monday

  1. Refresh the reference data with post-1 September pricing and re-run the full 500-ticket eval. It must come in at 1% or below (5 or fewer wrong).
  2. Add a refund guardrail that blocks or flags commitment language such as "we'll refund." It must show zero refund commitments on the re-run plus a targeted set of refund-request tickets.
  3. Kill switch and flag button. Ops can disable the feature instantly, and agents can flag bad drafts in one click.

Fallback

  • If condition 2 fails, no-go for Monday.
  • If only condition 1 fails, launch on the three non-billing queues (62% of tickets) and keep billing off until it passes.
  • Friday training goes ahead for all 42 agents either way. It should cover the refund rule and checking prices before sending.

Known issues, not blockers

  • Welsh tickets (~0.3%) get English replies. Agents will write these manually, and we'll suppress drafts for them if detection is feasible.
  • Table formatting is cosmetic, since agents edit every draft. We'll fix it after launch.

After launch

  • Sample wrong-draft rates weekly for the first month.
  • Track how often agents edit or discard drafts, as a check on whether review discipline slips.
  • Ask the one pilot agent who didn't want to keep it why.

I'd like the eval re-run results by Thursday so we can make the final call before Friday's training.

Check by check

Got wrong · 1
  • Limits the downside of being wrongIt describes post-launch monitoring but does not name a threshold or trigger after launch to pause/roll back, nor flag anything irreversible.
Got right · 9
  • Uses the supplied evidence correctlyAll current-state facts used are from the supplied context; no invented systems or numbers.
  • Addresses the actual decisionIt commits to 'Go with conditions' and says what would change the call: no-go if condition 2 fails, non-billing launch if only condition 1 fails.
  • Respects explicit constraintsUnder 400 words, and the proposed eval re-run, 1%/zero refund gates, and non-billing fallback would enforce the agreed criteria.
  • Identifies material uncertaintyIt identifies the re-run pricing/refund results as the open unknowns and maps them to specific decision changes.
  • Avoids unsupported claimsThe 0.6% estimate is clearly labelled a projection, and causes match the supplied context.
  • Produces the required deliverableA complete launch recommendation with criteria table, conditions, fallback and post-launch monitoring, usable as-is.
  • Checks every agreed criterionThe table explicitly checks all four agreed criteria.
  • Makes a clear callClear 'Go with conditions' call with named fallbacks.
  • Gets the base of every number rightAll derived rates are correct: 9/500=1.8%, 9−6=3/500=0.6%, 38% billing implies 62% non-billing, and 2.9/7.1=−41%.

Grades and run details

Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly43%
  • passAddresses the actual decision98%
  • passRespects explicit constraints39%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims69%
  • passProduces the required deliverable86%
  • passChecks every agreed criterion100%
  • passMakes a clear call100%
  • partialLimits the downside of being wrong84%
  • passGets the base of every number right76%
Run
Run
#1
API response time
17 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 86% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI93.788.32None
2GPT-6 AstrawithChatGPT89.888.32None
3GPT-6.1 SolwithAPI87.388.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI86.077.62None
6Gemini 3.5 Flash-LitewithGemini53.142.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook