Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Addresses the actual decision95% pass
    The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong32% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows50% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Uses the supplied evidence correctly54% pass
    The output states that pilot agents were hands-on, engaged and motivated, which is not supported by the supplied context (only that 7/8 wanted to keep it).
    Opus 5.5 · Claude · Go/no-go for AI-drafted support replies

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Tomorrow's launch meeting decides whether AI-drafted support replies go live to all 42 support agents on Monday. You're the PM. Using the material below, write your launch recommendation for the meeting: go, no-go, or go with conditions, and why. Keep it under 400 words.

What the model was given6 items: Scenario, Launch criteria agreed in the PRD, Pilot (8 agents, 3 weeks, 2,940 tickets), Eval set (500 tickets), Open bugs, Support operations
ScenarioLedgerly is accounting software for small businesses. The feature drafts a first reply to each support ticket; an agent reviews and edits every draft before it is sent.
Launch criteria agreed in the PRD1. At least 60% of drafts accepted as-is or with minor edits. 2. Wrong or harmful drafts (wrong account data, wrong policy, invented features) at most 1% on the 500-ticket eval set. 3. No draft may commit to a refund. 4. Median time to first response down at least 30% in the pilot.
Pilot (8 agents, 3 weeks, 2,940 tickets)Accepted as-is or with minor edits: 64%. Median time to first response: 7.1 hours before, 4.2 hours during (−41%). Seven of the eight agents want to keep it.
Eval set (500 tickets)Wrong or harmful drafts: 9 of 500 (1.8%). Six of the nine are billing tickets quoting the old prices, which changed on 1 September; the model's reference data stops in August. Two drafts said “we'll refund this charge”; the pilot agents caught and removed both.
Open bugsReplies to Welsh-language tickets come back in English (about 0.3% of tickets). Tables in drafts are badly formatted.
Support operationsTraining for all 42 agents is booked for Friday. Billing is one of four queues and handles 38% of tickets.
What a strong answer doesThe answer key the graders mark against

Not a full go: criterion 2 fails (1.8% against 1%) and criterion 3 is breached even though agents caught it. Strong answer: go with conditions, launching to the non-billing queues while billing waits for updated pricing data and a hard block on refund language, then re-running the eval to confirm it meets the bar; names exactly what would turn it into a full go.

Critical failures (cap the score at 40)
  • Recommends a full launch while the error rate misses the agreed 1% bar
  • Treats the refund commitments as acceptable
Case

Staff-level · v1.9 · synthetic · AI product, launch

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 10 checks passed

The recommendation is complete, evidence-based, and actionable; the only meaningful gap is a missing post-launch pause threshold, which is a quick edit.

Output

Recommendation: Go with conditions

Not a clean go. Two of the four agreed launch criteria are currently failed, and I don't think we should waive criteria we set ourselves.

CriterionResultStatus
≥60% accepted64%Met
≤1% wrong/harmful1.8% (9/500)Failed
No refund commitments2 drafts promised refundsFailed
≥30% faster first response−41% (7.1h → 4.2h)Met

The value is real: faster responses, good acceptance, and 7 of 8 pilot agents want to keep it. Both failures also have identifiable causes.

  • Wrong drafts: Six of the nine errors come from stale pricing, because the reference data stops in August. Fixing that should bring us to about 3/500 (0.6%), but that is a projection until we re-run the eval.
  • Refunds: Agents caught both drafts, but the criterion says no draft, and with 42 agents we shouldn't rely on catching them by habit.

Conditions for Monday

  1. Refresh the reference data with post-1 September pricing and re-run the full 500-ticket eval. It must come in at 1% or below (5 or fewer wrong).
  2. Add a refund guardrail that blocks or flags commitment language such as "we'll refund." It must show zero refund commitments on the re-run plus a targeted set of refund-request tickets.
  3. Kill switch and flag button. Ops can disable the feature instantly, and agents can flag bad drafts in one click.

Fallback

  • If condition 2 fails, no-go for Monday.
  • If only condition 1 fails, launch on the three non-billing queues (62% of tickets) and keep billing off until it passes.
  • Friday training goes ahead for all 42 agents either way. It should cover the refund rule and checking prices before sending.

Known issues, not blockers

  • Welsh tickets (~0.3%) get English replies. Agents will write these manually, and we'll suppress drafts for them if detection is feasible.
  • Table formatting is cosmetic, since agents edit every draft. We'll fix it after launch.

After launch

  • Sample wrong-draft rates weekly for the first month.
  • Track how often agents edit or discard drafts, as a check on whether review discipline slips.
  • Ask the one pilot agent who didn't want to keep it why.

I'd like the eval re-run results by Thursday so we can make the final call before Friday's training.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly44%
  • passAddresses the actual decision98%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable84%
  • passChecks every agreed criterion100%
  • passMakes a clear call100%
  • partialLimits the downside of being wrong85%
  • passGets the base of every number right79%
Run
Run
#1
API response time
17 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 10 checks passed

The call and evidence are sound; a PM would only need to add specific post-launch monitoring thresholds and a rollback trigger.

Output

Recommendation: Go with conditions; otherwise no-go for Monday.

The pilot met two launch criteria: 64% of drafts were accepted as-is or with minor edits, above the 60% target, and median time to first response fell 41%, exceeding the 30% target. Seven of eight agents also want to keep the feature.

But the launch does not currently meet the safety criteria. The eval found harmful drafts in 1.8% of cases, above the 1% ceiling. Six errors came from outdated billing prices—a significant exposure because billing handles 38% of tickets. And two drafts promised refunds, directly violating the “no draft may commit to a refund” criterion, even though agents caught them.

Before enabling drafts broadly, require: - Updated billing reference data and a rerun of the 500-ticket eval with no more than five wrong or harmful drafts. - A reliable safeguard against refund commitments, verified to produce zero such drafts. - Welsh-language tickets routed away from AI drafting until that issue is fixed.

If those gates pass before Monday, proceed with the launch and monitor closely. If they do not, delay the launch rather than rely on agents to catch known failure modes. The formatting issue should also be addressed or clearly covered in agent training, but it is not the main launch blocker.

Grades and run details

Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly87%
  • passAddresses the actual decision96%
  • passRespects explicit constraints70%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims60%
  • passProduces the required deliverable94%
  • passChecks every agreed criterion75%
  • passMakes a clear call95%
  • partialLimits the downside of being wrong49%
  • passGets the base of every number right86%
Run
Run
#1
API response time
11 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 10 checks passed

The reasoning is sound, evidence is correctly used, and the recommendation is clear. A PM could add a monitoring threshold in a quick edit without redoing any analysis.

Output

Recommendation: no-go for Monday’s rollout to all 42 agents. Reassess once the failed safety gates have been fixed and retested.

The pilot demonstrates value: 64% of drafts were accepted as-is or with minor edits, exceeding the 60% target. Median first-response time fell 41%, exceeding the 30% target, although the before-and-after comparison does not establish causality. Seven of eight agents want to keep it.

But we failed two explicit launch criteria:

  • Wrong or harmful drafts: 1.8% against a 1% ceiling. Nine failures is above the maximum five permitted on the 500-ticket set.
  • No refund commitments: failed. Two drafts promised refunds. Agents catching them prevented those messages reaching customers; it does not satisfy a criterion about what drafts may say. Human review was already part of the agreed design.

Six errors used outdated billing prices. That gives us a concrete fix, but billing represents 38% of tickets, so this is a substantial exposure. Updating the reference data alone would not establish that the remaining errors or refund promises are resolved.

Required before reconsidering launch:

  1. Engineering and the billing policy owner: update and verify current prices and policies, with a named owner and process for keeping reference data current. Investigate the other three harmful drafts.
  2. Engineering: add and test a control that prevents refund commitments from appearing in agent-facing drafts, including paraphrases.
  3. PM and evaluation owner: rerun the full 500-ticket evaluation after the fixes, meeting the original thresholds, and add targeted pricing and refund tests. Report those additional results separately.

Excluding billing is a possible narrower release, not an automatic pass: we lack queue-level results and do not know whether refund failures are confined to billing. It would require explicit scope approval and validation of both routing and the remaining queues.

Welsh-language tickets should bypass drafting until language handling is fixed. Disable tables or use plain text pending the formatting fix.

Keep Friday’s training as preparation, with launch clearly marked as pending. The booking should not determine readiness. After the gates pass, expand in stages with monitoring and a named owner able to disable drafting.

Grades and run details

Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly69%
  • passAddresses the actual decision94%
  • passRespects explicit constraints44%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims74%
  • passProduces the required deliverable86%
  • passChecks every agreed criterion49%
  • passMakes a clear call93%
  • partialLimits the downside of being wrong54%
  • passGets the base of every number right83%
Run
Run
#1
Time to output
34 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

All got wrong 1

Limits the downside of being wrongWrongWrongWrong
Sonnet 5.5 · API

It describes post-launch monitoring but does not name a threshold or trigger after launch to pause/roll back, nor flag anything irreversible.

GPT-6 Luna · API

It says monitor closely but does not give a specific post-launch signal, threshold, pause/rollback action, or flag irreversible exposures.

GPT-6 Astra · ChatGPT

No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.

All got right 9

Uses the supplied evidence correctlyRightRightRight
Sonnet 5.5 · API

All current-state facts used are from the supplied context; no invented systems or numbers.

GPT-6 Luna · API

All factual claims about the current situation come from the supplied pilot, eval, bug, or support-ops context or follow by arithmetic.

GPT-6 Astra · ChatGPT

All factual statements about the current situation are directly from the supplied brief or context, or follow from arithmetic; no invented facts.

Addresses the actual decisionRightRightRight
Sonnet 5.5 · API

It commits to 'Go with conditions' and says what would change the call: no-go if condition 2 fails, non-billing launch if only condition 1 fails.

GPT-6 Luna · API

It commits clearly to go with conditions for Monday and says the gates that would change the call to full go or delay.

GPT-6 Astra · ChatGPT

The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).

Respects explicit constraintsRightRightRight
Sonnet 5.5 · API

Under 400 words, and the proposed eval re-run, 1%/zero refund gates, and non-billing fallback would enforce the agreed criteria.

GPT-6 Luna · API

It stays well under 400 words, addresses the launch meeting, and its proposed conditions would actually enforce the failed criteria.

GPT-6 Astra · ChatGPT

The output is a launch recommendation, addresses the meeting, and stays within the 400-word limit (345 words).

Identifies material uncertaintyRightRightRight
Sonnet 5.5 · API

It identifies the re-run pricing/refund results as the open unknowns and maps them to specific decision changes.

GPT-6 Luna · API

It names the unresolved gates—billing data rerun, refund safeguard, Welsh routing—and says before-Monday pass/fail changes the call.

GPT-6 Astra · ChatGPT

It identifies the lack of queue-level results and the unknown scope of refund failures, and explains how those would be resolved.

Avoids unsupported claimsRightRightRight
Sonnet 5.5 · API

The 0.6% estimate is clearly labelled a projection, and causes match the supplied context.

GPT-6 Luna · API

It does not present agent catch as acceptable or invent causes; its billing-exposure point is supported by the 38% billing share.

GPT-6 Astra · ChatGPT

The note about causality is framed as a methodological observation, not as a factual claim; no unsupported causes or forecasts are presented as fact.

Produces the required deliverableRightRightRight
Sonnet 5.5 · API

A complete launch recommendation with criteria table, conditions, fallback and post-launch monitoring, usable as-is.

GPT-6 Luna · API

The recommendation is complete, under the length limit, aimed at the meeting, and usable with only light edits.

GPT-6 Astra · ChatGPT

The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.

Checks every agreed criterionRightRightRight
Sonnet 5.5 · API

The table explicitly checks all four agreed criteria.

GPT-6 Luna · API

It explicitly tests all four criteria: 64% acceptance, 1.8% error rate, refund language, and 41% time reduction.

GPT-6 Astra · ChatGPT

All four launch criteria are explicitly assessed: acceptance rate, error rate, refund commitment, and time reduction.

Makes a clear callRightRightRight
Sonnet 5.5 · API

Clear 'Go with conditions' call with named fallbacks.

GPT-6 Luna · API

It makes one clear call—go with conditions, otherwise no-go—and states what would change it.

GPT-6 Astra · ChatGPT

Clear 'no-go' call up front, with a statement that reassessment follows fixes and retesting.

Gets the base of every number rightRightRightRight
Sonnet 5.5 · API

All derived rates are correct: 9/500=1.8%, 9−6=3/500=0.6%, 38% billing implies 62% non-billing, and 2.9/7.1=−41%.

GPT-6 Luna · API

Its derived percentages use correct denominators and steps: 9/500, 64% pilot acceptance, and 40.8% median-time reduction.

GPT-6 Astra · ChatGPT

All derived figures (41% time reduction, 1.8% error rate, 9/500, 38% of tickets) are computed correctly and bases are clear.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.388.32None
2GPT-6 AstrawithChatGPT86.088.32None
3Sonnet 5.5withAPI85.488.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI82.177.62None
6Gemini 3.8 FlashwithAPI71.931.521 capped
7Gemini 3.5 Flash-LitewithGemini50.642.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook