Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Addresses the actual decision95% pass
    The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong32% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows50% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Uses the supplied evidence correctly54% pass
    The output states that pilot agents were hands-on, engaged and motivated, which is not supported by the supplied context (only that 7/8 wanted to keep it).
    Opus 5.5 · Claude · Go/no-go for AI-drafted support replies

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Tomorrow's launch meeting decides whether AI-drafted support replies go live to all 42 support agents on Monday. You're the PM. Using the material below, write your launch recommendation for the meeting: go, no-go, or go with conditions, and why. Keep it under 400 words.

What the model was given6 items: Scenario, Launch criteria agreed in the PRD, Pilot (8 agents, 3 weeks, 2,940 tickets), Eval set (500 tickets), Open bugs, Support operations
ScenarioLedgerly is accounting software for small businesses. The feature drafts a first reply to each support ticket; an agent reviews and edits every draft before it is sent.
Launch criteria agreed in the PRD1. At least 60% of drafts accepted as-is or with minor edits. 2. Wrong or harmful drafts (wrong account data, wrong policy, invented features) at most 1% on the 500-ticket eval set. 3. No draft may commit to a refund. 4. Median time to first response down at least 30% in the pilot.
Pilot (8 agents, 3 weeks, 2,940 tickets)Accepted as-is or with minor edits: 64%. Median time to first response: 7.1 hours before, 4.2 hours during (−41%). Seven of the eight agents want to keep it.
Eval set (500 tickets)Wrong or harmful drafts: 9 of 500 (1.8%). Six of the nine are billing tickets quoting the old prices, which changed on 1 September; the model's reference data stops in August. Two drafts said “we'll refund this charge”; the pilot agents caught and removed both.
Open bugsReplies to Welsh-language tickets come back in English (about 0.3% of tickets). Tables in drafts are badly formatted.
Support operationsTraining for all 42 agents is booked for Friday. Billing is one of four queues and handles 38% of tickets.
What a strong answer doesThe answer key the graders mark against

Not a full go: criterion 2 fails (1.8% against 1%) and criterion 3 is breached even though agents caught it. Strong answer: go with conditions, launching to the non-billing queues while billing waits for updated pricing data and a hard block on refund language, then re-running the eval to confirm it meets the bar; names exactly what would turn it into a full go.

Critical failures (cap the score at 40)
  • Recommends a full launch while the error rate misses the agreed 1% bar
  • Treats the refund commitments as acceptable
Case

Staff-level · v1.9 · synthetic · AI product, launch

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 10 checks passed

The evidence and reasoning are sound and the correct call is clear; only a small post-launch monitoring/rollback trigger is missing and can be added quickly.

Output

Recommendation: No-go for the all-agent launch on Monday

The pilot demonstrates value, but we have not met the agreed safety gates. Agent review is a safeguard—not a substitute for those gates.

PRD criterionResultStatus
≥60% accepted as-is or with minor edits64%Pass
≤1% wrong or harmful drafts on 500 tickets9/500 = 1.8%; maximum allowed is 5Fail
No draft commits to a refundTwo drafts promised refundsFail
≥30% reduction in median first-response time41% reductionPass

Seven of eight pilot agents want to keep the feature, which supports its usefulness. However, catching the refund promises before sending does not satisfy the requirement that no draft make that commitment.

What must change before launch

  • Refresh billing reference data with current prices and verify retrieval uses it. Six of nine harmful drafts used old prices, suggesting a concentrated, fixable problem—but we must also investigate the remaining three.
  • Enforce the refund restriction, with testing that shows prohibited commitments are blocked before drafts reach agents. Agent training alone is insufficient.
  • Rerun a representative 500-ticket evaluation after fixes. Require no more than five wrong or harmful drafts and zero refund commitments; explicitly cover current pricing and refund requests.
  • Handle known quality defects: route Welsh-language tickets to manual drafting until language support is verified; suppress tables or fix their formatting.

Monday plan

Keep Friday’s training, but frame it as readiness training, not a launch commitment. Until safety fixes are verified, pause AI drafting for billing and use manual replies there.

Billing represents 38% of tickets, so excluding it is a material scope change. Any continued non-billing pilot should be explicitly approved, safety-validated, monitored, and reversible—not presented as meeting the original launch criteria. Reassess full rollout once the gates pass.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly73%
  • passAddresses the actual decision91%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty93%
  • passAvoids unsupported claims42%
  • passProduces the required deliverable67%
  • passChecks every agreed criterion99%
  • passMakes a clear call89%
  • failLimits the downside of being wrong37%
  • passGets the base of every number right87%
Run
Run
#1
API response time
12 s
Submitted
29 Sept 2026

Needs more than a quick edit3 of 10 checks passed

A PM would need to add a plan to re-run the eval set after fixes and explicit post-launch monitoring triggers; the unsupported 'training cutoff' claim also needs correction.

Output

Launch Recommendation: Go with Conditions

As the Product Manager, my recommendation for Monday’s rollout to all 42 agents is Go with Conditions.

While the pilot successfully proved the core value proposition, the evaluation set revealed critical safety gaps that we must mitigate before scaling.3

---

Why We Are Ready (The "Go" Factors)

  • Efficiency Gains Exceeded Target: Median time to first response dropped by 41% (from 7.1 to 4.2 hours), surpassing our 30% goal.
  • High Team Adoption: 7 out of 8 pilot agents want to keep using the feature.
  • Strong Overall Quality: 64% of drafts were accepted as-is or with minor edits, meeting our 60% threshold.

---

Critical Conditions for Monday's Launch

To protect our customers and brand, the following four conditions must be met prior to Monday:

1. Fix the Billing Knowledge Gap (Mandatory): * The Issue: 6 out of the 9 harmful drafts on the eval set quoted old prices because the model's training data cutoff is August4, and prices changed September 1st. * Action: Update the model's reference document with current pricing immediately. 2. Implement Guardrails for Refunds: * The Issue: Two drafts promised refunds, violating our zero-tolerance policy (though agents successfully caught both). * Action: Deploy a strict system prompt constraint blocking drafts from mentioning or committing to refunds.2 3. Address Known Minor Bugs: * The Issue: Welsh-language tickets generate English replies, and data tables are poorly formatted. * Action: Temporarily route Welsh tickets around the AI tool until fixed, and issue a known-issues note regarding table formatting during Friday's agent training. 4. Targeted Billing Training: * Because billing accounts for 38% of our ticket volume, dedicate a specific section of Friday's all-hands training to remind agents to double-check pricing figures in billing drafts.

---

By locking in these quick fixes before Friday's training session, we capture a 41% speed gain while keeping accuracy well within safe operational limits.1

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisSubstantial rework

    Source: Eval set (500 tickets)Wrong or harmful drafts: 9 of 500 (1.8%).

    What we’d changeRequire a re-run of the 500-ticket eval after the fixes, and launch only if it passes. Nothing yet shows accuracy is within the limit.

  2. 2Constraint missedRestore the constraintTargeted repair

    Source: Launch criteria agreed in the PRDNo draft may commit to a refund.

    What we’d changeA prompt instruction isn't a guarantee. Test the control, and add a check that stops refund commitments reaching agents' drafts.

  3. 3Constraint missedRestore the constraintTargeted repair

    Source: Launch criteria agreed in the PRDWrong or harmful drafts (wrong account data, wrong policy, invented features) at most 1% on the 500-ticket eval set. 3. No draft may commit to a refund.

    What we’d changeCheck each agreed criterion and say plainly that two fail: wrong or harmful drafts at 1.8% against 1%, and two refund commitments against none.

  4. 4Invented evidenceVerify or remove the claimQuick edit

    Source: Eval set (500 tickets)the model's reference data stops in August

    What we’d changeThe brief says the model's reference data stops in August, not its training data. That's what makes a pricing refresh a quick fix.

Claims the judge couldn’t find in the brief

  • The model's training data cutoff is August

Grades and run details

Decision model 55 · LLM judge 3 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly55%
  • partialAddresses the actual decision45%
  • partialRespects explicit constraints13%
  • partialIdentifies material uncertainty73%
  • partialAvoids unsupported claims27%
  • partialProduces the required deliverable46%
  • partialChecks every agreed criterion69%
  • partialMakes a clear call76%
  • failLimits the downside of being wrong96%
  • passGets the base of every number right76%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Usable with a quick edit9 of 10 checks passed

The reasoning is sound, evidence is correctly used, and the recommendation is clear. A PM could add a monitoring threshold in a quick edit without redoing any analysis.

Output

Recommendation: no-go for Monday’s rollout to all 42 agents. Reassess once the failed safety gates have been fixed and retested.

The pilot demonstrates value: 64% of drafts were accepted as-is or with minor edits, exceeding the 60% target. Median first-response time fell 41%, exceeding the 30% target, although the before-and-after comparison does not establish causality. Seven of eight agents want to keep it.

But we failed two explicit launch criteria:

  • Wrong or harmful drafts: 1.8% against a 1% ceiling. Nine failures is above the maximum five permitted on the 500-ticket set.
  • No refund commitments: failed. Two drafts promised refunds. Agents catching them prevented those messages reaching customers; it does not satisfy a criterion about what drafts may say. Human review was already part of the agreed design.

Six errors used outdated billing prices. That gives us a concrete fix, but billing represents 38% of tickets, so this is a substantial exposure. Updating the reference data alone would not establish that the remaining errors or refund promises are resolved.

Required before reconsidering launch:

  1. Engineering and the billing policy owner: update and verify current prices and policies, with a named owner and process for keeping reference data current. Investigate the other three harmful drafts.
  2. Engineering: add and test a control that prevents refund commitments from appearing in agent-facing drafts, including paraphrases.
  3. PM and evaluation owner: rerun the full 500-ticket evaluation after the fixes, meeting the original thresholds, and add targeted pricing and refund tests. Report those additional results separately.

Excluding billing is a possible narrower release, not an automatic pass: we lack queue-level results and do not know whether refund failures are confined to billing. It would require explicit scope approval and validation of both routing and the remaining queues.

Welsh-language tickets should bypass drafting until language handling is fixed. Disable tables or use plain text pending the formatting fix.

Keep Friday’s training as preparation, with launch clearly marked as pending. The booking should not determine readiness. After the gates pass, expand in stages with monitoring and a named owner able to disable drafting.

Grades and run details

Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly69%
  • passAddresses the actual decision94%
  • passRespects explicit constraints44%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims74%
  • passProduces the required deliverable86%
  • passChecks every agreed criterion49%
  • passMakes a clear call93%
  • partialLimits the downside of being wrong54%
  • passGets the base of every number right83%
Run
Run
#1
Time to output
34 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyRightMixedRight
GPT-6.1 Sol · API

All factual claims about the current situation match the supplied context or follow from arithmetic; no current-state fact is invented.

Gemini 3.5 Flash-Lite · Gemini

The claim about a 'training data cutoff' is not supported; the context only mentions reference data stopping in August, so an unsupported fact about the current system was introduced.

GPT-6 Astra · ChatGPT

All factual statements about the current situation are directly from the supplied brief or context, or follow from arithmetic; no invented facts.

Addresses the actual decisionRightWrongRight
GPT-6.1 Sol · API

It commits to a no-go for the all-agent launch and identifies gate passage after fixes as what would change the call.

Gemini 3.5 Flash-Lite · Gemini

The output does not state what result or condition would change the 'Go with Conditions' call (e.g., failing to meet the conditions would make it a no-go).

GPT-6 Astra · ChatGPT

The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).

Identifies material uncertaintyRightWrongRight
GPT-6.1 Sol · API

It names the remaining three harmful drafts and unverified billing/refund/language fixes as unknowns and says a rerun eval would resolve them.

Gemini 3.5 Flash-Lite · Gemini

It does not name specific unknowns that could change the decision or say how they would be resolved, such as whether the pricing fix alone will bring error rate below 1%.

GPT-6 Astra · ChatGPT

It identifies the lack of queue-level results and the unknown scope of refund failures, and explains how those would be resolved.

Avoids unsupported claimsRightWrongRight
GPT-6.1 Sol · API

Hypotheses are marked as suggestions, and conclusions do not present unsupported current facts as established.

Gemini 3.5 Flash-Lite · Gemini

The output presents the 'training data cutoff' as a fact when the source only says reference data stops in August, turning an interpretation into an unsupported claim.

GPT-6 Astra · ChatGPT

The note about causality is framed as a methodological observation, not as a factual claim; no unsupported causes or forecasts are presented as fact.

Checks every agreed criterionRightWrongRight
GPT-6.1 Sol · API

The table checks all four agreed launch criteria against the evidence.

Gemini 3.5 Flash-Lite · Gemini

It does not explicitly compare the 1.8% harmful rate against the 1% bar, leaving the second launch criterion unchecked.

GPT-6 Astra · ChatGPT

All four launch criteria are explicitly assessed: acceptance rate, error rate, refund commitment, and time reduction.

Makes a clear callRightWrongRight
GPT-6.1 Sol · API

It makes a single clear call—no-go for the all-agent launch Monday—and states reassessment once gates pass.

Gemini 3.5 Flash-Lite · Gemini

The call is clear but does not state what would change it, e.g., that failing to meet the conditions would lead to a no-go.

GPT-6 Astra · ChatGPT

Clear 'no-go' call up front, with a statement that reassessment follows fixes and retesting.

All got wrong 1

Limits the downside of being wrongWrongWrongWrong
GPT-6.1 Sol · API

It does not specify a post-launch monitoring signal, threshold or trigger, rollback action, or irreversible exposure if the recommendation is wrong.

Gemini 3.5 Flash-Lite · Gemini

No post-launch signal, threshold to pause or roll back, or irreversible actions are identified.

GPT-6 Astra · ChatGPT

No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.

All got right 3

Respects explicit constraintsRightRightRight
GPT-6.1 Sol · API

The output respects the under-400-word brief, makes a go/no-go/conditions recommendation, and proposes concrete enforcement such as a refund block and rerun eval.

Gemini 3.5 Flash-Lite · Gemini

The output stays under 400 words and respects the requested form of a launch recommendation for the meeting.

GPT-6 Astra · ChatGPT

The output is a launch recommendation, addresses the meeting, and stays within the 400-word limit (345 words).

Produces the required deliverableRightRightRight
GPT-6.1 Sol · API

A usable launch recommendation is present, complete, and within length for the meeting.

Gemini 3.5 Flash-Lite · Gemini

The recommendation is under 400 words, addressed to the meeting, and usable with light edits, though it lacks a retest plan.

GPT-6 Astra · ChatGPT

The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.

Gets the base of every number rightRightRightRight
GPT-6.1 Sol · API

Derived rates such as 1.8%, maximum of 5, and 41% reduction use the correct denominators and steps.

Gemini 3.5 Flash-Lite · Gemini

The 41% time reduction uses the correct base (7.1 hours to 4.2 hours), and other percentages match the supplied data.

GPT-6 Astra · ChatGPT

All derived figures (41% time reduction, 1.8% error rate, 9/500, 38% of tickets) are computed correctly and bases are clear.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.388.32None
2GPT-6 AstrawithChatGPT86.088.32None
3Sonnet 5.5withAPI85.488.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI82.177.62None
6Gemini 3.8 FlashwithAPI71.931.521 capped
7Gemini 3.5 Flash-LitewithGemini50.642.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook