Tasks / Operate

Make the launch call

Can the model make a clear go / no-go call from mixed launch evidence, checked against the criteria agreed up front?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Finds the Scotland breach100% pass
    Clearly identifies Scotland's refund breach (66% above control, 143% after supplier change) and acceptance dip below 80%, correctly holds Scotland.
    Opus 5.5 · Claude · Go/no-go for AI substitutions, from the rollout data
  2. Produces the required deliverable98% pass
    The recommendation is in the requested form, for the meeting, within the length, and includes necessary steps a PM could act on.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  3. Addresses the actual decision95% pass
    The output gives a clear, unambiguous 'no-go' call early, and states what would change it (fixes and retest).
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies

Where it slips

  1. Limits the downside of being wrong32% pass
    No specific post-launch monitoring signal or threshold is named; only generic 'monitoring' and a kill switch are mentioned.
    GPT-6 Astra · ChatGPT · Go/no-go for AI-drafted support replies
  2. Catches the duplicate rows50% pass
    Finds and removes the duplicate rows but never states whether they change the results.
    GPT-6.1 Sol · API · Go/no-go for AI substitutions, from the rollout data
  3. Uses the supplied evidence correctly54% pass
    The output states that pilot agents were hands-on, engaged and motivated, which is not supported by the supplied context (only that 7/8 wanted to keep it).
    Opus 5.5 · Claude · Go/no-go for AI-drafted support replies

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Tomorrow's launch meeting decides whether AI-drafted support replies go live to all 42 support agents on Monday. You're the PM. Using the material below, write your launch recommendation for the meeting: go, no-go, or go with conditions, and why. Keep it under 400 words.

What the model was given6 items: Scenario, Launch criteria agreed in the PRD, Pilot (8 agents, 3 weeks, 2,940 tickets), Eval set (500 tickets), Open bugs, Support operations
ScenarioLedgerly is accounting software for small businesses. The feature drafts a first reply to each support ticket; an agent reviews and edits every draft before it is sent.
Launch criteria agreed in the PRD1. At least 60% of drafts accepted as-is or with minor edits. 2. Wrong or harmful drafts (wrong account data, wrong policy, invented features) at most 1% on the 500-ticket eval set. 3. No draft may commit to a refund. 4. Median time to first response down at least 30% in the pilot.
Pilot (8 agents, 3 weeks, 2,940 tickets)Accepted as-is or with minor edits: 64%. Median time to first response: 7.1 hours before, 4.2 hours during (−41%). Seven of the eight agents want to keep it.
Eval set (500 tickets)Wrong or harmful drafts: 9 of 500 (1.8%). Six of the nine are billing tickets quoting the old prices, which changed on 1 September; the model's reference data stops in August. Two drafts said “we'll refund this charge”; the pilot agents caught and removed both.
Open bugsReplies to Welsh-language tickets come back in English (about 0.3% of tickets). Tables in drafts are badly formatted.
Support operationsTraining for all 42 agents is booked for Friday. Billing is one of four queues and handles 38% of tickets.
What a strong answer doesThe answer key the graders mark against

Not a full go: criterion 2 fails (1.8% against 1%) and criterion 3 is breached even though agents caught it. Strong answer: go with conditions, launching to the non-billing queues while billing waits for updated pricing data and a hard block on refund language, then re-running the eval to confirm it meets the bar; names exactly what would turn it into a full go.

Critical failures (cap the score at 40)
  • Recommends a full launch while the error rate misses the agreed 1% bar
  • Treats the refund commitments as acceptable
Case

Staff-level · v1.9 · synthetic · AI product, launch

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit3 of 10 checks passed

A PM would need to re-check the non-billing error rate, refund-draft queue attribution, and monitoring/rollback plan before using it. The core recommendation is useful but not evidence-safe as written.

Output

Recommendation: Go with Conditions

While the pilot demonstrated strong operational impact—reducing median first-response time by 41% (beating our 30% target) with a 64% acceptance rate—we failed two mandatory launch criteria:

  1. Accuracy (Criterion 2): Harmful/wrong drafts hit 1.8% (9/500) versus the $\le$1.0% threshold.
  2. Refund commitments (Criterion 3): Two drafts explicitly promised refunds, violating our zero-tolerance policy.

However, these failures are isolated and addressable. Six of the nine errors were caused by August reference data lacking the September 1 price changes, and both refund commitments originated in billing workflows.

Rather than cancelling Friday’s 42-agent training or delaying the productivity benefits across the entire organization, we should launch on Monday under the following three conditions:

Launch Conditions

1. Disable AI drafting on the Billing queue for Monday: Billing represents 38% of volume and contains almost all our critical failure points (pricing errors and refund risks). We will launch exclusively to the remaining three queues (62% of ticket volume), where accuracy meets our threshold. 2. Implement deterministic refund guardrails by Friday: Deploy a hard system-prompt restriction and a post-generation regex/keyword filter that blocks or drops drafts containing refund commitments before reaching the agent. 3. Update agent training on Friday: Keep Friday’s training on schedule. Explicitly train agents on: - The pilot's phased rollout (Billing queue excluded initially). - Known minor bugs (Welsh-language tickets defaulting to English; manual table reformatting required). - Strict agent accountability as the final human reviewer for every sent message.

Path to Full Billing Rollout

Engineering will ingest the post-September 1 pricing documentation and run a targeted 200-ticket eval on billing drafts next week. Once billing tickets achieve $<1\%$ error and 0 refund promises, we will enable the feature for the billing queue.

This approach protects Ledgerly’s financial liability and customer trust while immediately capturing significant speed improvements for over 60% of our support volume.

Claims the judge couldn’t find in the brief

  • Both refund commitments originated in billing workflows.
  • Billing contains almost all critical failure points, including pricing errors and refund risks.
  • The remaining three queues meet the accuracy threshold.

Grades and run details

Decision model 90 · LLM judge 3 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly45%
  • passAddresses the actual decision100%
  • passRespects explicit constraints54%
  • passIdentifies material uncertainty96%
  • partialAvoids unsupported claims27%
  • passProduces the required deliverable96%
  • passChecks every agreed criterion24%
  • passMakes a clear call100%
  • partialLimits the downside of being wrong38%
  • passGets the base of every number right86%
Run
Run
#1
API response time
13 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 10 checks passed

The recommendation is complete, evidence-based, and actionable; the only meaningful gap is a missing post-launch pause threshold, which is a quick edit.

Output

Recommendation: Go with conditions

Not a clean go. Two of the four agreed launch criteria are currently failed, and I don't think we should waive criteria we set ourselves.

CriterionResultStatus
≥60% accepted64%Met
≤1% wrong/harmful1.8% (9/500)Failed
No refund commitments2 drafts promised refundsFailed
≥30% faster first response−41% (7.1h → 4.2h)Met

The value is real: faster responses, good acceptance, and 7 of 8 pilot agents want to keep it. Both failures also have identifiable causes.

  • Wrong drafts: Six of the nine errors come from stale pricing, because the reference data stops in August. Fixing that should bring us to about 3/500 (0.6%), but that is a projection until we re-run the eval.
  • Refunds: Agents caught both drafts, but the criterion says no draft, and with 42 agents we shouldn't rely on catching them by habit.

Conditions for Monday

  1. Refresh the reference data with post-1 September pricing and re-run the full 500-ticket eval. It must come in at 1% or below (5 or fewer wrong).
  2. Add a refund guardrail that blocks or flags commitment language such as "we'll refund." It must show zero refund commitments on the re-run plus a targeted set of refund-request tickets.
  3. Kill switch and flag button. Ops can disable the feature instantly, and agents can flag bad drafts in one click.

Fallback

  • If condition 2 fails, no-go for Monday.
  • If only condition 1 fails, launch on the three non-billing queues (62% of tickets) and keep billing off until it passes.
  • Friday training goes ahead for all 42 agents either way. It should cover the refund rule and checking prices before sending.

Known issues, not blockers

  • Welsh tickets (~0.3%) get English replies. Agents will write these manually, and we'll suppress drafts for them if detection is feasible.
  • Table formatting is cosmetic, since agents edit every draft. We'll fix it after launch.

After launch

  • Sample wrong-draft rates weekly for the first month.
  • Track how often agents edit or discard drafts, as a check on whether review discipline slips.
  • Ask the one pilot agent who didn't want to keep it why.

I'd like the eval re-run results by Thursday so we can make the final call before Friday's training.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly44%
  • passAddresses the actual decision98%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable84%
  • passChecks every agreed criterion100%
  • passMakes a clear call100%
  • partialLimits the downside of being wrong85%
  • passGets the base of every number right79%
Run
Run
#1
API response time
17 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 10 checks passed

The evidence and reasoning are sound and the correct call is clear; only a small post-launch monitoring/rollback trigger is missing and can be added quickly.

Output

Recommendation: No-go for the all-agent launch on Monday

The pilot demonstrates value, but we have not met the agreed safety gates. Agent review is a safeguard—not a substitute for those gates.

PRD criterionResultStatus
≥60% accepted as-is or with minor edits64%Pass
≤1% wrong or harmful drafts on 500 tickets9/500 = 1.8%; maximum allowed is 5Fail
No draft commits to a refundTwo drafts promised refundsFail
≥30% reduction in median first-response time41% reductionPass

Seven of eight pilot agents want to keep the feature, which supports its usefulness. However, catching the refund promises before sending does not satisfy the requirement that no draft make that commitment.

What must change before launch

  • Refresh billing reference data with current prices and verify retrieval uses it. Six of nine harmful drafts used old prices, suggesting a concentrated, fixable problem—but we must also investigate the remaining three.
  • Enforce the refund restriction, with testing that shows prohibited commitments are blocked before drafts reach agents. Agent training alone is insufficient.
  • Rerun a representative 500-ticket evaluation after fixes. Require no more than five wrong or harmful drafts and zero refund commitments; explicitly cover current pricing and refund requests.
  • Handle known quality defects: route Welsh-language tickets to manual drafting until language support is verified; suppress tables or fix their formatting.

Monday plan

Keep Friday’s training, but frame it as readiness training, not a launch commitment. Until safety fixes are verified, pause AI drafting for billing and use manual replies there.

Billing represents 38% of tickets, so excluding it is a material scope change. Any continued non-billing pilot should be explicitly approved, safety-validated, monitored, and reversible—not presented as meeting the original launch criteria. Reassess full rollout once the gates pass.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly73%
  • passAddresses the actual decision91%
  • passRespects explicit constraints37%
  • passIdentifies material uncertainty93%
  • passAvoids unsupported claims42%
  • passProduces the required deliverable67%
  • passChecks every agreed criterion99%
  • passMakes a clear call89%
  • failLimits the downside of being wrong37%
  • passGets the base of every number right87%
Run
Run
#1
API response time
12 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyMixedRightRight
Gemini 3.8 Flash · API

It invents or assumes that both refund drafts were billing-related, that non-billing queues meet the accuracy threshold, and that billing contains almost all critical failure points.

Sonnet 5.5 · API

All current-state facts used are from the supplied context; no invented systems or numbers.

GPT-6.1 Sol · API

All factual claims about the current situation match the supplied context or follow from arithmetic; no current-state fact is invented.

Respects explicit constraintsMixedRightRight
Gemini 3.8 Flash · API

The proposed non-billing launch is not shown to enforce the 1% wrong/harmful criterion, and the refund guardrail proposal is not evidenced as sufficient.

Sonnet 5.5 · API

Under 400 words, and the proposed eval re-run, 1%/zero refund gates, and non-billing fallback would enforce the agreed criteria.

GPT-6.1 Sol · API

The output respects the under-400-word brief, makes a go/no-go/conditions recommendation, and proposes concrete enforcement such as a refund block and rerun eval.

Identifies material uncertaintyMixedRightRight
Gemini 3.8 Flash · API

It does not name the key unknowns that could change the call, such as the non-billing error rate, whether refund drafts were billing-related, or whether guardrails actually prevent refund commitments.

Sonnet 5.5 · API

It identifies the re-run pricing/refund results as the open unknowns and maps them to specific decision changes.

GPT-6.1 Sol · API

It names the remaining three harmful drafts and unverified billing/refund/language fixes as unknowns and says a rerun eval would resolve them.

Avoids unsupported claimsWrongRightRight
Gemini 3.8 Flash · API

It presents unsupported causal and coverage claims as fact, including that refund commitments originated in billing workflows and that non-billing accuracy meets the threshold.

Sonnet 5.5 · API

The 0.6% estimate is clearly labelled a projection, and causes match the supplied context.

GPT-6.1 Sol · API

Hypotheses are marked as suggestions, and conclusions do not present unsupported current facts as established.

Checks every agreed criterionMixedRightRight
Gemini 3.8 Flash · API

It checks criteria 2, 3, and 4, but does not explicitly test criterion 1 against the 60% acceptance bar.

Sonnet 5.5 · API

The table explicitly checks all four agreed criteria.

GPT-6.1 Sol · API

The table checks all four agreed launch criteria against the evidence.

Gets the base of every number rightMixedRightRight
Gemini 3.8 Flash · API

It uses the 62% volume base correctly but treats non-billing accuracy as meeting the threshold without a verified denominator or error count.

Sonnet 5.5 · API

All derived rates are correct: 9/500=1.8%, 9−6=3/500=0.6%, 38% billing implies 62% non-billing, and 2.9/7.1=−41%.

GPT-6.1 Sol · API

Derived rates such as 1.8%, maximum of 5, and 41% reduction use the correct denominators and steps.

All got wrong 1

Limits the downside of being wrongWrongWrongWrong
Gemini 3.8 Flash · API

It does not specify post-launch monitoring, a pause or rollback threshold, or irreversible customer-communication risks.

Sonnet 5.5 · API

It describes post-launch monitoring but does not name a threshold or trigger after launch to pause/roll back, nor flag anything irreversible.

GPT-6.1 Sol · API

It does not specify a post-launch monitoring signal, threshold or trigger, rollback action, or irreversible exposure if the recommendation is wrong.

All got right 3

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

It clearly recommends go with conditions for Monday and states the billing re-evaluation result that would enable full rollout.

Sonnet 5.5 · API

It commits to 'Go with conditions' and says what would change the call: no-go if condition 2 fails, non-billing launch if only condition 1 fails.

GPT-6.1 Sol · API

It commits to a no-go for the all-agent launch and identifies gate passage after fixes as what would change the call.

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

It is a concise launch recommendation in the requested form and appears usable for the meeting with light edits.

Sonnet 5.5 · API

A complete launch recommendation with criteria table, conditions, fallback and post-launch monitoring, usable as-is.

GPT-6.1 Sol · API

A usable launch recommendation is present, complete, and within length for the meeting.

Makes a clear callRightRightRight
Gemini 3.8 Flash · API

It makes a clear go-with-conditions call and states the billing rollout condition.

Sonnet 5.5 · API

Clear 'Go with conditions' call with named fallbacks.

GPT-6.1 Sol · API

It makes a single clear call—no-go for the all-agent launch Monday—and states reassessment once gates pass.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.388.32None
2GPT-6 AstrawithChatGPT86.088.32None
3Sonnet 5.5withAPI85.488.32None
4Opus 5.5withClaude80.486.42None
5GPT-6 LunawithAPI82.177.62None
6Gemini 3.8 FlashwithAPI71.931.521 capped
7Gemini 3.5 Flash-LitewithGemini50.642.22None

About the task

The PM job

Deciding whether a feature ships on the planned date, and on what conditions.

Why it matters

Launch meetings reward optimism. A good call checks each agreed criterion, weighs a real risk against a real gain, and says exactly what would change the answer. A weak one rubber-stamps the launch or blocks it on noise.

What good looks like

  • Checks each agreed criterion against the evidence
  • Makes one clear call, with conditions if needed
  • Separates risks that block launch from risks that can be managed
  • Says what would change the call
  • Says how to limit the damage if the call is wrong

Deliberately not measured

  • Rollout engineering detail
  • Project-plan formatting
Capability tested

Deciding against agreed launch criteria

The failure we’re looking for

Rubber-stamps a launch that misses an agreed bar, or blocks it on noise

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

A launch memo from supplied evidence · Staff level: a regional call from an attached data workbook