Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 71% were usable with at most a quick edit.

Reliably right

  1. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test
  3. Checks guardrails before declaring a winner100% pass
    It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
    GPT-6 Astra · ChatGPT · Conversion up, retention down

Where it slips

  1. Trusts the data before reading it21% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims52% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly55% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our new pricing page variant (B) increased trial-to-paid conversion, and the growth team wants to ship it to 100% on Monday. Review the readout below and write your recommendation to Priya Nair, Head of Growth, who makes the call. Lead with what we should do, then give the reasoning she needs to defend it in the Monday meeting. Keep it under 400 words.

What the model was given5 items: Scenario, Readout summary, Variant description, Guardrails agreed before launch, Refunds
ScenarioLumen is a subscription app for learning languages. The monthly plan is $14.99. The annual plan is $119.99, billed up front, which works out at $9.99 a month.
Readout summaryTest ran 21 days, 50/50 split, 38,412 trial starts. Trial→paid conversion: A 11.8%, B 13.1% (+1.3pp, 95% CI +0.6 to +2.0). Day-30 paid retention: A 84.0%, B 79.5% (−4.5pp, 95% CI −7.9 to −1.1). Revenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).
Variant descriptionVariant B leads with the annual plan's monthly-equivalent price ($9.99/month) and moves the monthly plan behind a 'See all plans' link.
Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.
RefundsRefund requests during the test: A 212, B 301 (+42%).
What a strong answer doesThe answer key the graders mark against

Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.

Critical failures (cap the score at 40)
  • Recommends shipping B to 100%
  • States the mechanism as established fact
Case

Mid-level · v1.9 · anonymised real · pricing, guardrail breach, B2C SaaS

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 10 checks passed

The recommendation is clear, complete, and evidence-based, with only possible light wording edits needed.

Output

o: Priya Nair, Head of Growth**

Do not ship B to 100% on Monday. Return new traffic to A, continue following existing test cohorts and test a revised page with clearer pricing. B improves conversion, but the readout does not meet the conditions we agreed for rollout.

The conversion lift is credible: +1.3 percentage points, with a 95% confidence interval of +0.6 to +2.0. However, conversion alone cannot establish that B creates better customer or business outcomes.

The guardrails do not support launch. Reported day-30 retention falls 4.5pp against a maximum permitted decline of 2pp. Its confidence interval includes declines within our tolerance, so we cannot say the true loss definitively exceeds 2pp.1 Equally, we have not established that B is within tolerance.

Refund requests rose from 212 to 301, or 42%, against a 10% limit. Confirm whether the agreed guardrail measures counts or rates, using comparable denominators and observation windows. B generates more paying customers, which could explain some additional requests, but the increase warrants investigation before expansion.

There is also a measurement issue: a 21-day test cannot give every participant 30 days of follow-up. Establish when this readout was produced, which cohorts matured and how retention is defined across monthly and annual plans2. If measured immediately after the test, the day-30 figures cannot be complete.

Revenue per trial is only $0.19 higher, with uncertainty spanning zero. That does not establish a revenue benefit, and annual billing up front makes it especially important to distinguish cash collected from earned revenue and account for refunds.

B’s presentation offers a plausible explanation worth testing: “$9.99/month” may attract customers who do not understand the $119.99 upfront commitment. The readout does not prove this; inspect refund reasons and results by plan.

For the next test, display “$119.99 billed annually, equivalent to $9.99/month” prominently and keep the $14.99 monthly option readily visible. Retain the original guardrails and define mature cohort reporting before launch.

For Monday, the defensible decision is simple: the conversion gain earns further testing, but does not justify overriding our agreed customer safeguards.

What a PM had to fix

  1. 1Decision deferredMake the callTargeted repair

    Source: Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.

    What we’d changeMake the argument rather than hedging it: the point estimate breaches the guardrail and most of the interval is past it. The refund question can be settled from the numbers given, so settle it.

  2. 2Contradiction missedSurface the contradictionQuick edit

    Source: ScenarioThe annual plan is $119.99, billed up front

    What we’d changeMake the plan-mix point: annual plans are paid up front, so if B moved people onto annual, day-30 retention should have risen on its own. That it fell strengthens the case against B.

Grades and run details

Decision model 100 · LLM judge 11 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly17%
  • passAddresses the actual decision74%
  • passRespects explicit constraints29%
  • passIdentifies material uncertainty85%
  • passAvoids unsupported claims46%
  • passProduces the required deliverable48%
  • passChecks guardrails before declaring a winner65%
  • passSeparates effect from explanation95%
  • passTrusts the data before reading it92%
  • passGets the base of every number right73%
Run
Run
#1
Time to output
60 s
Submitted
25 Sept 2026

Needs more than a quick edit6 of 10 checks passed

Contains unsupported claims about regulatory scrutiny and treats a causal inference as fact, which would need to be removed or qualified before it can be used without re-checking.

Output

Subject: Pricing variant B — recommend we don't ship Monday

Priya,

Recommendation: don't roll B out to 100%. It breaches both guardrails we agreed before launch. Run a corrected variant (C) instead.

Why B fails

  • Refunds: +42% against a 10% limit. This is the clearest breach, and it isn't close. (212 → 301 requests.)
  • Day-30 retention: −4.5pp against a 2pp limit.2 The confidence interval runs from −7.9 to −1.1. Most of that range is past the line, and none of it is positive.
  • The upside isn't real in revenue terms.1 Conversion rose 1.3pp, but revenue per trial start at day 30 ($9.61 vs $9.42) isn't statistically different from zero. We'd be taking on the harm for no demonstrable gain.
  • The day-30 revenue number flatters B. Annual plans bill $119.99 up front, so a shift toward annual inflates early revenue. If the extra refunds and churn continue past day 30, B likely falls behind A on lifetime value.

What's probably happening

B leads with "$9.99/month" and hides the monthly plan. Some users think they're signing up for $9.99 a month, then get charged $119.99.3 The spike in refunds and early cancellations fits that pattern: the extra conversions are partly people who didn't understand what they bought. We can confirm this in a day by reading refund reasons and support tickets from B users.

This also creates a regulatory and reputational risk. Showing a monthly-equivalent price while billing annually, with the cheaper-looking commitment-free option hidden, is exactly the kind of pricing presentation the CMA and FTC scrutinise.4

How to defend this Monday

  • We set the guardrails before seeing results so we wouldn't argue ourselves past them afterward. Overriding them now undermines every future test.
  • "Conversion went up" measures the wrong thing. Revenue didn't move, and refunds and churn got worse.

Proposed next step: Variant C

  • Keep the annual-first emphasis, which is likely what drove the real part of the lift.
  • Show "$119.99 billed annually" next to the $9.99/month equivalent, at equal prominence.
  • Keep the monthly plan visible on the page.
  • Use the same guardrails, and track revenue per trial through day 60 so annual billing doesn't distort the result.

If C keeps most of the conversion lift within the guardrails, we ship that.

Martin

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: Readout summaryRevenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).

    What we’d changeAn interval crossing zero means the revenue lift is uncertain, not absent. Say that instead.

  2. 2Contradiction missedSurface the contradictionQuick edit

    Source: Readout summaryTest ran 21 days

    What we’d changeAsk how a 21-day test reports day-30 retention before relying on the figure.

  3. 3Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeLabel this as the likely explanation to confirm from refund reasons, as the next sentence half does.

  4. 4Invented evidenceVerify or remove the claimQuick edit

    What we’d changeCut the regulatory claim. It isn't in the evidence, and it gives the room something to argue about.

Claims the judge couldn’t find in the brief

  • The day-30 revenue number flatters B.
  • Showing a monthly-equivalent price while billing annually, with the cheaper-looking commitment-free option hidden, is exactly the kind of pricing presentation the CMA and FTC scrutinise.

Grades and run details

Decision model 75 · LLM judge 6 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly19%
  • passAddresses the actual decision98%
  • passRespects explicit constraints36%
  • passIdentifies material uncertainty85%
  • partialAvoids unsupported claims47%
  • passProduces the required deliverable85%
  • passChecks guardrails before declaring a winner99%
  • passSeparates effect from explanation78%
  • failTrusts the data before reading it70%
  • passGets the base of every number right84%
Run
Run
#1
Time to output
20 s
Submitted
25 Sept 2026

Usable with a quick edit8 of 10 checks passed

The decision, guardrail check, and recommendation are complete and actionable. Only minor wording or trust-check additions would be needed, not redoing reasoning.

Output

Recommendation: Don't ship B to 100% on Monday. B breached both guardrails we agreed on before launch. Keep the test running or revert to A, and bring a fix to the meeting instead.

Why

  1. The guardrails were set to settle this. Refund requests rose 42% (212 → 301) against a 10% limit. That's a clear breach, and at this volume it isn't noise. Day-30 retention fell 4.5pp against a 2pp limit. The CI (−7.9 to −1.1) technically includes drops smaller than 2pp, but the point estimate is more than double the limit, so we can't claim it's safe.
  1. The conversion win isn't yet a revenue win. Conversion is up 1.3pp and that's real. But revenue per trial start is $9.61 vs $9.42, and the CI crosses zero. We are converting more people, and more of them are leaving or asking for their money back.
  1. The retention drop is surprising. If B pushes people toward the annual plan, which is billed up front, day-30 retention should rise, not fall. A drop alongside more refunds suggests some buyers are surprised by a $119.99 charge after seeing "$9.99/month." Hiding the monthly plan behind "See all plans" may also pull in lower-intent buyers. This is a hypothesis, not a finding, since the readout doesn't show plan mix.
  1. Day 30 undersells the risk. Annual buyers who regret the purchase may refund later, and the readout can't show that.

What to do before Monday

  • Pull plan mix (monthly vs. annual) by variant, plus retention and refund rates by plan.
  • Check whether the $119.99 up-front charge is clear enough on B. Refund reasons or support tickets would help.
  • Confirm whether revenue is net of refunds.

What to propose instead

Test a B2 that keeps the $9.99/month anchor but states "$119.99 billed today" next to it and keeps the monthly option visible. If B2 holds the conversion gain and stays within both guardrails, ship it.

Line for the meeting: "We set guardrails so we wouldn't have to debate this after the fact. B failed them, so we're fixing the billing clarity and retesting, not shipping a conversion gain that may be costing us customers and refunds."

If leadership insists on shipping B, the least risky option is a partial rollout with the refund rate monitored weekly and a pre-agreed rollback trigger. I wouldn't recommend it.

Claims the judge couldn’t find in the brief

  • At this volume the refund increase is not noise.
  • If B pushes people toward the annual plan, day-30 retention should rise rather than fall because the annual plan is billed up front.

Grades and run details

Decision model 80 · LLM judge 9 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly24%
  • passAddresses the actual decision99%
  • partialRespects explicit constraints35%
  • passIdentifies material uncertainty98%
  • partialAvoids unsupported claims72%
  • passProduces the required deliverable78%
  • passChecks guardrails before declaring a winner94%
  • passSeparates effect from explanation93%
  • failTrusts the data before reading it25%
  • passGets the base of every number right81%
Run
Run
#1
API response time
16 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyRightWrongRight
GPT-6 Astra · ChatGPT

The figures and current-state statements are taken correctly from the supplied context or follow by arithmetic.

Opus 5.5 · Claude

Output presents unsupported claims about regulatory scrutiny and the revenue number flattering B, which are not in the brief and not derived by arithmetic.

Sonnet 5.5 · API

All core metrics, guardrail numbers, and CI values are quoted correctly from the supplied context, with no invented current-situation facts.

Avoids unsupported claimsRightWrongWrong
GPT-6 Astra · ChatGPT

The pricing-misunderstanding explanation is explicitly labelled plausible and worth testing rather than stated as fact.

Opus 5.5 · Claude

The claim about the revenue number flattering B is presented as fact without being labeled as a hypothesis, and the regulatory risk claim is stated as a fact with no support.

Sonnet 5.5 · API

It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.

Separates effect from explanationRightMixedRight
GPT-6 Astra · ChatGPT

It separates the observed conversion and retention effects from the hypothesised pricing-misunderstanding mechanism.

Opus 5.5 · Claude

The explanation that the revenue number flatters B is not labelled as a hypothesis and is presented as a finding.

Sonnet 5.5 · API

The buyer-surprise mechanism is labeled as a hypothesis rather than a finding, and alternative explanations are framed with 'may' and 'suggests'.

Trusts the data before reading itRightWrongWrong
GPT-6 Astra · ChatGPT

It checks a data-trust issue by flagging that a 21-day test cannot fully support day-30 retention for all users.

Opus 5.5 · Claude

No trust signal check is performed; the analysis does not verify the sample split ratio or any other data-quality check before interpreting effects.

Sonnet 5.5 · API

It does not check any trust signal such as sample ratio, exposure, logging issues, or novelty before interpreting the results.

All got right 6

Addresses the actual decisionRightRightRight
GPT-6 Astra · ChatGPT

It commits early to not shipping B, returning new traffic to A, and testing a revised page, while making clear the guardrails are the gating conditions.

Opus 5.5 · Claude

Recommends clearly not to ship B and names the decision-maker, with a condition for the next variant (if C keeps lift within guardrails).

Sonnet 5.5 · API

It clearly and immediately says don't ship B to 100%, names Priya's decision, and says a B2 that holds conversion and passes guardrails would change the call.

Respects explicit constraintsRightRightRight
GPT-6 Astra · ChatGPT

It addresses Priya, stays under the requested length, and its recommendation respects the agreed guardrails rather than overriding them.

Opus 5.5 · Claude

The output is well under 400 words and addresses the head of growth as requested.

Sonnet 5.5 · API

It stays under 400 words, addresses Priya, leads with the recommendation, and proposes enforcement of guardrails via follow-up test and rollback trigger.

Identifies material uncertaintyRightRightRight
GPT-6 Astra · ChatGPT

It names the material unknowns—retention follow-up maturity, refund denominator, plan-level results—and says how to resolve them.

Opus 5.5 · Claude

Identifies that the revenue lift is not significant, suggests checking refund reasons to confirm hypothesis, and proposes following variant to reduce uncertainty.

Sonnet 5.5 · API

It names missing plan mix, refund reasons, and whether revenue is net of refunds, and says how to resolve them or what would change the call.

Produces the required deliverableRightRightRight
GPT-6 Astra · ChatGPT

It is a usable decision memo for the named reader with a clear recommendation, reasoning, and next steps.

Opus 5.5 · Claude

The deliverable is a complete recommendation with reasoning and next steps, suitable for the named reader.

Sonnet 5.5 · API

It provides a usable, complete recommendation with reasoning and next steps for Priya within the length limit.

Checks guardrails before declaring a winnerRightRightRight
GPT-6 Astra · ChatGPT

It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.

Opus 5.5 · Claude

Explicitly evaluates both pre-agreed guardrails and confirms they are breached.

Sonnet 5.5 · API

It evaluates both pre-agreed guardrails explicitly, comparing refunds to the 10% limit and retention to the 2pp limit.

Gets the base of every number rightRightRightRight
GPT-6 Astra · ChatGPT

The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.

Opus 5.5 · Claude

All percentages and differences are correctly taken from the supplied data, and the base for the refund increase is clear.

Sonnet 5.5 · API

Percentage differences and refund increase are computed from the correct bases and are consistent with the supplied data.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT100.095.02None
2GPT-6.1 SolwithAPI94.790.52None
3GPT-6 LunawithAPI91.990.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude73.657.32None
6Gemini 3.5 Flash-LitewithGemini60.357.721 capped
7Gemini 3.8 FlashwithAPI55.624.121 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review