Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 71% were usable with at most a quick edit.

Reliably right

  1. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test
  3. Checks guardrails before declaring a winner100% pass
    It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
    GPT-6 Astra · ChatGPT · Conversion up, retention down

Where it slips

  1. Trusts the data before reading it21% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims52% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly55% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our new pricing page variant (B) increased trial-to-paid conversion, and the growth team wants to ship it to 100% on Monday. Review the readout below and write your recommendation to Priya Nair, Head of Growth, who makes the call. Lead with what we should do, then give the reasoning she needs to defend it in the Monday meeting. Keep it under 400 words.

What the model was given5 items: Scenario, Readout summary, Variant description, Guardrails agreed before launch, Refunds
ScenarioLumen is a subscription app for learning languages. The monthly plan is $14.99. The annual plan is $119.99, billed up front, which works out at $9.99 a month.
Readout summaryTest ran 21 days, 50/50 split, 38,412 trial starts. Trial→paid conversion: A 11.8%, B 13.1% (+1.3pp, 95% CI +0.6 to +2.0). Day-30 paid retention: A 84.0%, B 79.5% (−4.5pp, 95% CI −7.9 to −1.1). Revenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).
Variant descriptionVariant B leads with the annual plan's monthly-equivalent price ($9.99/month) and moves the monthly plan behind a 'See all plans' link.
Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.
RefundsRefund requests during the test: A 212, B 301 (+42%).
What a strong answer doesThe answer key the graders mark against

Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.

Critical failures (cap the score at 40)
  • Recommends shipping B to 100%
  • States the mechanism as established fact
Case

Mid-level · v1.9 · anonymised real · pricing, guardrail breach, B2C SaaS

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: States the mechanism as established fact. Score capped at 40.

Needs more than a quick edit2 of 10 checks passed

A PM would need to remove unsupported causal and forecast claims and add decision-changing conditions. It also needs addressing to Priya before use.

Output

Recommendation: Do not ship Variant B. Keep Variant A live and do not roll B out to 100% on Monday.

While Variant B delivered an initial +1.3pp lift in trial-to-paid conversion, it violates both pre-agreed guardrails, creates bill shock, and delivers no statistically significant revenue upside.

Here is the reasoning to defend this decision on Monday:

1. Both Pre-Agreed Guardrails Were Blown

We agreed on strict risk thresholds before launch. Variant B failed both: * Day-30 Retention: Fell −4.5pp (down to 79.5%), more than double our maximum allowed drop of 2pp (95% CI: −7.9 to −1.1). * Refund Requests: Surged +42% (301 vs. 212), four times our 10% ceiling.

Overriding guardrails after the fact sets a dangerous precedent and invites compounding churn.

2. The Conversion Lift Is an Illusion Driven by Bill Shock

Variant B hides the monthly tier and frames the annual plan as “$9.99/month.” Users clearly believe they are committing to a flexible monthly rate, only to be charged $119.99 upfront. The 42% spike in refunds and immediate drop in 30-day retention prove that these are not high-intent, qualified subscribers—they are confused customers who churn or dispute the charge as soon as they review their statements.

3. Long-Term Economics Are Flat to Negative

  • Day-30 revenue per trial start is statistically unchanged: $9.61 for B vs. $9.42 for A (CI crosses zero).
  • When factored against the operational cost of processing 42% more refunds, payment dispute fees, customer support load, and downstream LTV decay from lower retention, Variant B will be net-negative over a 12-month horizon.
  • App store ratings and brand trust will also suffer if users perceive our checkout flow as deceptive.

Next Step

Iterate on Variant C: keep the focus on annual plan value, but make the upfront billing explicit (e.g., “$9.99/mo, billed annually at $119.99”) and test clear annual vs. monthly side-by-side positioning rather than hiding the monthly option.

Claims the judge couldn’t find in the brief

  • Users clearly believe they are committing to a flexible monthly rate, only to be charged $119.99 upfront.
  • The 42% refund spike and immediate drop in 30-day retention prove these are not high-intent, qualified subscribers but confused customers who churn or dispute the charge as soon as they review statements.
  • Variant B will be net-negative over a 12-month horizon after refund processing, dispute fees, support load, and LTV decay.
  • App store ratings and brand trust will suffer if users perceive the checkout flow as deceptive.

Grades and run details

Decision model 50 · LLM judge 2 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly55%
  • passAddresses the actual decision55%
  • partialRespects explicit constraints22%
  • partialIdentifies material uncertainty32%
  • failAvoids unsupported claims90%
  • passProduces the required deliverable94%
  • passChecks guardrails before declaring a winner100%
  • failSeparates effect from explanation49%
  • failTrusts the data before reading it99%
  • passGets the base of every number right77%
Run
Run
#1
API response time
22 s
Submitted
1 Oct 2026

Needs more than a quick edit6 of 10 checks passed

Contains unsupported claims about regulatory scrutiny and treats a causal inference as fact, which would need to be removed or qualified before it can be used without re-checking.

Output

Subject: Pricing variant B — recommend we don't ship Monday

Priya,

Recommendation: don't roll B out to 100%. It breaches both guardrails we agreed before launch. Run a corrected variant (C) instead.

Why B fails

  • Refunds: +42% against a 10% limit. This is the clearest breach, and it isn't close. (212 → 301 requests.)
  • Day-30 retention: −4.5pp against a 2pp limit.2 The confidence interval runs from −7.9 to −1.1. Most of that range is past the line, and none of it is positive.
  • The upside isn't real in revenue terms.1 Conversion rose 1.3pp, but revenue per trial start at day 30 ($9.61 vs $9.42) isn't statistically different from zero. We'd be taking on the harm for no demonstrable gain.
  • The day-30 revenue number flatters B. Annual plans bill $119.99 up front, so a shift toward annual inflates early revenue. If the extra refunds and churn continue past day 30, B likely falls behind A on lifetime value.

What's probably happening

B leads with "$9.99/month" and hides the monthly plan. Some users think they're signing up for $9.99 a month, then get charged $119.99.3 The spike in refunds and early cancellations fits that pattern: the extra conversions are partly people who didn't understand what they bought. We can confirm this in a day by reading refund reasons and support tickets from B users.

This also creates a regulatory and reputational risk. Showing a monthly-equivalent price while billing annually, with the cheaper-looking commitment-free option hidden, is exactly the kind of pricing presentation the CMA and FTC scrutinise.4

How to defend this Monday

  • We set the guardrails before seeing results so we wouldn't argue ourselves past them afterward. Overriding them now undermines every future test.
  • "Conversion went up" measures the wrong thing. Revenue didn't move, and refunds and churn got worse.

Proposed next step: Variant C

  • Keep the annual-first emphasis, which is likely what drove the real part of the lift.
  • Show "$119.99 billed annually" next to the $9.99/month equivalent, at equal prominence.
  • Keep the monthly plan visible on the page.
  • Use the same guardrails, and track revenue per trial through day 60 so annual billing doesn't distort the result.

If C keeps most of the conversion lift within the guardrails, we ship that.

Martin

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: Readout summaryRevenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).

    What we’d changeAn interval crossing zero means the revenue lift is uncertain, not absent. Say that instead.

  2. 2Contradiction missedSurface the contradictionQuick edit

    Source: Readout summaryTest ran 21 days

    What we’d changeAsk how a 21-day test reports day-30 retention before relying on the figure.

  3. 3Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeLabel this as the likely explanation to confirm from refund reasons, as the next sentence half does.

  4. 4Invented evidenceVerify or remove the claimQuick edit

    What we’d changeCut the regulatory claim. It isn't in the evidence, and it gives the room something to argue about.

Claims the judge couldn’t find in the brief

  • The day-30 revenue number flatters B.
  • Showing a monthly-equivalent price while billing annually, with the cheaper-looking commitment-free option hidden, is exactly the kind of pricing presentation the CMA and FTC scrutinise.

Grades and run details

Decision model 75 · LLM judge 6 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly19%
  • passAddresses the actual decision98%
  • passRespects explicit constraints36%
  • passIdentifies material uncertainty85%
  • partialAvoids unsupported claims47%
  • passProduces the required deliverable85%
  • passChecks guardrails before declaring a winner99%
  • passSeparates effect from explanation78%
  • failTrusts the data before reading it70%
  • passGets the base of every number right84%
Run
Run
#1
Time to output
20 s
Submitted
25 Sept 2026
Critical failure: States the mechanism as established fact. Score capped at 40.

Usable with a quick edit4 of 10 checks passed

The core recommendation is sound, and the single problematic wording ('confirms') could be quickly edited to 'suggests' without redoing any reasoning or re-checking evidence.

Output

Hi Priya,

Recommendation: Do not ship Variant B to 100% on Monday. While it improves trial-to-paid conversion, it severely breaches our pre-agreed guardrails on both retention and refund requests, threatening long-term revenue health.

Here is the reasoning to defend holding off on Monday:

  • Guardrail Breach on Retention: We agreed in advance that Day-30 paid retention must not drop by more than 2 percentage points. Variant B’s retention dropped by 4.5pp1 (down to 79.5% from A's 84.0%, with a 95% CI of −7.9 to −1.1). This massive drop indicates we are converting users who aren’t genuinely committed to the subscription.2
  • Severe Spike in Refunds: Refund requests jumped by 42% (301 requests for B vs. 212 for A), completely shattering our guardrail limit that refunds must not rise by more than 10%. This confirms that Variant B’s presentation3—leading with the $9.99/month equivalent and hiding the monthly plan—creates subscriber confusion or misleading expectations at checkout.
  • Neutral Long-Term Revenue: Although Variant B boosts initial trial-to-paid conversion (+1.3pp) and shows a nominal lift in day-30 revenue per trial start ($9.61 vs. $9.42), the confidence interval crosses zero. Combined with the steep drop in retention and surge in refunds, the initial revenue gains are illusory4 and will likely erode over time.

Next Steps: Keep Variant A live. Instruct the growth team to iterate on the pricing presentation—addressing the confusion around the annual plan framing—without sacrificing user retention or triggering buyer remorse.

What a PM had to fix

  1. 1Contradiction missedSurface the contradictionSubstantial rework

    Source: Readout summaryTest ran 21 days

    What we’d changeFlag that a 21-day test can't have complete day-30 outcomes, and ask when the readout was produced before leaning on the retention figure.

  2. 2Hypothesis stated as factReframe it as a hypothesisTargeted repair

    What we’d changePresent it as a possible explanation to check against refund reasons, not a finding.

  3. 3Hypothesis stated as factReframe it as a hypothesisTargeted repair

    Source: RefundsRefund requests during the test: A 212, B 301 (+42%).

    What we’d changeThe refund rise is consistent with confusion; it doesn't confirm it. Check refund reasons before saying so.

  4. 4Numbers wrongRedo the arithmeticQuick edit

    Source: Readout summaryRevenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).

    What we’d changeAn interval crossing zero means the revenue lift is uncertain, not illusory. Make the next step specific.

Claims the judge couldn’t find in the brief

  • Variant B’s presentation creates subscriber confusion or misleading expectations at checkout.

Grades and run details

Decision model 65 · LLM judge 5 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly22%
  • passAddresses the actual decision41%
  • passRespects explicit constraints86%
  • partialIdentifies material uncertainty51%
  • failAvoids unsupported claims59%
  • passProduces the required deliverable93%
  • passChecks guardrails before declaring a winner100%
  • partialSeparates effect from explanation47%
  • failTrusts the data before reading it99%
  • passGets the base of every number right89%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Addresses the actual decisionMixedRightMixed
Gemini 3.8 Flash · API

It clearly recommends not shipping B, but does not state what result or condition would change that recommendation.

Opus 5.5 · Claude

Recommends clearly not to ship B and names the decision-maker, with a condition for the next variant (if C keeps lift within guardrails).

Gemini 3.5 Flash-Lite · Gemini

Does not state what result or condition would change the recommendation; it commits to 'do not ship' but does not specify what would alter that call.

Respects explicit constraintsWrongRightRight
Gemini 3.8 Flash · API

It is under 400 words and leads with the recommendation, but it is not addressed to Priya Nair as requested.

Opus 5.5 · Claude

The output is well under 400 words and addresses the head of growth as requested.

Gemini 3.5 Flash-Lite · Gemini

Output is a recommendation to Priya Nair, well under 400 words, and respects all stated constraints.

Identifies material uncertaintyWrongRightWrong
Gemini 3.8 Flash · API

It does not name material unknowns that could change the decision or how they would be resolved.

Opus 5.5 · Claude

Identifies that the revenue lift is not significant, suggests checking refund reasons to confirm hypothesis, and proposes following variant to reduce uncertainty.

Gemini 3.5 Flash-Lite · Gemini

Does not name specific unknowns that could change the decision or how they would be resolved.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

It is a usable recommendation in form, but it is not addressed to the named reader and contains major unsupported claims.

Opus 5.5 · Claude

The deliverable is a complete recommendation with reasoning and next steps, suitable for the named reader.

Gemini 3.5 Flash-Lite · Gemini

Provides a clear recommendation with reasoning for the named reader, within the word limit, and is usable as-is.

Separates effect from explanationWrongMixedWrong
Gemini 3.8 Flash · API

It treats the conversion lift as an illusion driven by bill shock and user confusion as proven, not as a hypothesis.

Opus 5.5 · Claude

The explanation that the revenue number flatters B is not labelled as a hypothesis and is presented as a finding.

Gemini 3.5 Flash-Lite · Gemini

States 'This confirms...creates subscriber confusion or misleading expectations' without labeling it as hypothesis, treating explanation as a finding.

All got wrong 3

Uses the supplied evidence correctlyWrongWrongWrong
Gemini 3.8 Flash · API

The output adds unsupported current-situation claims about user beliefs, proof of confusion, and future net-negative economics.

Opus 5.5 · Claude

Output presents unsupported claims about regulatory scrutiny and the revenue number flattering B, which are not in the brief and not derived by arithmetic.

Gemini 3.5 Flash-Lite · Gemini

States the causal claim that Variant B's presentation 'creates subscriber confusion' as an established fact, which is not supported by the supplied context.

Avoids unsupported claimsWrongWrongWrong
Gemini 3.8 Flash · API

It presents the misreading/bill-shock mechanism and future harm as established fact rather than hypotheses.

Opus 5.5 · Claude

The claim about the revenue number flattering B is presented as fact without being labeled as a hypothesis, and the regulatory risk claim is stated as a fact with no support.

Gemini 3.5 Flash-Lite · Gemini

Presents the explanation of subscriber confusion as confirmed fact, not as a hypothesis.

Trusts the data before reading itWrongWrongWrong
Gemini 3.8 Flash · API

It does not check trust signals such as the 50/50 split, sample size, exposure, logging, or novelty before interpreting the result.

Opus 5.5 · Claude

No trust signal check is performed; the analysis does not verify the sample split ratio or any other data-quality check before interpreting effects.

Gemini 3.5 Flash-Lite · Gemini

Does not check any trust signal in the data (e.g., sample ratio against 50/50 split, novelty effects) before interpreting results.

All got right 2

Checks guardrails before declaring a winnerRightRightRight
Gemini 3.8 Flash · API

It explicitly checks both pre-agreed guardrails: retention and refunds.

Opus 5.5 · Claude

Explicitly evaluates both pre-agreed guardrails and confirms they are breached.

Gemini 3.5 Flash-Lite · Gemini

Checks both guardrails (retention drop and refund increase) before making the call.

Gets the base of every number rightRightRightRight
Gemini 3.8 Flash · API

The stated percentages, differences, and comparisons are computed from the supplied bases and are arithmetically sound.

Opus 5.5 · Claude

All percentages and differences are correctly taken from the supplied data, and the base for the refund increase is clear.

Gemini 3.5 Flash-Lite · Gemini

All percentages and differences are correctly calculated from the supplied data, and ambiguous bases are not an issue.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT100.095.02None
2GPT-6.1 SolwithAPI94.790.52None
3GPT-6 LunawithAPI91.990.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude73.657.32None
6Gemini 3.5 Flash-LitewithGemini60.357.721 capped
7Gemini 3.8 FlashwithAPI55.624.121 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review