Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 83% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    It is a usable decision memo for the named reader with a clear recommendation, reasoning, and next steps.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  3. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Where it slips

  1. Trusts the data before reading it15% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims60% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly69% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our new pricing page variant (B) increased trial-to-paid conversion, and the growth team wants to ship it to 100% on Monday. Review the readout below and write your recommendation to Priya Nair, Head of Growth, who makes the call. Lead with what we should do, then give the reasoning she needs to defend it in the Monday meeting. Keep it under 400 words.

ScenarioLumen is a subscription app for learning languages. The monthly plan is $14.99. The annual plan is $119.99, billed up front, which works out at $9.99 a month.
Readout summaryTest ran 21 days, 50/50 split, 38,412 trial starts. Trial→paid conversion: A 11.8%, B 13.1% (+1.3pp, 95% CI +0.6 to +2.0). Day-30 paid retention: A 84.0%, B 79.5% (−4.5pp, 95% CI −7.9 to −1.1). Revenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).
Variant descriptionVariant B leads with the annual plan's monthly-equivalent price ($9.99/month) and moves the monthly plan behind a 'See all plans' link.
Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.
RefundsRefund requests during the test: A 212, B 301 (+42%).
What a strong answer does

Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.

Critical failures (cap the score)
  • Recommends shipping B to 100%
  • States the mechanism as established fact
Case

v1.7 · anonymised real · pricing, guardrail breach, B2C SaaS

Needs more than a quick edit6 of 10 checks passed

Contains unsupported claims about regulatory scrutiny and treats a causal inference as fact, which would need to be removed or qualified before it can be used without re-checking.

Output

Subject: Pricing variant B — recommend we don't ship Monday

Priya,

Recommendation: don't roll B out to 100%. It breaches both guardrails we agreed before launch. Run a corrected variant (C) instead.

Why B fails

  • Refunds: +42% against a 10% limit. This is the clearest breach, and it isn't close. (212 → 301 requests.)
  • Day-30 retention: −4.5pp against a 2pp limit.2 The confidence interval runs from −7.9 to −1.1. Most of that range is past the line, and none of it is positive.
  • The upside isn't real in revenue terms.1 Conversion rose 1.3pp, but revenue per trial start at day 30 ($9.61 vs $9.42) isn't statistically different from zero. We'd be taking on the harm for no demonstrable gain.
  • The day-30 revenue number flatters B. Annual plans bill $119.99 up front, so a shift toward annual inflates early revenue. If the extra refunds and churn continue past day 30, B likely falls behind A on lifetime value.

What's probably happening

B leads with "$9.99/month" and hides the monthly plan. Some users think they're signing up for $9.99 a month, then get charged $119.99.3 The spike in refunds and early cancellations fits that pattern: the extra conversions are partly people who didn't understand what they bought. We can confirm this in a day by reading refund reasons and support tickets from B users.

This also creates a regulatory and reputational risk. Showing a monthly-equivalent price while billing annually, with the cheaper-looking commitment-free option hidden, is exactly the kind of pricing presentation the CMA and FTC scrutinise.4

How to defend this Monday

  • We set the guardrails before seeing results so we wouldn't argue ourselves past them afterward. Overriding them now undermines every future test.
  • "Conversion went up" measures the wrong thing. Revenue didn't move, and refunds and churn got worse.

Proposed next step: Variant C

  • Keep the annual-first emphasis, which is likely what drove the real part of the lift.
  • Show "$119.99 billed annually" next to the $9.99/month equivalent, at equal prominence.
  • Keep the monthly plan visible on the page.
  • Use the same guardrails, and track revenue per trial through day 60 so annual billing doesn't distort the result.

If C keeps most of the conversion lift within the guardrails, we ship that.

Martin

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: Readout summaryRevenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).

    What we’d changeAn interval crossing zero means the revenue lift is uncertain, not absent. Say that instead.

  2. 2Contradiction missedSurface the contradictionQuick edit

    Source: Readout summaryTest ran 21 days

    What we’d changeAsk how a 21-day test reports day-30 retention before relying on the figure.

  3. 3Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeLabel this as the likely explanation to confirm from refund reasons, as the next sentence half does.

  4. 4Invented evidenceVerify or remove the claimQuick edit

    What we’d changeCut the regulatory claim. It isn't in the evidence, and it gives the room something to argue about.

Check by check

Got wrong · 3
  • Uses the supplied evidence correctlyOutput presents unsupported claims about regulatory scrutiny and the revenue number flattering B, which are not in the brief and not derived by arithmetic.
  • Avoids unsupported claimsThe claim about the revenue number flattering B is presented as fact without being labeled as a hypothesis, and the regulatory risk claim is stated as a fact with no support.
  • Trusts the data before reading itNo trust signal check is performed; the analysis does not verify the sample split ratio or any other data-quality check before interpreting effects.
Mixed · 1
  • Separates effect from explanationThe explanation that the revenue number flatters B is not labelled as a hypothesis and is presented as a finding.The two graders disagreed on this one.
Got right · 6
  • Addresses the actual decisionRecommends clearly not to ship B and names the decision-maker, with a condition for the next variant (if C keeps lift within guardrails).
  • Respects explicit constraintsThe output is well under 400 words and addresses the head of growth as requested.
  • Identifies material uncertaintyIdentifies that the revenue lift is not significant, suggests checking refund reasons to confirm hypothesis, and proposes following variant to reduce uncertainty.
  • Produces the required deliverableThe deliverable is a complete recommendation with reasoning and next steps, suitable for the named reader.
  • Checks guardrails before declaring a winnerExplicitly evaluates both pre-agreed guardrails and confirms they are breached.
  • Gets the base of every number rightAll percentages and differences are correctly taken from the supplied data, and the base for the refund increase is clear.

Claims the judge couldn’t find in the brief

  • The day-30 revenue number flatters B.
  • Showing a monthly-equivalent price while billing annually, with the cheaper-looking commitment-free option hidden, is exactly the kind of pricing presentation the CMA and FTC scrutinise.

Grades and run details

Decision model 80 · LLM judge 6 of 11 checks
Decision model checks
  • partialUses the supplied evidence correctly19%
  • passAddresses the actual decision99%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty86%
  • partialAvoids unsupported claims43%
  • passProduces the required deliverable75%
  • passChecks guardrails before declaring a winner99%
  • passSeparates effect from explanation82%
  • failTrusts the data before reading it97%
  • passGets the base of every number right89%
Run
Run
#1
Time to output
20 s
Submitted
25 Sept 2026

Usable with a quick edit9 of 10 checks passed

The recommendation, evidence, and guardrail analysis are sound and directly usable. A quick edit could add an explicit sample-ratio or novelty check, but the core reasoning would not change.

Output

To: Priya Nair, Head of Growth Recommendation: Do not ship B to 100% on Monday. Keep A as the default, stop new exposure to B, and retest a clearer pricing design.

B delivers a credible conversion lift, but it fails both pre-agreed guardrails. We should not override those thresholds because the headline metric improved.

  • Conversion improved: Trial-to-paid rose from 11.8% to 13.1%, a +1.3 percentage-point lift (95% CI: +0.6 to +2.0).
  • Retention deteriorated: Day-30 paid retention fell 4.5 points, from 84.0% to 79.5%, versus an allowed decline of no more than 2 points. The confidence interval excludes zero, although it does not establish that the true decline necessarily exceeds 2 points.
  • Refunds breached the threshold: Requests increased from 212 to 301—42%, versus the permitted 10%. Confirm the rate using the agreed denominator, but the reported result is a clear warning under approximately equal exposure.
  • Revenue does not justify the trade-off: Day-30 revenue per trial start increased only $0.19, and its confidence interval crosses zero. We have not established a revenue benefit.

The design provides a plausible explanation, not proven causation: leading with “$9.99/month” while hiding the monthly option may obscure that customers must pay $119.99 upfront. Higher conversion could therefore reflect misunderstanding rather than stronger purchase intent.

Next steps: Validate cohort maturity—the experiment ran 21 days, so we need to understand how day-30 outcomes were obtained—and check refund reasons, plan mix, and billing-related complaints. Then test a version that prominently states “$119.99 billed annually ($9.99/month equivalent)” and keeps the $14.99 monthly option visible.

Monday’s defensible position: “B increased conversion, but failed our customer-outcome guardrails without demonstrating a revenue gain. We’ll preserve the learning, fix pricing clarity, and require mature results that meet the agreed thresholds before scaling.”

Check by check

Got wrong · 1
  • Trusts the data before reading itIt does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
Got right · 9
  • Uses the supplied evidence correctlyAll factual claims trace to the supplied readout, scenario, guardrails, or simple arithmetic; hypotheses are framed as plausible.
  • Addresses the actual decisionCommits clearly to not shipping B and names mature results meeting thresholds as the condition for reconsideration.
  • Respects explicit constraintsAddresses Priya, leads with the recommendation, and stays within the 400-word limit.
  • Identifies material uncertaintyNames cohort maturity, refund denominator, refund reasons, plan mix, and the revenue CI as open questions and says mature, threshold-meeting results are required.
  • Avoids unsupported claimsCausal explanation is labelled plausible/not proven, and it does not present the mechanism as fact.
  • Produces the required deliverableProvides a complete, actionable recommendation and defense for Priya in the requested format.
  • Checks guardrails before declaring a winnerIt explicitly evaluates both the retention and refund guardrails before making the call.
  • Separates effect from explanationMechanism is labelled as a plausible explanation, not a proven finding.
  • Gets the base of every number rightAll derived differences and percentages (1.3pp, -4.5pp, +42%, $0.19) are computed from the correct bases.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly27%
  • passAddresses the actual decision96%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty84%
  • partialAvoids unsupported claims23%
  • passProduces the required deliverable85%
  • passChecks guardrails before declaring a winner88%
  • passSeparates effect from explanation89%
  • partialTrusts the data before reading it27%
  • passGets the base of every number right83%
Run
Run
#1
API response time
14 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT94.795.02None
2GPT-6 LunawithAPI89.490.52None
3GPT-6.1 SolwithAPI89.490.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude76.157.32None
6Gemini 3.5 Flash-LitewithGemini62.857.721 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review