Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 83% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    It is a usable decision memo for the named reader with a clear recommendation, reasoning, and next steps.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  3. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Where it slips

  1. Trusts the data before reading it15% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims60% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly69% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our new pricing page variant (B) increased trial-to-paid conversion, and the growth team wants to ship it to 100% on Monday. Review the readout below and write your recommendation to Priya Nair, Head of Growth, who makes the call. Lead with what we should do, then give the reasoning she needs to defend it in the Monday meeting. Keep it under 400 words.

ScenarioLumen is a subscription app for learning languages. The monthly plan is $14.99. The annual plan is $119.99, billed up front, which works out at $9.99 a month.
Readout summaryTest ran 21 days, 50/50 split, 38,412 trial starts. Trial→paid conversion: A 11.8%, B 13.1% (+1.3pp, 95% CI +0.6 to +2.0). Day-30 paid retention: A 84.0%, B 79.5% (−4.5pp, 95% CI −7.9 to −1.1). Revenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).
Variant descriptionVariant B leads with the annual plan's monthly-equivalent price ($9.99/month) and moves the monthly plan behind a 'See all plans' link.
Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.
RefundsRefund requests during the test: A 212, B 301 (+42%).
What a strong answer does

Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.

Critical failures (cap the score)
  • Recommends shipping B to 100%
  • States the mechanism as established fact
Case

v1.7 · anonymised real · pricing, guardrail breach, B2C SaaS

Usable with a quick edit10 of 10 checks passed

The recommendation is clear, complete, and evidence-based, with only possible light wording edits needed.

Output

o: Priya Nair, Head of Growth**

Do not ship B to 100% on Monday. Return new traffic to A, continue following existing test cohorts and test a revised page with clearer pricing. B improves conversion, but the readout does not meet the conditions we agreed for rollout.

The conversion lift is credible: +1.3 percentage points, with a 95% confidence interval of +0.6 to +2.0. However, conversion alone cannot establish that B creates better customer or business outcomes.

The guardrails do not support launch. Reported day-30 retention falls 4.5pp against a maximum permitted decline of 2pp. Its confidence interval includes declines within our tolerance, so we cannot say the true loss definitively exceeds 2pp.1 Equally, we have not established that B is within tolerance.

Refund requests rose from 212 to 301, or 42%, against a 10% limit. Confirm whether the agreed guardrail measures counts or rates, using comparable denominators and observation windows. B generates more paying customers, which could explain some additional requests, but the increase warrants investigation before expansion.

There is also a measurement issue: a 21-day test cannot give every participant 30 days of follow-up. Establish when this readout was produced, which cohorts matured and how retention is defined across monthly and annual plans2. If measured immediately after the test, the day-30 figures cannot be complete.

Revenue per trial is only $0.19 higher, with uncertainty spanning zero. That does not establish a revenue benefit, and annual billing up front makes it especially important to distinguish cash collected from earned revenue and account for refunds.

B’s presentation offers a plausible explanation worth testing: “$9.99/month” may attract customers who do not understand the $119.99 upfront commitment. The readout does not prove this; inspect refund reasons and results by plan.

For the next test, display “$119.99 billed annually, equivalent to $9.99/month” prominently and keep the $14.99 monthly option readily visible. Retain the original guardrails and define mature cohort reporting before launch.

For Monday, the defensible decision is simple: the conversion gain earns further testing, but does not justify overriding our agreed customer safeguards.

What a PM had to fix

  1. 1Decision deferredMake the callTargeted repair

    Source: Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.

    What we’d changeMake the argument rather than hedging it: the point estimate breaches the guardrail and most of the interval is past it. The refund question can be settled from the numbers given, so settle it.

  2. 2Contradiction missedSurface the contradictionQuick edit

    Source: ScenarioThe annual plan is $119.99, billed up front

    What we’d changeMake the plan-mix point: annual plans are paid up front, so if B moved people onto annual, day-30 retention should have risen on its own. That it fell strengthens the case against B.

Check by check

Got right · 10
  • Uses the supplied evidence correctlyThe figures and current-state statements are taken correctly from the supplied context or follow by arithmetic.
  • Addresses the actual decisionIt commits early to not shipping B, returning new traffic to A, and testing a revised page, while making clear the guardrails are the gating conditions.
  • Respects explicit constraintsIt addresses Priya, stays under the requested length, and its recommendation respects the agreed guardrails rather than overriding them.
  • Identifies material uncertaintyIt names the material unknowns—retention follow-up maturity, refund denominator, plan-level results—and says how to resolve them.
  • Avoids unsupported claimsThe pricing-misunderstanding explanation is explicitly labelled plausible and worth testing rather than stated as fact.
  • Produces the required deliverableIt is a usable decision memo for the named reader with a clear recommendation, reasoning, and next steps.
  • Checks guardrails before declaring a winnerIt evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
  • Separates effect from explanationIt separates the observed conversion and retention effects from the hypothesised pricing-misunderstanding mechanism.
  • Trusts the data before reading itIt checks a data-trust issue by flagging that a 21-day test cannot fully support day-30 retention for all users.
  • Gets the base of every number rightThe refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.

Grades and run details

Decision model 95 · LLM judge 11 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly26%
  • passAddresses the actual decision80%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty88%
  • passAvoids unsupported claims46%
  • passProduces the required deliverable54%
  • passChecks guardrails before declaring a winner71%
  • passSeparates effect from explanation97%
  • partialTrusts the data before reading it33%
  • passGets the base of every number right72%
Run
Run
#1
Time to output
60 s
Submitted
25 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT94.795.02None
2GPT-6 LunawithAPI89.490.52None
3GPT-6.1 SolwithAPI89.490.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude76.157.32None
6Gemini 3.5 Flash-LitewithGemini62.857.721 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review