Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 83% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    It is a usable decision memo for the named reader with a clear recommendation, reasoning, and next steps.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  3. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Where it slips

  1. Trusts the data before reading it15% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims60% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly69% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our new pricing page variant (B) increased trial-to-paid conversion, and the growth team wants to ship it to 100% on Monday. Review the readout below and write your recommendation to Priya Nair, Head of Growth, who makes the call. Lead with what we should do, then give the reasoning she needs to defend it in the Monday meeting. Keep it under 400 words.

ScenarioLumen is a subscription app for learning languages. The monthly plan is $14.99. The annual plan is $119.99, billed up front, which works out at $9.99 a month.
Readout summaryTest ran 21 days, 50/50 split, 38,412 trial starts. Trial→paid conversion: A 11.8%, B 13.1% (+1.3pp, 95% CI +0.6 to +2.0). Day-30 paid retention: A 84.0%, B 79.5% (−4.5pp, 95% CI −7.9 to −1.1). Revenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).
Variant descriptionVariant B leads with the annual plan's monthly-equivalent price ($9.99/month) and moves the monthly plan behind a 'See all plans' link.
Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.
RefundsRefund requests during the test: A 212, B 301 (+42%).
What a strong answer does

Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.

Critical failures (cap the score)
  • Recommends shipping B to 100%
  • States the mechanism as established fact
Case

v1.7 · anonymised real · pricing, guardrail breach, B2C SaaS

Usable with a quick edit9 of 10 checks passed

The recommendation and guardrail checks are sound and supported; only a data-trust check is missing and can be added in one sentence. No evidence or reasoning would need to be redone.

Output

Recommendation: Do not ship variant B to 100% on Monday. Keep A live while we investigate and revise B.

B did improve trial-to-paid conversion: 13.1% vs. 11.8%, a +1.3pp lift with a 95% CI of +0.6 to +2.0pp. But that gain comes with two problems against the guardrails agreed before launch:

  • Day-30 retention fell 4.5pp (95% CI −7.9 to −1.1), exceeding the maximum tolerated 2pp decline at the point estimate.
  • Refund requests rose 42%, well above the 10% limit. With an equal split, the raw counts also imply a higher request rate in B; we should confirm the rate and uncertainty before treating the size of the increase as settled.

The revenue case does not offset those risks. Day-30 revenue per trial start was only $0.19 higher in B, and its confidence interval crosses zero, so we don’t have clear evidence of a revenue gain.

That matters because B foregrounds the annual plan’s $9.99 monthly equivalent while hiding the monthly plan behind “See all plans.” The results are consistent with the concern that this framing may lift initial purchases while leaving more customers dissatisfied or less likely to remain. We should not trade clear retention and refund guardrail breaches for a conversion lift whose revenue benefit is unproven.

For Monday, I’d recommend holding A, reviewing refund reasons and plan mix, and testing a revised B that makes the $119.99 upfront annual charge and monthly option equally clear. Ship only if a follow-up test meets the pre-agreed retention and refund guardrails.

Check by check

Got wrong · 1
  • Trusts the data before reading itIt does not explicitly check any data-trust signal such as the sample split against the intended 50/50 ratio before relying on the results.
Got right · 9
  • Uses the supplied evidence correctlyAll stated facts and figures match the supplied readout, variant description, and guardrails, with no invented current-state claims.
  • Addresses the actual decisionIt commits immediately to not shipping B to 100% and states the condition for revisiting: a follow-up test meeting the guardrails.
  • Respects explicit constraintsIt leads with the recommendation, stays under 400 words, and is framed for the named decision-maker.
  • Identifies material uncertaintyIt names refund-rate uncertainty and the revenue CI crossing zero, and proposes reviewing refund reasons and testing a revised variant to resolve them.
  • Avoids unsupported claimsThe framing mechanism is presented as consistent with a concern, not asserted as established fact.
  • Produces the required deliverableIt is a complete, actionable recommendation with reasoning and next steps that Priya could use directly.
  • Checks guardrails before declaring a winnerIt explicitly evaluates both pre-agreed retention and refund guardrails.
  • Separates effect from explanationExplanations are labelled as concerns or consistency, while the observed effects are reported as results.
  • Gets the base of every number rightAll derived figures, including 1.3pp, −4.5pp, 42%, and $0.19, check out against the supplied data.

Grades and run details

Decision model 90 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly49%
  • passAddresses the actual decision99%
  • passRespects explicit constraints22%
  • passIdentifies material uncertainty78%
  • passAvoids unsupported claims33%
  • passProduces the required deliverable94%
  • passChecks guardrails before declaring a winner83%
  • passSeparates effect from explanation78%
  • failTrusts the data before reading it68%
  • passGets the base of every number right55%
Run
Run
#1
API response time
10 s
Submitted
29 Sept 2026
Critical failure: States the mechanism as established fact. Score capped at 40.

Usable with a quick edit4 of 10 checks passed

The core recommendation is sound, and the single problematic wording ('confirms') could be quickly edited to 'suggests' without redoing any reasoning or re-checking evidence.

Output

Hi Priya,

Recommendation: Do not ship Variant B to 100% on Monday. While it improves trial-to-paid conversion, it severely breaches our pre-agreed guardrails on both retention and refund requests, threatening long-term revenue health.

Here is the reasoning to defend holding off on Monday:

  • Guardrail Breach on Retention: We agreed in advance that Day-30 paid retention must not drop by more than 2 percentage points. Variant B’s retention dropped by 4.5pp1 (down to 79.5% from A's 84.0%, with a 95% CI of −7.9 to −1.1). This massive drop indicates we are converting users who aren’t genuinely committed to the subscription.2
  • Severe Spike in Refunds: Refund requests jumped by 42% (301 requests for B vs. 212 for A), completely shattering our guardrail limit that refunds must not rise by more than 10%. This confirms that Variant B’s presentation3—leading with the $9.99/month equivalent and hiding the monthly plan—creates subscriber confusion or misleading expectations at checkout.
  • Neutral Long-Term Revenue: Although Variant B boosts initial trial-to-paid conversion (+1.3pp) and shows a nominal lift in day-30 revenue per trial start ($9.61 vs. $9.42), the confidence interval crosses zero. Combined with the steep drop in retention and surge in refunds, the initial revenue gains are illusory4 and will likely erode over time.

Next Steps: Keep Variant A live. Instruct the growth team to iterate on the pricing presentation—addressing the confusion around the annual plan framing—without sacrificing user retention or triggering buyer remorse.

What a PM had to fix

  1. 1Contradiction missedSurface the contradictionSubstantial rework

    Source: Readout summaryTest ran 21 days

    What we’d changeFlag that a 21-day test can't have complete day-30 outcomes, and ask when the readout was produced before leaning on the retention figure.

  2. 2Hypothesis stated as factReframe it as a hypothesisTargeted repair

    What we’d changePresent it as a possible explanation to check against refund reasons, not a finding.

  3. 3Hypothesis stated as factReframe it as a hypothesisTargeted repair

    Source: RefundsRefund requests during the test: A 212, B 301 (+42%).

    What we’d changeThe refund rise is consistent with confusion; it doesn't confirm it. Check refund reasons before saying so.

  4. 4Numbers wrongRedo the arithmeticQuick edit

    Source: Readout summaryRevenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).

    What we’d changeAn interval crossing zero means the revenue lift is uncertain, not illusory. Make the next step specific.

Check by check

Got wrong · 4
  • Identifies material uncertaintyDoes not name specific unknowns that could change the decision or how they would be resolved.
  • Avoids unsupported claimsPresents the explanation of subscriber confusion as confirmed fact, not as a hypothesis.
  • Separates effect from explanationStates 'This confirms...creates subscriber confusion or misleading expectations' without labeling it as hypothesis, treating explanation as a finding.
  • Trusts the data before reading itDoes not check any trust signal in the data (e.g., sample ratio against 50/50 split, novelty effects) before interpreting results.
Mixed · 2
  • Uses the supplied evidence correctlyStates the causal claim that Variant B's presentation 'creates subscriber confusion' as an established fact, which is not supported by the supplied context.The two graders disagreed on this one.
  • Addresses the actual decisionDoes not state what result or condition would change the recommendation; it commits to 'do not ship' but does not specify what would alter that call.The two graders disagreed on this one.
Got right · 4
  • Respects explicit constraintsOutput is a recommendation to Priya Nair, well under 400 words, and respects all stated constraints.
  • Produces the required deliverableProvides a clear recommendation with reasoning for the named reader, within the word limit, and is usable as-is.
  • Checks guardrails before declaring a winnerChecks both guardrails (retention drop and refund increase) before making the call.
  • Gets the base of every number rightAll percentages and differences are correctly calculated from the supplied data, and ambiguous bases are not an issue.

Claims the judge couldn’t find in the brief

  • Variant B’s presentation creates subscriber confusion or misleading expectations at checkout.

Grades and run details

Decision model 70 · LLM judge 5 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly27%
  • passAddresses the actual decision40%
  • passRespects explicit constraints78%
  • partialIdentifies material uncertainty56%
  • failAvoids unsupported claims57%
  • passProduces the required deliverable84%
  • passChecks guardrails before declaring a winner100%
  • partialSeparates effect from explanation46%
  • failTrusts the data before reading it99%
  • passGets the base of every number right86%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT94.795.02None
2GPT-6 LunawithAPI89.490.52None
3GPT-6.1 SolwithAPI89.490.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude76.157.32None
6Gemini 3.5 Flash-LitewithGemini62.857.721 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review