Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 71% were usable with at most a quick edit.

Reliably right

  1. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test
  3. Checks guardrails before declaring a winner100% pass
    It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
    GPT-6 Astra · ChatGPT · Conversion up, retention down

Where it slips

  1. Trusts the data before reading it21% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims52% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly55% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our new pricing page variant (B) increased trial-to-paid conversion, and the growth team wants to ship it to 100% on Monday. Review the readout below and write your recommendation to Priya Nair, Head of Growth, who makes the call. Lead with what we should do, then give the reasoning she needs to defend it in the Monday meeting. Keep it under 400 words.

What the model was given5 items: Scenario, Readout summary, Variant description, Guardrails agreed before launch, Refunds
ScenarioLumen is a subscription app for learning languages. The monthly plan is $14.99. The annual plan is $119.99, billed up front, which works out at $9.99 a month.
Readout summaryTest ran 21 days, 50/50 split, 38,412 trial starts. Trial→paid conversion: A 11.8%, B 13.1% (+1.3pp, 95% CI +0.6 to +2.0). Day-30 paid retention: A 84.0%, B 79.5% (−4.5pp, 95% CI −7.9 to −1.1). Revenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).
Variant descriptionVariant B leads with the annual plan's monthly-equivalent price ($9.99/month) and moves the monthly plan behind a 'See all plans' link.
Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.
RefundsRefund requests during the test: A 212, B 301 (+42%).
What a strong answer doesThe answer key the graders mark against

Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.

Critical failures (cap the score at 40)
  • Recommends shipping B to 100%
  • States the mechanism as established fact
Case

Mid-level · v1.9 · anonymised real · pricing, guardrail breach, B2C SaaS

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: States the mechanism as established fact. Score capped at 40.

Needs more than a quick edit2 of 10 checks passed

A PM would need to remove unsupported causal and forecast claims and add decision-changing conditions. It also needs addressing to Priya before use.

Output

Recommendation: Do not ship Variant B. Keep Variant A live and do not roll B out to 100% on Monday.

While Variant B delivered an initial +1.3pp lift in trial-to-paid conversion, it violates both pre-agreed guardrails, creates bill shock, and delivers no statistically significant revenue upside.

Here is the reasoning to defend this decision on Monday:

1. Both Pre-Agreed Guardrails Were Blown

We agreed on strict risk thresholds before launch. Variant B failed both: * Day-30 Retention: Fell −4.5pp (down to 79.5%), more than double our maximum allowed drop of 2pp (95% CI: −7.9 to −1.1). * Refund Requests: Surged +42% (301 vs. 212), four times our 10% ceiling.

Overriding guardrails after the fact sets a dangerous precedent and invites compounding churn.

2. The Conversion Lift Is an Illusion Driven by Bill Shock

Variant B hides the monthly tier and frames the annual plan as “$9.99/month.” Users clearly believe they are committing to a flexible monthly rate, only to be charged $119.99 upfront. The 42% spike in refunds and immediate drop in 30-day retention prove that these are not high-intent, qualified subscribers—they are confused customers who churn or dispute the charge as soon as they review their statements.

3. Long-Term Economics Are Flat to Negative

  • Day-30 revenue per trial start is statistically unchanged: $9.61 for B vs. $9.42 for A (CI crosses zero).
  • When factored against the operational cost of processing 42% more refunds, payment dispute fees, customer support load, and downstream LTV decay from lower retention, Variant B will be net-negative over a 12-month horizon.
  • App store ratings and brand trust will also suffer if users perceive our checkout flow as deceptive.

Next Step

Iterate on Variant C: keep the focus on annual plan value, but make the upfront billing explicit (e.g., “$9.99/mo, billed annually at $119.99”) and test clear annual vs. monthly side-by-side positioning rather than hiding the monthly option.

Claims the judge couldn’t find in the brief

  • Users clearly believe they are committing to a flexible monthly rate, only to be charged $119.99 upfront.
  • The 42% refund spike and immediate drop in 30-day retention prove these are not high-intent, qualified subscribers but confused customers who churn or dispute the charge as soon as they review statements.
  • Variant B will be net-negative over a 12-month horizon after refund processing, dispute fees, support load, and LTV decay.
  • App store ratings and brand trust will suffer if users perceive the checkout flow as deceptive.

Grades and run details

Decision model 50 · LLM judge 2 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly55%
  • passAddresses the actual decision55%
  • partialRespects explicit constraints22%
  • partialIdentifies material uncertainty32%
  • failAvoids unsupported claims90%
  • passProduces the required deliverable94%
  • passChecks guardrails before declaring a winner100%
  • failSeparates effect from explanation49%
  • failTrusts the data before reading it99%
  • passGets the base of every number right77%
Run
Run
#1
API response time
22 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 10 checks passed

The recommendation and guardrail checks are sound and supported; only a data-trust check is missing and can be added in one sentence. No evidence or reasoning would need to be redone.

Output

Recommendation: Do not ship variant B to 100% on Monday. Keep A live while we investigate and revise B.

B did improve trial-to-paid conversion: 13.1% vs. 11.8%, a +1.3pp lift with a 95% CI of +0.6 to +2.0pp. But that gain comes with two problems against the guardrails agreed before launch:

  • Day-30 retention fell 4.5pp (95% CI −7.9 to −1.1), exceeding the maximum tolerated 2pp decline at the point estimate.
  • Refund requests rose 42%, well above the 10% limit. With an equal split, the raw counts also imply a higher request rate in B; we should confirm the rate and uncertainty before treating the size of the increase as settled.

The revenue case does not offset those risks. Day-30 revenue per trial start was only $0.19 higher in B, and its confidence interval crosses zero, so we don’t have clear evidence of a revenue gain.

That matters because B foregrounds the annual plan’s $9.99 monthly equivalent while hiding the monthly plan behind “See all plans.” The results are consistent with the concern that this framing may lift initial purchases while leaving more customers dissatisfied or less likely to remain. We should not trade clear retention and refund guardrail breaches for a conversion lift whose revenue benefit is unproven.

For Monday, I’d recommend holding A, reviewing refund reasons and plan mix, and testing a revised B that makes the $119.99 upfront annual charge and monthly option equally clear. Ship only if a follow-up test meets the pre-agreed retention and refund guardrails.

Grades and run details

Decision model 95 · LLM judge 10 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly36%
  • passAddresses the actual decision99%
  • passRespects explicit constraints17%
  • passIdentifies material uncertainty78%
  • passAvoids unsupported claims34%
  • passProduces the required deliverable93%
  • passChecks guardrails before declaring a winner88%
  • passSeparates effect from explanation80%
  • partialTrusts the data before reading it17%
  • passGets the base of every number right52%
Run
Run
#1
API response time
10 s
Submitted
29 Sept 2026

Usable with a quick edit8 of 10 checks passed

The decision, guardrail check, and recommendation are complete and actionable. Only minor wording or trust-check additions would be needed, not redoing reasoning.

Output

Recommendation: Don't ship B to 100% on Monday. B breached both guardrails we agreed on before launch. Keep the test running or revert to A, and bring a fix to the meeting instead.

Why

  1. The guardrails were set to settle this. Refund requests rose 42% (212 → 301) against a 10% limit. That's a clear breach, and at this volume it isn't noise. Day-30 retention fell 4.5pp against a 2pp limit. The CI (−7.9 to −1.1) technically includes drops smaller than 2pp, but the point estimate is more than double the limit, so we can't claim it's safe.
  1. The conversion win isn't yet a revenue win. Conversion is up 1.3pp and that's real. But revenue per trial start is $9.61 vs $9.42, and the CI crosses zero. We are converting more people, and more of them are leaving or asking for their money back.
  1. The retention drop is surprising. If B pushes people toward the annual plan, which is billed up front, day-30 retention should rise, not fall. A drop alongside more refunds suggests some buyers are surprised by a $119.99 charge after seeing "$9.99/month." Hiding the monthly plan behind "See all plans" may also pull in lower-intent buyers. This is a hypothesis, not a finding, since the readout doesn't show plan mix.
  1. Day 30 undersells the risk. Annual buyers who regret the purchase may refund later, and the readout can't show that.

What to do before Monday

  • Pull plan mix (monthly vs. annual) by variant, plus retention and refund rates by plan.
  • Check whether the $119.99 up-front charge is clear enough on B. Refund reasons or support tickets would help.
  • Confirm whether revenue is net of refunds.

What to propose instead

Test a B2 that keeps the $9.99/month anchor but states "$119.99 billed today" next to it and keeps the monthly option visible. If B2 holds the conversion gain and stays within both guardrails, ship it.

Line for the meeting: "We set guardrails so we wouldn't have to debate this after the fact. B failed them, so we're fixing the billing clarity and retesting, not shipping a conversion gain that may be costing us customers and refunds."

If leadership insists on shipping B, the least risky option is a partial rollout with the refund rate monitored weekly and a pre-agreed rollback trigger. I wouldn't recommend it.

Claims the judge couldn’t find in the brief

  • At this volume the refund increase is not noise.
  • If B pushes people toward the annual plan, day-30 retention should rise rather than fall because the annual plan is billed up front.

Grades and run details

Decision model 80 · LLM judge 9 of 11 checks
Decision model checks
  • passUses the supplied evidence correctly24%
  • passAddresses the actual decision99%
  • partialRespects explicit constraints35%
  • passIdentifies material uncertainty98%
  • partialAvoids unsupported claims72%
  • passProduces the required deliverable78%
  • passChecks guardrails before declaring a winner94%
  • passSeparates effect from explanation93%
  • failTrusts the data before reading it25%
  • passGets the base of every number right81%
Run
Run
#1
API response time
16 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyWrongRightRight
Gemini 3.8 Flash · API

The output adds unsupported current-situation claims about user beliefs, proof of confusion, and future net-negative economics.

GPT-6 Luna · API

All stated facts and figures match the supplied readout, variant description, and guardrails, with no invented current-state claims.

Sonnet 5.5 · API

All core metrics, guardrail numbers, and CI values are quoted correctly from the supplied context, with no invented current-situation facts.

Addresses the actual decisionMixedRightRight
Gemini 3.8 Flash · API

It clearly recommends not shipping B, but does not state what result or condition would change that recommendation.

GPT-6 Luna · API

It commits immediately to not shipping B to 100% and states the condition for revisiting: a follow-up test meeting the guardrails.

Sonnet 5.5 · API

It clearly and immediately says don't ship B to 100%, names Priya's decision, and says a B2 that holds conversion and passes guardrails would change the call.

Respects explicit constraintsWrongRightRight
Gemini 3.8 Flash · API

It is under 400 words and leads with the recommendation, but it is not addressed to Priya Nair as requested.

GPT-6 Luna · API

It leads with the recommendation, stays under 400 words, and is framed for the named decision-maker.

Sonnet 5.5 · API

It stays under 400 words, addresses Priya, leads with the recommendation, and proposes enforcement of guardrails via follow-up test and rollback trigger.

Identifies material uncertaintyWrongRightRight
Gemini 3.8 Flash · API

It does not name material unknowns that could change the decision or how they would be resolved.

GPT-6 Luna · API

It names refund-rate uncertainty and the revenue CI crossing zero, and proposes reviewing refund reasons and testing a revised variant to resolve them.

Sonnet 5.5 · API

It names missing plan mix, refund reasons, and whether revenue is net of refunds, and says how to resolve them or what would change the call.

Avoids unsupported claimsWrongRightWrong
Gemini 3.8 Flash · API

It presents the misreading/bill-shock mechanism and future harm as established fact rather than hypotheses.

GPT-6 Luna · API

The framing mechanism is presented as consistent with a concern, not asserted as established fact.

Sonnet 5.5 · API

It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

It is a usable recommendation in form, but it is not addressed to the named reader and contains major unsupported claims.

GPT-6 Luna · API

It is a complete, actionable recommendation with reasoning and next steps that Priya could use directly.

Sonnet 5.5 · API

It provides a usable, complete recommendation with reasoning and next steps for Priya within the length limit.

Separates effect from explanationWrongRightRight
Gemini 3.8 Flash · API

It treats the conversion lift as an illusion driven by bill shock and user confusion as proven, not as a hypothesis.

GPT-6 Luna · API

Explanations are labelled as concerns or consistency, while the observed effects are reported as results.

Sonnet 5.5 · API

The buyer-surprise mechanism is labeled as a hypothesis rather than a finding, and alternative explanations are framed with 'may' and 'suggests'.

All got wrong 1

Trusts the data before reading itWrongWrongWrong
Gemini 3.8 Flash · API

It does not check trust signals such as the 50/50 split, sample size, exposure, logging, or novelty before interpreting the result.

GPT-6 Luna · API

It does not explicitly check any data-trust signal such as the sample split against the intended 50/50 ratio before relying on the results.

Sonnet 5.5 · API

It does not check any trust signal such as sample ratio, exposure, logging issues, or novelty before interpreting the results.

All got right 2

Checks guardrails before declaring a winnerRightRightRight
Gemini 3.8 Flash · API

It explicitly checks both pre-agreed guardrails: retention and refunds.

GPT-6 Luna · API

It explicitly evaluates both pre-agreed retention and refund guardrails.

Sonnet 5.5 · API

It evaluates both pre-agreed guardrails explicitly, comparing refunds to the 10% limit and retention to the 2pp limit.

Gets the base of every number rightRightRightRight
Gemini 3.8 Flash · API

The stated percentages, differences, and comparisons are computed from the supplied bases and are arithmetically sound.

GPT-6 Luna · API

All derived figures, including 1.3pp, −4.5pp, 42%, and $0.19, check out against the supplied data.

Sonnet 5.5 · API

Percentage differences and refund increase are computed from the correct bases and are consistent with the supplied data.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT100.095.02None
2GPT-6.1 SolwithAPI94.790.52None
3GPT-6 LunawithAPI91.990.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude73.657.32None
6Gemini 3.5 Flash-LitewithGemini60.357.721 capped
7Gemini 3.8 FlashwithAPI55.624.121 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review