Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 71% were usable with at most a quick edit.

Reliably right

  1. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test
  3. Checks guardrails before declaring a winner100% pass
    It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
    GPT-6 Astra · ChatGPT · Conversion up, retention down

Where it slips

  1. Trusts the data before reading it21% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims52% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly55% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We tested a shorter onboarding checklist and the result came back 'not significant'. The team wants to call it a failure and move on. Write a short note to the onboarding team saying what we should conclude from this test and what we should do next. Keep it under 300 words.

What the model was given4 items: Scenario, Readout, Qualitative, Options on the table
ScenarioTrellis is a project-management tool for small teams. The shorter checklist cuts onboarding from nine steps to five. It is already built: shipping it to everyone would take about a day of engineering. Activation means creating a first project and inviting at least one teammate within seven days.
ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8). n = 3,960 per arm. The pre-registered minimum detectable effect was 3pp.
QualitativeSupport tickets tagged 'onboarding confusion' fell from 44 to 29 during the test.
Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.
What a strong answer doesThe answer key the graders mark against

Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.

Critical failures (cap the score at 40)
  • Concludes the change has no effect
Case

Mid-level · v1.8 · synthetic · null result, onboarding

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit6 of 9 checks passed

The evidence and reasoning are solid; only light edits are needed to remove or label 'easy to reverse' and add a data-trust check.

Output

Subject: Shorter onboarding checklist: what the test tells us and what to do

Bottom line: This test didn't fail. It was inconclusive, and I recommend we ship the shorter checklist.

What the result says - Activation was 32.9% vs. 31.2% (+1.7pp). The 95% CI runs from −1.4 to +4.8pp, so "not significant" means we can't rule out zero. It doesn't mean there's no effect. - The data are compatible with a real gain of several points, and also with a small loss. The test was powered to detect 3pp, so a true effect of 1-2pp would probably look just like this. - Support tickets tagged "onboarding confusion" dropped from 44 to 29. The counts are small, so this isn't proof, but it points the same way as the activation estimate.

Why ship rather than keep testing - Shipping costs about a day of engineering, and it's easy to reverse. - The plausible downside is small (about −1.4pp at worst), and the point estimate and the ticket data both lean positive. - Four more weeks would detect ~2pp, but our observed effect is 1.7pp. We could easily wait a month and end up inconclusive again, while half of new users get the longer flow.

Next steps 1. Ship the five-step checklist to everyone. 2. Track seven-day activation against the pre-test baseline for the next 4-6 weeks. If it drops well below ~31%, we revert. 3. Keep tracking the onboarding-confusion tickets. 4. Record the test as "inconclusive, shipped on cost and risk grounds" rather than "failed," so nobody later cites it as evidence that shorter onboarding doesn't work.

If we were weighing a costly or hard-to-reverse change, I'd argue for the longer test. Here the cheaper move is to ship and monitor.

Claims the judge couldn’t find in the brief

  • Shipping is easy to reverse.

Grades and run details

Decision model 83 · LLM judge 7 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly34%
  • passAddresses the actual decision100%
  • passRespects explicit constraints63%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims30%
  • passProduces the required deliverable97%
  • passInterprets power correctly100%
  • failTrusts the data before reading it90%
  • passGets the base of every number right77%
Run
Run
#1
API response time
11 s
Submitted
29 Sept 2026

Needs more than a quick edit6 of 9 checks passed

The unsupported claim about ticket splitting and the missing trust check require verification before a PM could safely act.

Output

Shorter checklist test — what it tells us and what's next

Hi all,

The shorter onboarding checklist came back "not significant." That doesn't make it a failure, and I don't think we should drop it.

What the test actually says

  • B activated at 32.9% vs 31.2% for A, a +1.7pp lift.
  • The 95% interval runs from −1.4pp to +4.8pp. Most of that range is positive, and the worst plausible case is a small loss.
  • We sized the test to detect 3pp. A real 1–2pp gain was always likely to come back "not significant." The result means the effect is smaller than 3pp or zero.1 It doesn't mean the change didn't work.

Supporting signal

Support tickets tagged "onboarding confusion" fell from 44 to 29 during the test. We didn't split tickets by arm, so this is directional only. It does point the same way as the activation result.

Recommendation: ship it

  • It's already built, and rolling it out costs about a day.
  • The likely effect is neutral to positive, and five steps are simpler to maintain than nine.
  • Running four more weeks would only detect effects of about 2pp. Our point estimate is 1.7pp, so we would probably still get an ambiguous answer, a month later.

Guardrails

  • Watch 7-day activation for the four weeks after rollout. If it drops more than about 1.5pp below the current baseline, roll back.
  • Split onboarding-confusion tickets by checklist version from now on, so next time we can tell whether a ticket drop is real.

For future tests

Before launch, agree what we'll do if the result comes back inconclusive. For a cheap, low-risk change like this, "not significant but probably not harmful" should default to ship.

Happy to discuss.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

    What we’d changeCorrect the reading: the interval runs from −1.4 to +4.8pp, so an effect above 3pp isn't ruled out. The test just wasn't sized to reliably detect smaller ones.

Claims the judge couldn’t find in the brief

  • We didn't split tickets by arm.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly45%
  • passAddresses the actual decision92%
  • passRespects explicit constraints52%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable89%
  • passInterprets power correctly99%
  • failTrusts the data before reading it65%
  • passGets the base of every number right80%
Run
Run
#1
Time to output
13 s
Submitted
25 Sept 2026

Usable with a quick edit4 of 9 checks passed

The note provides clear direction; a PM could quickly add a condition for reverting and verify SRM, but the core reasoning is sound and evidence is accurate.

Output

Hi Team,

Before we write off the shorter onboarding checklist as a failure, let's look at the full picture.

While the headline result was statistically non-significant, the numbers and qualitative feedback tell a more nuanced story:

Positive Directional Lift: Group B (5 steps) saw a 32.9% activation rate compared to 31.2% for Group A (9 steps)—a 1.7 percentage point increase.

Underpowered Test: Our pre-registered minimum detectable effect (MDE) was 3pp, meaning the test wasn't powered to reliably catch a 1.7pp or 2pp shift.

Reduced Support Friction: Support tickets tagged "onboarding confusion" dropped significantly from 44 to 29 during the test.

What we should conclude: The test wasn't a failure; it was inconclusive. We likely have a small positive gain accompanied by a meaningful reduction in user confusion.1

What we should do next: Given that shipping this to everyone only takes about a day of engineering, we should ship it. Continuing to run the test for another four weeks just to chase statistical significance on a small effect isn't a wise use of time2 when the engineering cost to launch is so low and support friction has already improved. Let's roll out the 5-step checklist and monitor downstream retention.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisStart again

    Source: ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

    What we’d changeGive the interval: anywhere from a 1.4pp loss to a 4.8pp gain, so a gain isn't established. And the ticket drop wasn't split by arm, so it can't be credited to the new checklist.

  2. 2OtherFix itTargeted repair

    Source: Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.

    What we’d changeWeigh the risk of a small activation loss, not only the one day of engineering, and say what result would justify the extra four weeks.

Grades and run details

Decision model 56 · LLM judge 7 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly28%
  • partialAddresses the actual decision36%
  • passRespects explicit constraints85%
  • partialIdentifies material uncertainty69%
  • failAvoids unsupported claims44%
  • passProduces the required deliverable88%
  • passInterprets power correctly81%
  • failTrusts the data before reading it88%
  • passGets the base of every number right50%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyMixedWrongMixed
Sonnet 5.5 · API

It invents 'easy to reverse', which is not in the supplied context and is presented as fact.

Opus 5.5 · Claude

The claim 'We didn't split tickets by arm' is not stated in the supplied context and has no support.

Gemini 3.5 Flash-Lite · Gemini

All facts are drawn directly from the supplied context without invention.

Addresses the actual decisionRightRightWrong
Sonnet 5.5 · API

Commits to shipping the shorter checklist and says it would revert if activation drops well below ~31% or argue longer if the change were costly.

Opus 5.5 · Claude

Commits unambiguously to 'ship it', specifies monitoring and rollback conditions.

Gemini 3.5 Flash-Lite · Gemini

Output commits to shipping but does not state what result or condition would change that decision.

Identifies material uncertaintyRightRightWrong
Sonnet 5.5 · API

Names effect-size uncertainty and small support-ticket counts, and says 4-6 week monitoring would resolve or trigger revert.

Opus 5.5 · Claude

Names the uncertainty around effect size, bounds it with the confidence interval, and defines a rollback trigger.

Gemini 3.5 Flash-Lite · Gemini

Output does not name the specific unknowns that could change the decision (e.g., true effect could be zero or negative) and does not bound the uncertainty; it presents the effect as likely positive without acknowledging the CI includes negative values.

Avoids unsupported claimsWrongWrongMixed
Sonnet 5.5 · API

Presents 'easy to reverse' as established fact without support; other inferences are hedged.

Opus 5.5 · Claude

Presents 'We didn't split tickets by arm' as an established fact without evidence or qualification.

Gemini 3.5 Flash-Lite · Gemini

Interpretations are labelled as 'likely' and 'meaningful reduction' is supported by ticket data; no claims presented as established fact that aren't.

All got wrong 1

Trusts the data before reading itWrongWrongWrong
Sonnet 5.5 · API

Does not check any trust signal such as sample ratio or logging before interpreting results.

Opus 5.5 · Claude

No check of sample ratio, logging issues, or any trust signal before interpreting the result.

Gemini 3.5 Flash-Lite · Gemini

No trust signal (e.g., sample ratio, SRM check) is examined before interpreting the results.

All got right 4

Respects explicit constraintsRightRightRight
Sonnet 5.5 · API

Respects the form, reader, and under-300-word limit.

Opus 5.5 · Claude

Delivers a short note under 300 words to the onboarding team addressing conclusions and next steps.

Gemini 3.5 Flash-Lite · Gemini

The note is under 300 words and respects the requested short note format.

Produces the required deliverableRightRightRight
Sonnet 5.5 · API

It is a complete, actionable note to the onboarding team within length.

Opus 5.5 · Claude

The output is a complete, usable note in the requested form and within length, ready with light edits.

Gemini 3.5 Flash-Lite · Gemini

The output is a complete note to the onboarding team, within word limit, and could be acted on with light edits.

Interprets power correctlyRightRightRight
Sonnet 5.5 · API

Explains MDE 3pp, the CI, and that a true 1-2pp effect would likely be missed.

Opus 5.5 · Claude

Explains the test was powered to detect 3pp and that a 1–2pp effect would likely be non-significant.

Gemini 3.5 Flash-Lite · Gemini

Correctly explains that the test was not powered to detect a 1.7pp or 2pp shift given the 3pp MDE.

Gets the base of every number rightRightRightRight
Sonnet 5.5 · API

All cited percentages and differences match the supplied data, and ticket base is clear as tagged tickets.

Opus 5.5 · Claude

All percentages and differences are derived correctly from the supplied activation rates and counts.

Gemini 3.5 Flash-Lite · Gemini

Percentage differences and ticket counts are presented accurately with clear bases.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT100.095.02None
2GPT-6.1 SolwithAPI94.790.52None
3GPT-6 LunawithAPI91.990.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude73.657.32None
6Gemini 3.5 Flash-LitewithGemini60.357.721 capped
7Gemini 3.8 FlashwithAPI55.624.121 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review