Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 71% were usable with at most a quick edit.

Reliably right

  1. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test
  3. Checks guardrails before declaring a winner100% pass
    It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
    GPT-6 Astra · ChatGPT · Conversion up, retention down

Where it slips

  1. Trusts the data before reading it21% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims52% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly55% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We tested a shorter onboarding checklist and the result came back 'not significant'. The team wants to call it a failure and move on. Write a short note to the onboarding team saying what we should conclude from this test and what we should do next. Keep it under 300 words.

What the model was given4 items: Scenario, Readout, Qualitative, Options on the table
ScenarioTrellis is a project-management tool for small teams. The shorter checklist cuts onboarding from nine steps to five. It is already built: shipping it to everyone would take about a day of engineering. Activation means creating a first project and inviting at least one teammate within seven days.
ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8). n = 3,960 per arm. The pre-registered minimum detectable effect was 3pp.
QualitativeSupport tickets tagged 'onboarding confusion' fell from 44 to 29 during the test.
Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.
What a strong answer doesThe answer key the graders mark against

Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.

Critical failures (cap the score at 40)
  • Concludes the change has no effect
Case

Mid-level · v1.8 · synthetic · null result, onboarding

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit3 of 9 checks passed

A PM would need to re-check the evidence and add rollback or monitoring thresholds before acting. The note overstates certainty and omits what would change the ship decision.

Output

Team,

Before we write off the shorter checklist as a failure, let’s reframe what the data actually tells us.

“Not statistically significant” does not mean “no impact”—it means our test was underpowered to detect an effect smaller than 3pp. Here is the full picture:

  1. The directional signals are uniformly positive: Activation increased by +1.7pp (31.2% → 32.9%). While the 95% CI (−1.4pp to +4.8pp) crosses zero, it leans heavily positive. The downside risk is small, while the potential upside is meaningful.
  2. Clear reduction in user friction: Support tickets tagged “onboarding confusion” dropped by 34% (44 to 29). Even if activation were completely flat, cutting nearly half the steps improved user experience and reduced support burden.
  3. The cost/benefit is strongly asymmetric: Shipping this takes one day of engineering.

Recommendation: Ship the 5-step checklist to 100% and move on.

Running the test for another four weeks to hit an MDE of 2pp isn’t the best use of our time or traffic. We would burn a month of runway just to confirm what looks like a positive, low-risk change with virtually zero engineering cost.

Let’s bank the operational win (fewer support tickets, leaner UX, likely small activation lift) by rolling it out tomorrow. We can monitor activation metrics in production and direct our experimentation bandwidth toward bigger, higher-leverage bets for Trellis.

Claims the judge couldn’t find in the brief

  • The CI “leans heavily positive.”
  • The downside risk is small and the potential upside is meaningful.
  • Cutting nearly half the steps improved user experience and reduced support burden.
  • The change is positive and low-risk.
  • The engineering cost is virtually zero.
  • There is likely a small activation lift.

Grades and run details

Decision model 61 · LLM judge 3 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly29%
  • passAddresses the actual decision30%
  • passRespects explicit constraints76%
  • partialIdentifies material uncertainty67%
  • failAvoids unsupported claims79%
  • passProduces the required deliverable87%
  • passInterprets power correctly92%
  • failTrusts the data before reading it98%
  • passGets the base of every number right62%
Run
Run
#1
API response time
11 s
Submitted
1 Oct 2026

Needs more than a quick edit6 of 9 checks passed

The unsupported claim about ticket splitting and the missing trust check require verification before a PM could safely act.

Output

Shorter checklist test — what it tells us and what's next

Hi all,

The shorter onboarding checklist came back "not significant." That doesn't make it a failure, and I don't think we should drop it.

What the test actually says

  • B activated at 32.9% vs 31.2% for A, a +1.7pp lift.
  • The 95% interval runs from −1.4pp to +4.8pp. Most of that range is positive, and the worst plausible case is a small loss.
  • We sized the test to detect 3pp. A real 1–2pp gain was always likely to come back "not significant." The result means the effect is smaller than 3pp or zero.1 It doesn't mean the change didn't work.

Supporting signal

Support tickets tagged "onboarding confusion" fell from 44 to 29 during the test. We didn't split tickets by arm, so this is directional only. It does point the same way as the activation result.

Recommendation: ship it

  • It's already built, and rolling it out costs about a day.
  • The likely effect is neutral to positive, and five steps are simpler to maintain than nine.
  • Running four more weeks would only detect effects of about 2pp. Our point estimate is 1.7pp, so we would probably still get an ambiguous answer, a month later.

Guardrails

  • Watch 7-day activation for the four weeks after rollout. If it drops more than about 1.5pp below the current baseline, roll back.
  • Split onboarding-confusion tickets by checklist version from now on, so next time we can tell whether a ticket drop is real.

For future tests

Before launch, agree what we'll do if the result comes back inconclusive. For a cheap, low-risk change like this, "not significant but probably not harmful" should default to ship.

Happy to discuss.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

    What we’d changeCorrect the reading: the interval runs from −1.4 to +4.8pp, so an effect above 3pp isn't ruled out. The test just wasn't sized to reliably detect smaller ones.

Claims the judge couldn’t find in the brief

  • We didn't split tickets by arm.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly45%
  • passAddresses the actual decision92%
  • passRespects explicit constraints52%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable89%
  • passInterprets power correctly99%
  • failTrusts the data before reading it65%
  • passGets the base of every number right80%
Run
Run
#1
Time to output
13 s
Submitted
25 Sept 2026

Usable with a quick edit8 of 9 checks passed

The recommendation is clear and the evidence is used correctly; a quick edit adding a data-quality check would make it complete.

Output

Team—

This test is inconclusive, not a failure. Activation was 32.9% with the shorter checklist versus 31.2% with the original—a +1.7 percentage-point estimate. But the 95% confidence interval (−1.4 to +4.8 points) leaves room for a modest decline, no effect, or a useful improvement. We haven’t established that the shorter checklist helps—or that it doesn’t.

The test was designed to detect a 3-point effect, so it had limited ability to resolve a smaller gain. That 3-point minimum detectable effect is a planning parameter, not a threshold below which improvements are worthless.

Confusion-tagged support tickets fell from 44 to 29. That’s encouraging, but without comparable exposure counts and uncertainty estimates, it’s supporting evidence rather than proof.

My recommendation: proceed with a staged rollout, retaining a randomized holdout. The shorter flow is already built and takes about a day to ship, making a monitored rollout reasonable—not risk-free. Track seven-day activation and confusion tickets per onboarded user, and agree on rollback criteria before expanding.

Another four weeks would improve sensitivity to roughly a 2-point effect, but would not guarantee a definitive answer. Continue the full test instead if distinguishing a gain of that size would materially change our shipping decision.

Let’s decide based on likely value, downside risk, and implementation cost—not whether the result crossed a significance threshold.

Grades and run details

Decision model 94 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly36%
  • passAddresses the actual decision40%
  • passRespects explicit constraints80%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable76%
  • passInterprets power correctly100%
  • partialTrusts the data before reading it51%
  • passGets the base of every number right84%
Run
Run
#1
API response time
10 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyWrongWrongRight
Gemini 3.8 Flash · API

It presents several unsupported interpretations and causal claims as facts, such as the CI leaning heavily positive, low risk, improved UX, and reduced support burden.

Opus 5.5 · Claude

The claim 'We didn't split tickets by arm' is not stated in the supplied context and has no support.

GPT-6.1 Sol · API

All current-situation facts and figures match the readout, scenario, and options; methodological caveats are labelled rather than invented.

Addresses the actual decisionMixedRightRight
Gemini 3.8 Flash · API

It commits to shipping, but does not state what result or condition would change that call.

Opus 5.5 · Claude

Commits unambiguously to 'ship it', specifies monitoring and rollback conditions.

GPT-6.1 Sol · API

It commits early to a staged rollout and states the condition under which it would continue the full test instead.

Identifies material uncertaintyWrongRightRight
Gemini 3.8 Flash · API

It notes the CI crosses zero but does not name the unknowns that would change the decision or how they would be resolved.

Opus 5.5 · Claude

Names the uncertainty around effect size, bounds it with the confidence interval, and defines a rollback trigger.

GPT-6.1 Sol · API

It names the CI, the 3pp MDE, qualitative uncertainty, and the condition that would change the call.

Avoids unsupported claimsWrongWrongRight
Gemini 3.8 Flash · API

It treats friction reduction, UX improvement, support-burden reduction, low risk, and likely activation lift as established rather than labelled hypotheses.

Opus 5.5 · Claude

Presents 'We didn't split tickets by arm' as an established fact without evidence or qualification.

GPT-6.1 Sol · API

Interpretations are labelled as such, and confident claims are limited to what the supplied evidence establishes.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

Although it is a note within length, it is not safely usable because the recommendation lacks decision thresholds and overstates the evidence.

Opus 5.5 · Claude

The output is a complete, usable note in the requested form and within length, ready with light edits.

GPT-6.1 Sol · API

The note provides a clear conclusion and actionable next step that the onboarding team could act on with light edits.

All got wrong 1

Trusts the data before reading itWrongWrongWrong
Gemini 3.8 Flash · API

It does not check trust signals such as sample ratio, exposure, logging, or novelty before interpreting the result.

Opus 5.5 · Claude

No check of sample ratio, logging issues, or any trust signal before interpreting the result.

GPT-6.1 Sol · API

It interprets the activation result without explicitly checking any trust signal such as the sample split against the intended ratio or exposure issues.

All got right 3

Respects explicit constraintsRightRightRight
Gemini 3.8 Flash · API

It is a short note to the onboarding team and is under 300 words.

Opus 5.5 · Claude

Delivers a short note under 300 words to the onboarding team addressing conclusions and next steps.

GPT-6.1 Sol · API

The note is well under 300 words and addressed to the onboarding team as requested.

Interprets power correctlyRightRightRight
Gemini 3.8 Flash · API

It correctly says the test was underpowered for effects smaller than about 3pp and that the result is not evidence of no effect.

Opus 5.5 · Claude

Explains the test was powered to detect 3pp and that a 1–2pp effect would likely be non-significant.

GPT-6.1 Sol · API

It correctly explains that the test could detect about 3pp and that the CI leaves smaller effects unresolved.

Gets the base of every number rightRightRightRight
Gemini 3.8 Flash · API

The activation difference, CI, ticket percentage, and step reduction are computed from the correct bases.

Opus 5.5 · Claude

All percentages and differences are derived correctly from the supplied activation rates and counts.

GPT-6.1 Sol · API

All derived percentages and differences match the supplied data, and raw ticket counts are not misread as rates.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT100.095.02None
2GPT-6.1 SolwithAPI94.790.52None
3GPT-6 LunawithAPI91.990.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude73.657.32None
6Gemini 3.5 Flash-LitewithGemini60.357.721 capped
7Gemini 3.8 FlashwithAPI55.624.121 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review