Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 71% were usable with at most a quick edit.

Reliably right

  1. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test
  3. Checks guardrails before declaring a winner100% pass
    It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
    GPT-6 Astra · ChatGPT · Conversion up, retention down

Where it slips

  1. Trusts the data before reading it21% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims52% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly55% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We tested a shorter onboarding checklist and the result came back 'not significant'. The team wants to call it a failure and move on. Write a short note to the onboarding team saying what we should conclude from this test and what we should do next. Keep it under 300 words.

What the model was given4 items: Scenario, Readout, Qualitative, Options on the table
ScenarioTrellis is a project-management tool for small teams. The shorter checklist cuts onboarding from nine steps to five. It is already built: shipping it to everyone would take about a day of engineering. Activation means creating a first project and inviting at least one teammate within seven days.
ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8). n = 3,960 per arm. The pre-registered minimum detectable effect was 3pp.
QualitativeSupport tickets tagged 'onboarding confusion' fell from 44 to 29 during the test.
Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.
What a strong answer doesThe answer key the graders mark against

Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.

Critical failures (cap the score at 40)
  • Concludes the change has no effect
Case

Mid-level · v1.8 · synthetic · null result, onboarding

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit8 of 9 checks passed

The note is actionable and well-structured; the one unsupported claim about missing per-arm data can be fixed by rephrasing as a check to perform, without redoing the reasoning.

Output

Team,

This test is inconclusive, not a failure. Activation increased from 31.2% to 32.9%, but the confidence interval spans a 1.4 percentage-point decline to a 4.8-point improvement. We haven’t established a benefit or ruled out meaningful harm.

The pre-registered 3pp minimum detectable effect describes what we designed the test to detect reliably. It isn’t a threshold below which improvements have no value. A smaller lift could still be worthwhile.

Confusion tickets fell from 44 to 29, which is encouraging supporting evidence. Without comparable rates by experiment arm, consistent tagging and exposure counts, we can’t attribute that reduction to the shorter checklist.

My recommendation is a staged rollout, provided basic checks reveal no downstream harm. With only a day of engineering needed, I wouldn’t automatically spend another four weeks pursuing statistical significance. The remaining decision is whether the likely benefit justifies the remaining cost and risk.

Before rollout, confirm every included user has had the full seven-day activation window and check experiment allocation, tracking and any available downstream engagement measures. Compare confusion-ticket rates between arms.

Then increase exposure gradually, retaining a randomised holdout and agreeing monitoring dates and a rollback threshold upfront. This is a practical decision under uncertainty, not a declaration that B won.

Extend the experiment instead if a possible 1.4pp activation loss is commercially unacceptable1 or if establishing a roughly 2pp improvement would change our decision. If we extend, pre-specify the analysis and stopping rule; four more weeks offers greater precision, not a guaranteed answer.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticQuick edit

    Source: ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

    What we’d changeDescribe 1.4pp as the lower end of the interval, not the largest possible loss, and qualify the one-day cost assumption.

Claims the judge couldn’t find in the brief

  • Comparable confusion-ticket rates by experiment arm, consistent tagging and exposure counts are not available.

Grades and run details

Decision model 100 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly20%
  • passAddresses the actual decision71%
  • passRespects explicit constraints51%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims36%
  • passProduces the required deliverable69%
  • passInterprets power correctly99%
  • passTrusts the data before reading it50%
  • passGets the base of every number right93%
Run
Run
#1
Time to output
21 s
Submitted
25 Sept 2026

Usable with a quick edit8 of 9 checks passed

The core reasoning and use of evidence are sound, and the only meaningful omissions are explicit trust checks and the original 3pp MDE, which could be added in a quick edit. No rechecking or reanalysis is required.

Output

Team—

This is inconclusive, not a failure. Activation was 1.7 percentage points higher with the shorter checklist, but the 95% CI (−1.4 to +4.8 points) includes no effect and possible harm—as well as a benefit of at least 3 points. The test wasn’t precise enough to settle the question.

I recommend running the test for another four weeks. That should let us assess effects around 2 points, which this change may plausibly deliver. Before extending, agree on what activation gain would justify rollout and what activation decline would stop it. Track onboarding-confusion tickets too, but compare rates per user: the raw drop from 44 to 29 is encouraging, not conclusive.

After the extension, decide using both activation and confusion—not a “significant/not significant” label alone.

Grades and run details

Decision model 89 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly78%
  • passAddresses the actual decision97%
  • passRespects explicit constraints95%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims46%
  • passProduces the required deliverable64%
  • passInterprets power correctly84%
  • failTrusts the data before reading it29%
  • passGets the base of every number right96%
Run
Run
#1
API response time
9 s
Submitted
29 Sept 2026

Needs more than a quick edit6 of 9 checks passed

The unsupported claim about ticket splitting and the missing trust check require verification before a PM could safely act.

Output

Shorter checklist test — what it tells us and what's next

Hi all,

The shorter onboarding checklist came back "not significant." That doesn't make it a failure, and I don't think we should drop it.

What the test actually says

  • B activated at 32.9% vs 31.2% for A, a +1.7pp lift.
  • The 95% interval runs from −1.4pp to +4.8pp. Most of that range is positive, and the worst plausible case is a small loss.
  • We sized the test to detect 3pp. A real 1–2pp gain was always likely to come back "not significant." The result means the effect is smaller than 3pp or zero.1 It doesn't mean the change didn't work.

Supporting signal

Support tickets tagged "onboarding confusion" fell from 44 to 29 during the test. We didn't split tickets by arm, so this is directional only. It does point the same way as the activation result.

Recommendation: ship it

  • It's already built, and rolling it out costs about a day.
  • The likely effect is neutral to positive, and five steps are simpler to maintain than nine.
  • Running four more weeks would only detect effects of about 2pp. Our point estimate is 1.7pp, so we would probably still get an ambiguous answer, a month later.

Guardrails

  • Watch 7-day activation for the four weeks after rollout. If it drops more than about 1.5pp below the current baseline, roll back.
  • Split onboarding-confusion tickets by checklist version from now on, so next time we can tell whether a ticket drop is real.

For future tests

Before launch, agree what we'll do if the result comes back inconclusive. For a cheap, low-risk change like this, "not significant but probably not harmful" should default to ship.

Happy to discuss.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

    What we’d changeCorrect the reading: the interval runs from −1.4 to +4.8pp, so an effect above 3pp isn't ruled out. The test just wasn't sized to reliably detect smaller ones.

Claims the judge couldn’t find in the brief

  • We didn't split tickets by arm.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly45%
  • passAddresses the actual decision92%
  • passRespects explicit constraints52%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable89%
  • passInterprets power correctly99%
  • failTrusts the data before reading it65%
  • passGets the base of every number right80%
Run
Run
#1
Time to output
13 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Uses the supplied evidence correctlyMixedRightWrong
GPT-6 Astra · ChatGPT

Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.

GPT-6 Luna · API

Every stated fact or figure is drawn correctly from the supplied readout/options, and no current-situation facts are invented.

Opus 5.5 · Claude

The claim 'We didn't split tickets by arm' is not stated in the supplied context and has no support.

Avoids unsupported claimsRightRightWrong
GPT-6 Astra · ChatGPT

Does not present interpretations as fact; treats the ticket drop as encouraging but not attributed, and does not overstate the statistical result.

GPT-6 Luna · API

Interpretive claims such as 'not precise enough' and 'encouraging, not conclusive' are appropriately hedged and follow from the supplied data.

Opus 5.5 · Claude

Presents 'We didn't split tickets by arm' as an established fact without evidence or qualification.

Trusts the data before reading itRightWrongWrong
GPT-6 Astra · ChatGPT

Recommends checking experiment allocation and tracking before rollout, which acts as a trust-signal check.

GPT-6 Luna · API

It does not explicitly check any trust signal such as sample ratio, exposure or logging quality, or novelty before interpreting the activation result.

Opus 5.5 · Claude

No check of sample ratio, logging issues, or any trust signal before interpreting the result.

All got right 6

Addresses the actual decisionRightRightRight
GPT-6 Astra · ChatGPT

Commits early to a staged rollout (conditional on checks) and states that a commercially unacceptable 1.4pp loss or a desire for a 2pp improvement would change the call.

GPT-6 Luna · API

It unambiguously answers that the result is inconclusive, not a failure, and says the next step is a four-week extension with pre-agreed thresholds before deciding.

Opus 5.5 · Claude

Commits unambiguously to 'ship it', specifies monitoring and rollback conditions.

Respects explicit constraintsRightRightRight
GPT-6 Astra · ChatGPT

Respects the word limit, addresses the onboarding team, and does not call the variant a failure.

GPT-6 Luna · API

The output is a note to the onboarding team, stays well under 300 words, and respects the requested short-note form.

Opus 5.5 · Claude

Delivers a short note under 300 words to the onboarding team addressing conclusions and next steps.

Identifies material uncertaintyRightRightRight
GPT-6 Astra · ChatGPT

Identifies that the CI includes harm, that ticket attribution is uncertain without per-arm data, and says extending the test would resolve the effect-size uncertainty.

GPT-6 Luna · API

It names the unresolved true activation effect and support-ticket rate issue, bounds the activation uncertainty with the CI, and says extension plus pre-agreed thresholds would resolve it.

Opus 5.5 · Claude

Names the uncertainty around effect size, bounds it with the confidence interval, and defines a rollback trigger.

Produces the required deliverableRightRightRight
GPT-6 Astra · ChatGPT

Delivers a short note to the onboarding team with an actionable recommendation and next steps, within 300 words.

GPT-6 Luna · API

The requested note is present, actionable, and complete enough for the onboarding team to proceed.

Opus 5.5 · Claude

The output is a complete, usable note in the requested form and within length, ready with light edits.

Interprets power correctlyRightRightRight
GPT-6 Astra · ChatGPT

Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.

GPT-6 Luna · API

It communicates that the original test could not distinguish the observed 1.7pp effect from no effect or harm, and that more data would allow detection around 2pp.

Opus 5.5 · Claude

Explains the test was powered to detect 3pp and that a 1–2pp effect would likely be non-significant.

Gets the base of every number rightRightRightRight
GPT-6 Astra · ChatGPT

All percentages are clearly from the arm-level activation rates, and ticket numbers are reported as raw counts without a confusing base.

GPT-6 Luna · API

The activation figures and differences are computed from the right values, and the output flags that the ticket drop should be compared per user rather than as a raw count.

Opus 5.5 · Claude

All percentages and differences are derived correctly from the supplied activation rates and counts.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT100.095.02None
2GPT-6.1 SolwithAPI94.790.52None
3GPT-6 LunawithAPI91.990.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude73.657.32None
6Gemini 3.5 Flash-LitewithGemini60.357.721 capped
7Gemini 3.8 FlashwithAPI55.624.121 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review