Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 71% were usable with at most a quick edit.

Reliably right

  1. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test
  3. Checks guardrails before declaring a winner100% pass
    It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
    GPT-6 Astra · ChatGPT · Conversion up, retention down

Where it slips

  1. Trusts the data before reading it21% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims52% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly55% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We tested a shorter onboarding checklist and the result came back 'not significant'. The team wants to call it a failure and move on. Write a short note to the onboarding team saying what we should conclude from this test and what we should do next. Keep it under 300 words.

What the model was given4 items: Scenario, Readout, Qualitative, Options on the table
ScenarioTrellis is a project-management tool for small teams. The shorter checklist cuts onboarding from nine steps to five. It is already built: shipping it to everyone would take about a day of engineering. Activation means creating a first project and inviting at least one teammate within seven days.
ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8). n = 3,960 per arm. The pre-registered minimum detectable effect was 3pp.
QualitativeSupport tickets tagged 'onboarding confusion' fell from 44 to 29 during the test.
Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.
What a strong answer doesThe answer key the graders mark against

Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.

Critical failures (cap the score at 40)
  • Concludes the change has no effect
Case

Mid-level · v1.8 · synthetic · null result, onboarding

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit8 of 9 checks passed

The core reasoning and use of evidence are sound, and the only meaningful omissions are explicit trust checks and the original 3pp MDE, which could be added in a quick edit. No rechecking or reanalysis is required.

Output

Team—

This is inconclusive, not a failure. Activation was 1.7 percentage points higher with the shorter checklist, but the 95% CI (−1.4 to +4.8 points) includes no effect and possible harm—as well as a benefit of at least 3 points. The test wasn’t precise enough to settle the question.

I recommend running the test for another four weeks. That should let us assess effects around 2 points, which this change may plausibly deliver. Before extending, agree on what activation gain would justify rollout and what activation decline would stop it. Track onboarding-confusion tickets too, but compare rates per user: the raw drop from 44 to 29 is encouraging, not conclusive.

After the extension, decide using both activation and confusion—not a “significant/not significant” label alone.

Grades and run details

Decision model 89 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly78%
  • passAddresses the actual decision97%
  • passRespects explicit constraints95%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims46%
  • passProduces the required deliverable64%
  • passInterprets power correctly84%
  • failTrusts the data before reading it29%
  • passGets the base of every number right96%
Run
Run
#1
API response time
9 s
Submitted
29 Sept 2026

Usable with a quick edit8 of 9 checks passed

The recommendation is clear and the evidence is used correctly; a quick edit adding a data-quality check would make it complete.

Output

Team—

This test is inconclusive, not a failure. Activation was 32.9% with the shorter checklist versus 31.2% with the original—a +1.7 percentage-point estimate. But the 95% confidence interval (−1.4 to +4.8 points) leaves room for a modest decline, no effect, or a useful improvement. We haven’t established that the shorter checklist helps—or that it doesn’t.

The test was designed to detect a 3-point effect, so it had limited ability to resolve a smaller gain. That 3-point minimum detectable effect is a planning parameter, not a threshold below which improvements are worthless.

Confusion-tagged support tickets fell from 44 to 29. That’s encouraging, but without comparable exposure counts and uncertainty estimates, it’s supporting evidence rather than proof.

My recommendation: proceed with a staged rollout, retaining a randomized holdout. The shorter flow is already built and takes about a day to ship, making a monitored rollout reasonable—not risk-free. Track seven-day activation and confusion tickets per onboarded user, and agree on rollback criteria before expanding.

Another four weeks would improve sensitivity to roughly a 2-point effect, but would not guarantee a definitive answer. Continue the full test instead if distinguishing a gain of that size would materially change our shipping decision.

Let’s decide based on likely value, downside risk, and implementation cost—not whether the result crossed a significance threshold.

Grades and run details

Decision model 94 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly36%
  • passAddresses the actual decision40%
  • passRespects explicit constraints80%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable76%
  • passInterprets power correctly100%
  • partialTrusts the data before reading it51%
  • passGets the base of every number right84%
Run
Run
#1
API response time
10 s
Submitted
29 Sept 2026

Usable with a quick edit6 of 9 checks passed

The evidence and reasoning are solid; only light edits are needed to remove or label 'easy to reverse' and add a data-trust check.

Output

Subject: Shorter onboarding checklist: what the test tells us and what to do

Bottom line: This test didn't fail. It was inconclusive, and I recommend we ship the shorter checklist.

What the result says - Activation was 32.9% vs. 31.2% (+1.7pp). The 95% CI runs from −1.4 to +4.8pp, so "not significant" means we can't rule out zero. It doesn't mean there's no effect. - The data are compatible with a real gain of several points, and also with a small loss. The test was powered to detect 3pp, so a true effect of 1-2pp would probably look just like this. - Support tickets tagged "onboarding confusion" dropped from 44 to 29. The counts are small, so this isn't proof, but it points the same way as the activation estimate.

Why ship rather than keep testing - Shipping costs about a day of engineering, and it's easy to reverse. - The plausible downside is small (about −1.4pp at worst), and the point estimate and the ticket data both lean positive. - Four more weeks would detect ~2pp, but our observed effect is 1.7pp. We could easily wait a month and end up inconclusive again, while half of new users get the longer flow.

Next steps 1. Ship the five-step checklist to everyone. 2. Track seven-day activation against the pre-test baseline for the next 4-6 weeks. If it drops well below ~31%, we revert. 3. Keep tracking the onboarding-confusion tickets. 4. Record the test as "inconclusive, shipped on cost and risk grounds" rather than "failed," so nobody later cites it as evidence that shorter onboarding doesn't work.

If we were weighing a costly or hard-to-reverse change, I'd argue for the longer test. Here the cheaper move is to ship and monitor.

Claims the judge couldn’t find in the brief

  • Shipping is easy to reverse.

Grades and run details

Decision model 83 · LLM judge 7 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly34%
  • passAddresses the actual decision100%
  • passRespects explicit constraints63%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims30%
  • passProduces the required deliverable97%
  • passInterprets power correctly100%
  • failTrusts the data before reading it90%
  • passGets the base of every number right77%
Run
Run
#1
API response time
11 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 2

Uses the supplied evidence correctlyRightRightMixed
GPT-6 Luna · API

Every stated fact or figure is drawn correctly from the supplied readout/options, and no current-situation facts are invented.

GPT-6.1 Sol · API

All current-situation facts and figures match the readout, scenario, and options; methodological caveats are labelled rather than invented.

Sonnet 5.5 · API

It invents 'easy to reverse', which is not in the supplied context and is presented as fact.

Avoids unsupported claimsRightRightWrong
GPT-6 Luna · API

Interpretive claims such as 'not precise enough' and 'encouraging, not conclusive' are appropriately hedged and follow from the supplied data.

GPT-6.1 Sol · API

Interpretations are labelled as such, and confident claims are limited to what the supplied evidence establishes.

Sonnet 5.5 · API

Presents 'easy to reverse' as established fact without support; other inferences are hedged.

All got wrong 1

Trusts the data before reading itWrongWrongWrong
GPT-6 Luna · API

It does not explicitly check any trust signal such as sample ratio, exposure or logging quality, or novelty before interpreting the activation result.

GPT-6.1 Sol · API

It interprets the activation result without explicitly checking any trust signal such as the sample split against the intended ratio or exposure issues.

Sonnet 5.5 · API

Does not check any trust signal such as sample ratio or logging before interpreting results.

All got right 6

Addresses the actual decisionRightRightRight
GPT-6 Luna · API

It unambiguously answers that the result is inconclusive, not a failure, and says the next step is a four-week extension with pre-agreed thresholds before deciding.

GPT-6.1 Sol · API

It commits early to a staged rollout and states the condition under which it would continue the full test instead.

Sonnet 5.5 · API

Commits to shipping the shorter checklist and says it would revert if activation drops well below ~31% or argue longer if the change were costly.

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

The output is a note to the onboarding team, stays well under 300 words, and respects the requested short-note form.

GPT-6.1 Sol · API

The note is well under 300 words and addressed to the onboarding team as requested.

Sonnet 5.5 · API

Respects the form, reader, and under-300-word limit.

Identifies material uncertaintyRightRightRight
GPT-6 Luna · API

It names the unresolved true activation effect and support-ticket rate issue, bounds the activation uncertainty with the CI, and says extension plus pre-agreed thresholds would resolve it.

GPT-6.1 Sol · API

It names the CI, the 3pp MDE, qualitative uncertainty, and the condition that would change the call.

Sonnet 5.5 · API

Names effect-size uncertainty and small support-ticket counts, and says 4-6 week monitoring would resolve or trigger revert.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The requested note is present, actionable, and complete enough for the onboarding team to proceed.

GPT-6.1 Sol · API

The note provides a clear conclusion and actionable next step that the onboarding team could act on with light edits.

Sonnet 5.5 · API

It is a complete, actionable note to the onboarding team within length.

Interprets power correctlyRightRightRight
GPT-6 Luna · API

It communicates that the original test could not distinguish the observed 1.7pp effect from no effect or harm, and that more data would allow detection around 2pp.

GPT-6.1 Sol · API

It correctly explains that the test could detect about 3pp and that the CI leaves smaller effects unresolved.

Sonnet 5.5 · API

Explains MDE 3pp, the CI, and that a true 1-2pp effect would likely be missed.

Gets the base of every number rightRightRightRight
GPT-6 Luna · API

The activation figures and differences are computed from the right values, and the output flags that the ticket drop should be compared per user rather than as a raw count.

GPT-6.1 Sol · API

All derived percentages and differences match the supplied data, and raw ticket counts are not misread as rates.

Sonnet 5.5 · API

All cited percentages and differences match the supplied data, and ticket base is clear as tagged tickets.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT100.095.02None
2GPT-6.1 SolwithAPI94.790.52None
3GPT-6 LunawithAPI91.990.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude73.657.32None
6Gemini 3.5 Flash-LitewithGemini60.357.721 capped
7Gemini 3.8 FlashwithAPI55.624.121 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review