Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 83% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    It is a usable decision memo for the named reader with a clear recommendation, reasoning, and next steps.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  3. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Where it slips

  1. Trusts the data before reading it15% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims60% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly69% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We tested a shorter onboarding checklist and the result came back 'not significant'. The team wants to call it a failure and move on. Write a short note to the onboarding team saying what we should conclude from this test and what we should do next. Keep it under 300 words.

ScenarioTrellis is a project-management tool for small teams. The shorter checklist cuts onboarding from nine steps to five. It is already built: shipping it to everyone would take about a day of engineering. Activation means creating a first project and inviting at least one teammate within seven days.
ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8). n = 3,960 per arm. The pre-registered minimum detectable effect was 3pp.
QualitativeSupport tickets tagged 'onboarding confusion' fell from 44 to 29 during the test.
Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.
What a strong answer does

Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.

Critical failures (cap the score)
  • Concludes the change has no effect
Case

v1.6 · synthetic · null result, onboarding

Usable with a quick edit4 of 9 checks passed

The note provides clear direction; a PM could quickly add a condition for reverting and verify SRM, but the core reasoning is sound and evidence is accurate.

Output

Hi Team,

Before we write off the shorter onboarding checklist as a failure, let's look at the full picture.

While the headline result was statistically non-significant, the numbers and qualitative feedback tell a more nuanced story:

Positive Directional Lift: Group B (5 steps) saw a 32.9% activation rate compared to 31.2% for Group A (9 steps)—a 1.7 percentage point increase.

Underpowered Test: Our pre-registered minimum detectable effect (MDE) was 3pp, meaning the test wasn't powered to reliably catch a 1.7pp or 2pp shift.

Reduced Support Friction: Support tickets tagged "onboarding confusion" dropped significantly from 44 to 29 during the test.

What we should conclude: The test wasn't a failure; it was inconclusive. We likely have a small positive gain accompanied by a meaningful reduction in user confusion.1

What we should do next: Given that shipping this to everyone only takes about a day of engineering, we should ship it. Continuing to run the test for another four weeks just to chase statistical significance on a small effect isn't a wise use of time2 when the engineering cost to launch is so low and support friction has already improved. Let's roll out the 5-step checklist and monitor downstream retention.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisStart again

    Source: ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

    What we’d changeGive the interval: anywhere from a 1.4pp loss to a 4.8pp gain, so a gain isn't established. And the ticket drop wasn't split by arm, so it can't be credited to the new checklist.

  2. 2OtherFix itTargeted repair

    Source: Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.

    What we’d changeWeigh the risk of a small activation loss, not only the one day of engineering, and say what result would justify the extra four weeks.

Check by check

Got wrong · 3
  • Addresses the actual decisionOutput commits to shipping but does not state what result or condition would change that decision.
  • Identifies material uncertaintyOutput does not name the specific unknowns that could change the decision (e.g., true effect could be zero or negative) and does not bound the uncertainty; it presents the effect as likely positive without acknowledging the CI includes negative values.
  • Trusts the data before reading itNo trust signal (e.g., sample ratio, SRM check) is examined before interpreting the results.
Mixed · 2
  • Uses the supplied evidence correctlyAll facts are drawn directly from the supplied context without invention.The two graders disagreed on this one.
  • Avoids unsupported claimsInterpretations are labelled as 'likely' and 'meaningful reduction' is supported by ticket data; no claims presented as established fact that aren't.The two graders disagreed on this one.
Got right · 4
  • Respects explicit constraintsThe note is under 300 words and respects the requested short note format.
  • Produces the required deliverableThe output is a complete note to the onboarding team, within word limit, and could be acted on with light edits.
  • Interprets power correctlyCorrectly explains that the test was not powered to detect a 1.7pp or 2pp shift given the 3pp MDE.
  • Gets the base of every number rightPercentage differences and ticket counts are presented accurately with clear bases.

Grades and run details

Decision model 56 · LLM judge 7 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly7%
  • partialAddresses the actual decision39%
  • passRespects explicit constraints87%
  • partialIdentifies material uncertainty73%
  • failAvoids unsupported claims40%
  • passProduces the required deliverable89%
  • passInterprets power correctly74%
  • failTrusts the data before reading it97%
  • passGets the base of every number right55%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Usable with a quick edit8 of 9 checks passed

The core reasoning and use of evidence are sound, and the only meaningful omissions are explicit trust checks and the original 3pp MDE, which could be added in a quick edit. No rechecking or reanalysis is required.

Output

Team—

This is inconclusive, not a failure. Activation was 1.7 percentage points higher with the shorter checklist, but the 95% CI (−1.4 to +4.8 points) includes no effect and possible harm—as well as a benefit of at least 3 points. The test wasn’t precise enough to settle the question.

I recommend running the test for another four weeks. That should let us assess effects around 2 points, which this change may plausibly deliver. Before extending, agree on what activation gain would justify rollout and what activation decline would stop it. Track onboarding-confusion tickets too, but compare rates per user: the raw drop from 44 to 29 is encouraging, not conclusive.

After the extension, decide using both activation and confusion—not a “significant/not significant” label alone.

Check by check

Got wrong · 1
  • Trusts the data before reading itIt does not explicitly check any trust signal such as sample ratio, exposure or logging quality, or novelty before interpreting the activation result.
Got right · 8
  • Uses the supplied evidence correctlyEvery stated fact or figure is drawn correctly from the supplied readout/options, and no current-situation facts are invented.
  • Addresses the actual decisionIt unambiguously answers that the result is inconclusive, not a failure, and says the next step is a four-week extension with pre-agreed thresholds before deciding.
  • Respects explicit constraintsThe output is a note to the onboarding team, stays well under 300 words, and respects the requested short-note form.
  • Identifies material uncertaintyIt names the unresolved true activation effect and support-ticket rate issue, bounds the activation uncertainty with the CI, and says extension plus pre-agreed thresholds would resolve it.
  • Avoids unsupported claimsInterpretive claims such as 'not precise enough' and 'encouraging, not conclusive' are appropriately hedged and follow from the supplied data.
  • Produces the required deliverableThe requested note is present, actionable, and complete enough for the onboarding team to proceed.
  • Interprets power correctlyIt communicates that the original test could not distinguish the observed 1.7pp effect from no effect or harm, and that more data would allow detection around 2pp.
  • Gets the base of every number rightThe activation figures and differences are computed from the right values, and the output flags that the ticket drop should be compared per user rather than as a raw count.

Grades and run details

Decision model 89 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly79%
  • passAddresses the actual decision97%
  • passRespects explicit constraints95%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims47%
  • passProduces the required deliverable66%
  • passInterprets power correctly86%
  • failTrusts the data before reading it79%
  • passGets the base of every number right95%
Run
Run
#1
API response time
9 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT94.795.02None
2GPT-6 LunawithAPI89.490.52None
3GPT-6.1 SolwithAPI89.490.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude76.157.32None
6Gemini 3.5 Flash-LitewithGemini62.857.721 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review