Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 71% were usable with at most a quick edit.

Reliably right

  1. Gets the base of every number right100% pass
    The refund increase and revenue difference are computed correctly, and the retention CI interpretation uses the correct 2pp threshold.
    GPT-6 Astra · ChatGPT · Conversion up, retention down
  2. Interprets power correctly100% pass
    Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test
  3. Checks guardrails before declaring a winner100% pass
    It evaluates both the day-30 retention guardrail and the refund-request guardrail before making a call.
    GPT-6 Astra · ChatGPT · Conversion up, retention down

Where it slips

  1. Trusts the data before reading it21% pass
    It does not explicitly check a pre-interpretation trust signal such as sample ratio or logging before relying on the lift; the denominator and maturity caveats come after interpreting.
    GPT-6.1 Sol · API · Conversion up, retention down
  2. Avoids unsupported claims52% pass
    It asserts the refund surge 'isn't noise' and that annual up-front billing means retention 'should rise' without supporting evidence or clear labeling.
    Sonnet 5.5 · API · Conversion up, retention down
  3. Uses the supplied evidence correctly55% pass
    Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.
    GPT-6 Astra · ChatGPT · The underpowered onboarding test

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We tested a shorter onboarding checklist and the result came back 'not significant'. The team wants to call it a failure and move on. Write a short note to the onboarding team saying what we should conclude from this test and what we should do next. Keep it under 300 words.

What the model was given4 items: Scenario, Readout, Qualitative, Options on the table
ScenarioTrellis is a project-management tool for small teams. The shorter checklist cuts onboarding from nine steps to five. It is already built: shipping it to everyone would take about a day of engineering. Activation means creating a first project and inviting at least one teammate within seven days.
ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8). n = 3,960 per arm. The pre-registered minimum detectable effect was 3pp.
QualitativeSupport tickets tagged 'onboarding confusion' fell from 44 to 29 during the test.
Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.
What a strong answer doesThe answer key the graders mark against

Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.

Critical failures (cap the score at 40)
  • Concludes the change has no effect
Case

Mid-level · v1.8 · synthetic · null result, onboarding

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit4 of 9 checks passed

The note provides clear direction; a PM could quickly add a condition for reverting and verify SRM, but the core reasoning is sound and evidence is accurate.

Output

Hi Team,

Before we write off the shorter onboarding checklist as a failure, let's look at the full picture.

While the headline result was statistically non-significant, the numbers and qualitative feedback tell a more nuanced story:

Positive Directional Lift: Group B (5 steps) saw a 32.9% activation rate compared to 31.2% for Group A (9 steps)—a 1.7 percentage point increase.

Underpowered Test: Our pre-registered minimum detectable effect (MDE) was 3pp, meaning the test wasn't powered to reliably catch a 1.7pp or 2pp shift.

Reduced Support Friction: Support tickets tagged "onboarding confusion" dropped significantly from 44 to 29 during the test.

What we should conclude: The test wasn't a failure; it was inconclusive. We likely have a small positive gain accompanied by a meaningful reduction in user confusion.1

What we should do next: Given that shipping this to everyone only takes about a day of engineering, we should ship it. Continuing to run the test for another four weeks just to chase statistical significance on a small effect isn't a wise use of time2 when the engineering cost to launch is so low and support friction has already improved. Let's roll out the 5-step checklist and monitor downstream retention.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisStart again

    Source: ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

    What we’d changeGive the interval: anywhere from a 1.4pp loss to a 4.8pp gain, so a gain isn't established. And the ticket drop wasn't split by arm, so it can't be credited to the new checklist.

  2. 2OtherFix itTargeted repair

    Source: Options on the tableRunning the test for another four weeks at current traffic would detect an effect of about 2pp.

    What we’d changeWeigh the risk of a small activation loss, not only the one day of engineering, and say what result would justify the extra four weeks.

Grades and run details

Decision model 56 · LLM judge 7 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly28%
  • partialAddresses the actual decision36%
  • passRespects explicit constraints85%
  • partialIdentifies material uncertainty69%
  • failAvoids unsupported claims44%
  • passProduces the required deliverable88%
  • passInterprets power correctly81%
  • failTrusts the data before reading it88%
  • passGets the base of every number right50%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Usable with a quick edit8 of 9 checks passed

The recommendation is clear and the evidence is used correctly; a quick edit adding a data-quality check would make it complete.

Output

Team—

This test is inconclusive, not a failure. Activation was 32.9% with the shorter checklist versus 31.2% with the original—a +1.7 percentage-point estimate. But the 95% confidence interval (−1.4 to +4.8 points) leaves room for a modest decline, no effect, or a useful improvement. We haven’t established that the shorter checklist helps—or that it doesn’t.

The test was designed to detect a 3-point effect, so it had limited ability to resolve a smaller gain. That 3-point minimum detectable effect is a planning parameter, not a threshold below which improvements are worthless.

Confusion-tagged support tickets fell from 44 to 29. That’s encouraging, but without comparable exposure counts and uncertainty estimates, it’s supporting evidence rather than proof.

My recommendation: proceed with a staged rollout, retaining a randomized holdout. The shorter flow is already built and takes about a day to ship, making a monitored rollout reasonable—not risk-free. Track seven-day activation and confusion tickets per onboarded user, and agree on rollback criteria before expanding.

Another four weeks would improve sensitivity to roughly a 2-point effect, but would not guarantee a definitive answer. Continue the full test instead if distinguishing a gain of that size would materially change our shipping decision.

Let’s decide based on likely value, downside risk, and implementation cost—not whether the result crossed a significance threshold.

Grades and run details

Decision model 94 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly36%
  • passAddresses the actual decision40%
  • passRespects explicit constraints80%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable76%
  • passInterprets power correctly100%
  • partialTrusts the data before reading it51%
  • passGets the base of every number right84%
Run
Run
#1
API response time
10 s
Submitted
29 Sept 2026

Usable with a quick edit8 of 9 checks passed

The note is actionable and well-structured; the one unsupported claim about missing per-arm data can be fixed by rephrasing as a check to perform, without redoing the reasoning.

Output

Team,

This test is inconclusive, not a failure. Activation increased from 31.2% to 32.9%, but the confidence interval spans a 1.4 percentage-point decline to a 4.8-point improvement. We haven’t established a benefit or ruled out meaningful harm.

The pre-registered 3pp minimum detectable effect describes what we designed the test to detect reliably. It isn’t a threshold below which improvements have no value. A smaller lift could still be worthwhile.

Confusion tickets fell from 44 to 29, which is encouraging supporting evidence. Without comparable rates by experiment arm, consistent tagging and exposure counts, we can’t attribute that reduction to the shorter checklist.

My recommendation is a staged rollout, provided basic checks reveal no downstream harm. With only a day of engineering needed, I wouldn’t automatically spend another four weeks pursuing statistical significance. The remaining decision is whether the likely benefit justifies the remaining cost and risk.

Before rollout, confirm every included user has had the full seven-day activation window and check experiment allocation, tracking and any available downstream engagement measures. Compare confusion-ticket rates between arms.

Then increase exposure gradually, retaining a randomised holdout and agreeing monitoring dates and a rollback threshold upfront. This is a practical decision under uncertainty, not a declaration that B won.

Extend the experiment instead if a possible 1.4pp activation loss is commercially unacceptable1 or if establishing a roughly 2pp improvement would change our decision. If we extend, pre-specify the analysis and stopping rule; four more weeks offers greater precision, not a guaranteed answer.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticQuick edit

    Source: ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8).

    What we’d changeDescribe 1.4pp as the lower end of the interval, not the largest possible loss, and qualify the one-day cost assumption.

Claims the judge couldn’t find in the brief

  • Comparable confusion-ticket rates by experiment arm, consistent tagging and exposure counts are not available.

Grades and run details

Decision model 100 · LLM judge 9 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly20%
  • passAddresses the actual decision71%
  • passRespects explicit constraints51%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims36%
  • passProduces the required deliverable69%
  • passInterprets power correctly99%
  • passTrusts the data before reading it50%
  • passGets the base of every number right93%
Run
Run
#1
Time to output
21 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyMixedRightMixed
Gemini 3.5 Flash-Lite · Gemini

All facts are drawn directly from the supplied context without invention.

GPT-6.1 Sol · API

All current-situation facts and figures match the readout, scenario, and options; methodological caveats are labelled rather than invented.

GPT-6 Astra · ChatGPT

Asserts that comparable per-arm ticket rates and exposure counts are unavailable, which is not in the supplied context and cannot be assumed.

Addresses the actual decisionWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

Output commits to shipping but does not state what result or condition would change that decision.

GPT-6.1 Sol · API

It commits early to a staged rollout and states the condition under which it would continue the full test instead.

GPT-6 Astra · ChatGPT

Commits early to a staged rollout (conditional on checks) and states that a commercially unacceptable 1.4pp loss or a desire for a 2pp improvement would change the call.

Identifies material uncertaintyWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

Output does not name the specific unknowns that could change the decision (e.g., true effect could be zero or negative) and does not bound the uncertainty; it presents the effect as likely positive without acknowledging the CI includes negative values.

GPT-6.1 Sol · API

It names the CI, the 3pp MDE, qualitative uncertainty, and the condition that would change the call.

GPT-6 Astra · ChatGPT

Identifies that the CI includes harm, that ticket attribution is uncertain without per-arm data, and says extending the test would resolve the effect-size uncertainty.

Avoids unsupported claimsMixedRightRight
Gemini 3.5 Flash-Lite · Gemini

Interpretations are labelled as 'likely' and 'meaningful reduction' is supported by ticket data; no claims presented as established fact that aren't.

GPT-6.1 Sol · API

Interpretations are labelled as such, and confident claims are limited to what the supplied evidence establishes.

GPT-6 Astra · ChatGPT

Does not present interpretations as fact; treats the ticket drop as encouraging but not attributed, and does not overstate the statistical result.

Trusts the data before reading itWrongWrongRight
Gemini 3.5 Flash-Lite · Gemini

No trust signal (e.g., sample ratio, SRM check) is examined before interpreting the results.

GPT-6.1 Sol · API

It interprets the activation result without explicitly checking any trust signal such as the sample split against the intended ratio or exposure issues.

GPT-6 Astra · ChatGPT

Recommends checking experiment allocation and tracking before rollout, which acts as a trust-signal check.

All got right 4

Respects explicit constraintsRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The note is under 300 words and respects the requested short note format.

GPT-6.1 Sol · API

The note is well under 300 words and addressed to the onboarding team as requested.

GPT-6 Astra · ChatGPT

Respects the word limit, addresses the onboarding team, and does not call the variant a failure.

Produces the required deliverableRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The output is a complete note to the onboarding team, within word limit, and could be acted on with light edits.

GPT-6.1 Sol · API

The note provides a clear conclusion and actionable next step that the onboarding team could act on with light edits.

GPT-6 Astra · ChatGPT

Delivers a short note to the onboarding team with an actionable recommendation and next steps, within 300 words.

Interprets power correctlyRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Correctly explains that the test was not powered to detect a 1.7pp or 2pp shift given the 3pp MDE.

GPT-6.1 Sol · API

It correctly explains that the test could detect about 3pp and that the CI leaves smaller effects unresolved.

GPT-6 Astra · ChatGPT

Explains that the test was designed to detect ~3pp effects, so a smaller but worthwhile lift could still exist and remain undetected.

Gets the base of every number rightRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Percentage differences and ticket counts are presented accurately with clear bases.

GPT-6.1 Sol · API

All derived percentages and differences match the supplied data, and raw ticket counts are not misread as rates.

GPT-6 Astra · ChatGPT

All percentages are clearly from the arm-level activation rates, and ticket numbers are reported as raw counts without a confusing base.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT100.095.02None
2GPT-6.1 SolwithAPI94.790.52None
3GPT-6 LunawithAPI91.990.52None
4Sonnet 5.5withAPI81.775.92None
5Opus 5.5withClaude73.657.32None
6Gemini 3.5 Flash-LitewithGemini60.357.721 capped
7Gemini 3.8 FlashwithAPI55.624.121 capped

About the task

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Checks the result can be trusted before interpreting it
  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Decision model and LLM judge, calibrated against a blind PM review