Analyse experiment results
Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?
The PM job
Reading a test readout and deciding what to do next.
Why it matters
Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.
What good looks like
- Reports effects with their uncertainty
- Checks guardrail metrics before declaring a winner
- Separates what the data shows from plausible explanations
- Recommends a next step proportionate to the evidence
Deliberately not measured
- Re-running the statistics from raw data
- Chart production
Interpreting results within their limits
Claims causality, ignores guardrails or recommends arbitrary testing
Decision model, LLM judge and blind PM review
Results
Every setup we’ve tested on this task, across all cases and repeats.
| # | Model · Harness | Task score | Decision model | LLM judge | PM review | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.
We tested a shorter onboarding checklist and the result came back 'not significant'. The team wants to call it a failure and move on. Write a short note to the onboarding team saying what we should conclude from this test and what we should do next. Keep it under 300 words.
Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.
- Concludes the change has no effect
v1.2 · synthetic · null result, onboarding