Analyse experiment results
Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?
The PM job
Reading a test readout and deciding what to do next.
Why it matters
Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.
What good looks like
- Reports effects with their uncertainty
- Checks guardrail metrics before declaring a winner
- Separates what the data shows from plausible explanations
- Recommends a next step proportionate to the evidence
Deliberately not measured
- Re-running the statistics from raw data
- Chart production
Interpreting results within their limits
Claims causality, ignores guardrails or recommends arbitrary testing
Decision model, LLM judge and blind PM review
Results
Every setup we’ve tested on this task, across all cases and repeats.
| # | Model · Harness | Task score | Decision model | LLM judge | PM review | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then put up to three outputs side by side. The outputs are the point; the scores just tell you where to look.
Our new pricing page variant (B) increased trial-to-paid conversion, and the growth team wants to ship it to 100% on Monday. Review the readout below and write your recommendation to Priya Nair, Head of Growth, who makes the call. Lead with what we should do, then give the reasoning she needs to defend it in the Monday meeting. Keep it under 400 words.
Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.
- Recommends shipping B to 100%
- States the mechanism as established fact
v1.3 · anonymised real · pricing, guardrail breach, B2C SaaS