Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 9 graded outputs by 5 models. 56% were usable with at most a quick edit.

Reliably right

  1. A realistic plan that beats the freeze100% pass
    It converts 740 windows to about 9 days implicitly, explains that correlation and weekly cycles require a longer test (20 days, ~3 weeks), and provides dates that fit before 6 November with Porto notice given on 30 September.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Addresses the actual decision94% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  3. Identifies material uncertainty92% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. An unambiguous primary metric53% pass
    The spec names courier cost per order as the primary metric with a rationale, but it does not include a planned trust check such as a sample ratio check or verification that the arms receive the scheduled share of windows and similar order volumes.
    Sonnet 5.5 · API · Batching deliveries before peak season
  2. Guardrails with thresholds58% pass
    Guardrails are named but lack specific numeric thresholds; 'materially worse' and 'drops meaningfully' are too vague to block rollout unambiguously.
    Sonnet 5.5 · API · Showing the delivery fee up front
  3. Sized from the real traffic58% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for checkout at Basketful. We want to test showing the delivery fee in the basket, instead of only at the last step of checkout. Write the experiment spec for Ravi, the growth engineer who'll build it, Ines, the analyst who'll read it out, and Chloe, our Head of Growth, who approves it. Keep it under 700 words. The team's draft plan is below. Fix what needs fixing.

About BasketfulAn online grocery service. Delivery costs £3.99, or is free on orders over £60. Today the fee first appears on the final checkout step.
Why we're testing it12% of support contacts last quarter were about delivery fees the customer hadn't expected. Design believes showing the fee early will build trust; Chloe worries it will put people off before they start checkout.
Traffic and baselinesAbout 40,000 different customers view their basket each week. About half of each week's basket viewers are new that week; the rest came back from earlier weeks. 8% of basket viewers place an order. Of customers who start checkout, 62% complete it. Average order value: £47.
Smallest change worth acting onChloe: “Anything smaller than half a percentage point on orders isn't worth arguing about.”
The team's draft planPrimary metric: checkout completion rate (orders ÷ checkout starts). Split: 50/50 by session. Duration: we'll check the dashboard every day and stop as soon as it's significant. Chloe's addition: “While we're at it, let's put the free-delivery-over-£60 banner in the variant too.”
What a strong answer does

A spec that tests one change and fixes the draft's three problems. The primary metric is orders per basket-viewing customer, not checkout completion rate: showing the fee earlier changes who starts checkout, so completion can rise while orders fall. Customers, not sessions, are randomised, so returning customers see one version. The free-delivery banner stays out of the variant (or gets its own test), or the result can't be attributed. The duration is fixed in advance, with no daily stopping: detecting 0.5pp on an 8% base needs roughly 47,000 customers per arm (about 95,000 in total), and because half of each week's viewers are returning customers, new customers accumulate at about 40,000, 60,000, 80,000 and 100,000 after four weeks, so it runs about four full weeks, not two and a half. Guardrails with thresholds: average order value, fee-related support contacts, and the checkout-start rate as a diagnostic. A sample ratio check before reading results. A decision rule written up front: ship if orders are no worse than 0.5pp down and fee contacts fall; don't ship if orders drop by 0.5pp or more; with a stated plan for an inconclusive result.

Critical failures (cap the score)
  • Uses checkout completion rate (orders ÷ checkout starts) as the primary metric
  • Bundles the free-delivery banner into the variant
Case

v1.0 · synthetic · B2C, ecommerce, pricing, metric choice

Critical failure: Bundles the free-delivery banner into the variant. Score capped at 40.

Needs more than a quick edit1 of 12 checks passed

The spec bundles the banner, uses an inadequate duration, and lacks a decision rule and guardrail thresholds, so a PM would need to redo core parts.

Output

To: Chloe (Head of Growth), Ravi (Growth Engineer), Ines (Data Analyst) From: PM, Checkout Subject: Experiment Spec: Early Delivery Fee Display in Basket

Team, here is the polished spec for our upcoming checkout experiment. I’ve incorporated Chloe’s threshold, fixed a few methodological traps in the draft, and scoped this for a clean, reliable read.

---

1. Objective & Hypothesis

  • Current State: Delivery fee (£3.99, free over £60) is only revealed on the final checkout step. 12% of last quarter’s support contacts were fee-related surprises.
  • Hypothesis: Showing the delivery fee (and progress to free delivery) directly in the basket increases customer trust and sets clear expectations upfront.
  • Expected Impact: While Chloe rightly worries this might cause some upfront drop-off among sub-£60 baskets, we believe the resulting higher-intent checkout starts will lift overall conversion and reduce cart abandonment due to "sticker shock."

---

2. Variant Definition

  • Control (A): Existing basket experience. Delivery fee hidden until final checkout step.
  • Variant (B):
  • Displays the delivery fee clearly in the basket summary.
  • Note on scope: Per Chloe’s suggestion, we are bundling the "free delivery over £60" progress banner into Variant B. Since the fee and the threshold are conceptually tied, showing one without the context of the other would skew the test.

---

3. Traffic Allocation & Unit of Randomization

  • Unit of Randomization: User-level (logged-in state or persistent cookie). Correction from draft: Session-level splitting would cause a jarring experience if a user refreshes or returns and sees fluctuating fee visibility.
  • Split: 50/50 (Control / Variant).
  • Audience: All customers viewing the basket (~40,000 unique users/week, split roughly 50/50 between new and returning).

---

4. Metrics

  • Primary Metric (Guardrail for Business Value):
  • Order Conversion Rate: Completed Orders ÷ Total Basket Views.
  • (Correction from draft: The draft proposed "Checkout Completion Rate (Orders ÷ Checkout Starts)." That misses top-of-funnel drop-off. If Variant B scares people away before checkout, measuring only checkout-starts will falsely make the variant look better. We must measure from the basket view).
  • Secondary Metrics:
  • Checkout Start Rate: Checkout Starts ÷ Basket Views.
  • Checkout Completion Rate: Orders ÷ Checkout Starts (to isolate downstream behavior).
  • Average Order Value (AOV): To monitor if the £60 threshold banner successfully drives basket-building behavior (baseline: £47).
  • Support Metric:
  • Volume of delivery-fee-related support tickets (Ines, let’s pull this baseline post-launch).

---

5. Success Criteria & Statistical Guardrails

  • Minimum Detectable Effect (MDE): Per Chloe, we care about a ≥ +0.5 percentage point absolute change in the Primary Metric (Order Conversion Rate from basket view, baseline ~8%).
  • Duration & Peeking:
  • Correction from draft: Do not stop the test the moment p < 0.05. Early peeking yields false positives.
  • We will run the test for a fixed duration of 2 full weeks to capture day-of-week and new vs. returning user cycles (~80,000 total unique users), yielding adequate statistical power for our MDE.

---

Next Steps:

  • Ravi: Let’s sync on tracking implementation for the basket-level fee and banner. Target build completion: [Insert Date].
  • Ines: Please set up the tracking dashboard, ensuring we segment by New vs. Returning users and Basket Value (<£60 vs. >£60).
  • Chloe: Good to approve?

Check by check

Got wrong · 10
  • Uses the supplied evidence correctlyThe claim that 2 weeks yields ~80,000 unique users and adequate power is not supported by the brief; the brief's traffic and returning-user dynamics do not support that sample size or power.
  • Respects explicit constraintsThe brief says 'Fix what needs fixing' and the draft bundles the free-delivery banner; the output keeps it bundled, failing to fix that problem.
  • Identifies material uncertaintyThe output does not name unknowns that could change the decision, nor does it specify what would happen if the result is inconclusive.
  • Avoids unsupported claimsThe output presents the 2-week duration as yielding adequate power as fact, without labelling it as an assumption, and the evidence does not support it.
  • Produces the required deliverableThe spec is missing a decision rule, guardrail thresholds, proper sample-size justification, and a sample ratio check, so it is not usable as is.
  • Tests one change at a timeThe spec bundles the free-delivery banner into the variant, arguing it is conceptually tied, instead of testing the fee change alone.
  • An unambiguous primary metricThe spec names one primary metric but does not include a trust check such as a sample ratio check.
  • Decision rule written before the testNo decision rule is stated; the spec does not map outcomes (including inconclusive) to actions with thresholds.
  • Sized from the real trafficThe 2-week duration is not worked out from the traffic and MDE; the required sample size is much larger given the 8% baseline and 0.5pp effect.
  • Guardrails with thresholdsGuardrail metrics are listed but no thresholds are set that would block a rollout.
Mixed · 1
  • Addresses the actual decisionThe output commits to a specific experiment design and asks Chloe for approval, which is the deliverable requested.The two graders disagreed on this one.
Got right · 1
  • Fixes the duration up frontThe spec replaces 'stop as soon as significant' with a fixed 2-week duration and explains that early peeking yields false positives.

Claims the judge couldn’t find in the brief

  • 2 full weeks yields ~80,000 total unique users and adequate statistical power for the MDE.

Grades and run details

Decision model 33 · LLM judge 2 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly18%
  • failAddresses the actual decision66%
  • failRespects explicit constraints73%
  • partialIdentifies material uncertainty76%
  • partialAvoids unsupported claims29%
  • partialProduces the required deliverable73%
  • failTests one change at a time99%
  • partialFixes the duration up front81%
  • partialAn unambiguous primary metric84%
  • partialDecision rule written before the test44%
  • partialSized from the real traffic77%
  • partialGuardrails with thresholds44%
Run
Run
#1
Time to output
5 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI85.488.52None
2Sonnet 5.5withAPI87.588.52None
3Opus 5.5withClaude79.276.92None
4GPT-6 LunawithAPI81.373.12None
5Gemini 3.5 Flash-LitewithGemini33.315.411 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review