Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for checkout at Basketful. We want to test showing the delivery fee in the basket, instead of only at the last step of checkout. Write the experiment spec for Ravi, the growth engineer who'll build it, Ines, the analyst who'll read it out, and Chloe, our Head of Growth, who approves it. Keep it under 700 words. The team's draft plan is below. Fix what needs fixing.

What the model was given5 items: About Basketful, Why we're testing it, Traffic and baselines, Smallest change worth acting on, The team's draft plan
About BasketfulAn online grocery service. Delivery costs £3.99, or is free on orders over £60. Today the fee first appears on the final checkout step.
Why we're testing it12% of support contacts last quarter were about delivery fees the customer hadn't expected. Design believes showing the fee early will build trust; Chloe worries it will put people off before they start checkout.
Traffic and baselinesAbout 40,000 different customers view their basket each week. About half of each week's basket viewers are new that week; the rest came back from earlier weeks. 8% of basket viewers place an order. Of customers who start checkout, 62% complete it. Average order value: £47.
Smallest change worth acting onChloe: “Anything smaller than half a percentage point on orders isn't worth arguing about.”
The team's draft planPrimary metric: checkout completion rate (orders ÷ checkout starts). Split: 50/50 by session. Duration: we'll check the dashboard every day and stop as soon as it's significant. Chloe's addition: “While we're at it, let's put the free-delivery-over-£60 banner in the variant too.”
What a strong answer doesThe answer key the graders mark against

A spec that tests one change and fixes the draft's three problems. The primary metric is orders per basket-viewing customer, not checkout completion rate: showing the fee earlier changes who starts checkout, so completion can rise while orders fall. Customers, not sessions, are randomised, so returning customers see one version. The free-delivery banner stays out of the variant (or gets its own test), or the result can't be attributed. The duration is fixed in advance, with no daily stopping: detecting 0.5pp on an 8% base needs roughly 47,000 customers per arm (about 95,000 in total), and because half of each week's viewers are returning customers, new customers accumulate at about 40,000, 60,000, 80,000 and 100,000 after four weeks, so it runs about four full weeks, not two and a half. Guardrails with thresholds: average order value, fee-related support contacts, and the checkout-start rate as a diagnostic. A sample ratio check before reading results. A decision rule written up front: ship if orders are no worse than 0.5pp down and fee contacts fall; don't ship if orders drop by 0.5pp or more; with a stated plan for an inconclusive result.

Critical failures (cap the score at 40)
  • Uses checkout completion rate (orders ÷ checkout starts) as the primary metric
  • Bundles the free-delivery banner into the variant
Case

Mid-level · v1.1 · synthetic · B2C, ecommerce, pricing, metric choice

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: Bundles the free-delivery banner into the variant. Score capped at 40.

Needs more than a quick edit1 of 12 checks passed

The spec bundles the banner, uses an inadequate duration, and lacks a decision rule and guardrail thresholds, so a PM would need to redo core parts.

Output

To: Chloe (Head of Growth), Ravi (Growth Engineer), Ines (Data Analyst) From: PM, Checkout Subject: Experiment Spec: Early Delivery Fee Display in Basket

Team, here is the polished spec for our upcoming checkout experiment. I’ve incorporated Chloe’s threshold, fixed a few methodological traps in the draft, and scoped this for a clean, reliable read.

---

1. Objective & Hypothesis

  • Current State: Delivery fee (£3.99, free over £60) is only revealed on the final checkout step. 12% of last quarter’s support contacts were fee-related surprises.
  • Hypothesis: Showing the delivery fee (and progress to free delivery) directly in the basket increases customer trust and sets clear expectations upfront.
  • Expected Impact: While Chloe rightly worries this might cause some upfront drop-off among sub-£60 baskets, we believe the resulting higher-intent checkout starts will lift overall conversion and reduce cart abandonment due to "sticker shock."

---

2. Variant Definition

  • Control (A): Existing basket experience. Delivery fee hidden until final checkout step.
  • Variant (B):
  • Displays the delivery fee clearly in the basket summary.
  • Note on scope: Per Chloe’s suggestion, we are bundling the "free delivery over £60" progress banner into Variant B. Since the fee and the threshold are conceptually tied, showing one without the context of the other would skew the test.

---

3. Traffic Allocation & Unit of Randomization

  • Unit of Randomization: User-level (logged-in state or persistent cookie). Correction from draft: Session-level splitting would cause a jarring experience if a user refreshes or returns and sees fluctuating fee visibility.
  • Split: 50/50 (Control / Variant).
  • Audience: All customers viewing the basket (~40,000 unique users/week, split roughly 50/50 between new and returning).

---

4. Metrics

  • Primary Metric (Guardrail for Business Value):
  • Order Conversion Rate: Completed Orders ÷ Total Basket Views.
  • (Correction from draft: The draft proposed "Checkout Completion Rate (Orders ÷ Checkout Starts)." That misses top-of-funnel drop-off. If Variant B scares people away before checkout, measuring only checkout-starts will falsely make the variant look better. We must measure from the basket view).
  • Secondary Metrics:
  • Checkout Start Rate: Checkout Starts ÷ Basket Views.
  • Checkout Completion Rate: Orders ÷ Checkout Starts (to isolate downstream behavior).
  • Average Order Value (AOV): To monitor if the £60 threshold banner successfully drives basket-building behavior (baseline: £47).
  • Support Metric:
  • Volume of delivery-fee-related support tickets (Ines, let’s pull this baseline post-launch).

---

5. Success Criteria & Statistical Guardrails

  • Minimum Detectable Effect (MDE): Per Chloe, we care about a ≥ +0.5 percentage point absolute change in the Primary Metric (Order Conversion Rate from basket view, baseline ~8%).
  • Duration & Peeking:
  • Correction from draft: Do not stop the test the moment p < 0.05. Early peeking yields false positives.
  • We will run the test for a fixed duration of 2 full weeks to capture day-of-week and new vs. returning user cycles (~80,000 total unique users), yielding adequate statistical power for our MDE.

---

Next Steps:

  • Ravi: Let’s sync on tracking implementation for the basket-level fee and banner. Target build completion: [Insert Date].
  • Ines: Please set up the tracking dashboard, ensuring we segment by New vs. Returning users and Basket Value (<£60 vs. >£60).
  • Chloe: Good to approve?

Claims the judge couldn’t find in the brief

  • 2 full weeks yields ~80,000 total unique users and adequate statistical power for the MDE.

Grades and run details

Decision model 33 · LLM judge 2 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly10%
  • failAddresses the actual decision65%
  • failRespects explicit constraints16%
  • partialIdentifies material uncertainty73%
  • partialAvoids unsupported claims30%
  • partialProduces the required deliverable75%
  • failTests one change at a time99%
  • partialFixes the duration up front76%
  • partialAn unambiguous primary metric69%
  • partialDecision rule written before the test47%
  • partialSized from the real traffic69%
  • partialGuardrails with thresholds45%
Run
Run
#1
Time to output
5 s
Submitted
30 Sept 2026

Needs more than a quick edit8 of 12 checks passed

The spec exceeds the 700-word limit and lacks guardrail thresholds, so a PM would need to trim it and add those thresholds before use.

Output

Experiment Spec: Delivery Fee Shown in Basket

Owner: Checkout PM · Build: Ravi · Analysis: Ines · Approver: Chloe

Why we're testing this

Today the £3.99 delivery fee (free over £60) first appears on the final checkout step. Last quarter, 12% of support contacts were about unexpected delivery fees. Design expects that showing the fee earlier will build trust. Chloe's concern is that it will put people off before they start checkout. This test tells us which effect is bigger.

Hypothesis

Showing the delivery fee in the basket will change the share of basket viewers who place an order. We don't assume a direction. The test is two-sided because either outcome is plausible.

Variants

  • Control: Current experience. The fee appears only at the final checkout step.
  • Variant: The basket shows a "Delivery: £3.99" line, or "Free" when the basket is over £60. The same fee still appears at checkout.

Change from draft: The free-delivery-over-£60 banner is removed from this test. If both changes ship together, we can't tell which one caused any result. The banner could also raise orders while the fee line lowers them, so the two effects would cancel and we'd learn nothing. We'll test the banner as a follow-up once this result is in.

Primary metric

Order conversion: the share of customers who view their basket during the test and go on to place at least one order.

Change from draft: The draft used checkout completion (orders ÷ checkout starts). That metric would be misleading here. If the variant puts off fee-sensitive customers before checkout, the people who remain in checkout are the ones already willing to pay. Completion would then rise even if total orders fell. Order conversion counts everyone who saw the basket, so it catches losses at either stage.

Secondary and guardrail metrics

  • Checkout start rate and checkout completion rate. These are diagnostic only and show where any change happens.
  • Average order value. Seeing the fee may push some customers over £60.
  • Revenue per basket viewer. This includes delivery fee income.
  • Fee-related support contacts per 1,000 orders. This is the problem we set out to fix.

Randomisation

  • Unit: customer (logged-in account ID). The draft split by session, which needs to change. About half of each week's basket viewers are returning customers. With a session split, they would see the fee in the basket on one visit and not on the next. That would contaminate both groups and confuse customers.
  • Split: 50/50, sticky for the whole test.
  • Ravi: Please log the assignment when a customer first views their basket, and confirm that logged-out traffic is either excluded or bucketed consistently.

Sample size and duration

  • Assumptions: baseline 8% order conversion, two-sided test at α = 0.05, 80% power.
  • Minimum effect we want to detect: 0.5 percentage points (8.0% → 8.5%). We're reading Chloe's threshold as absolute points on order conversion. That is roughly a 6% relative change.
  • Customers needed: about 47,500 per arm, or about 95,000 total.

How long that takes: About 40,000 customers view their basket in week 1, then about 20,000 new customers each week after. That gives roughly 40k, 60k, 80k and 100k cumulative customers. The test therefore runs for 4 full weeks.

Caveat for Ines: The 8% baseline is per basket viewer per week. Over a 4-week window, each customer has more chances to order, so the per-customer baseline will be higher. At 12%, for example, we'd need about 135,000 customers, which means about 6 weeks. Before launch, Ines will calculate the 4-week per-customer baseline from historical data. The final duration will be fixed from that number and recorded here.

Stopping rule

Change from draft: We won't stop the test as soon as the dashboard shows significance. If we check every day and stop on the first significant result, the false-positive rate climbs well above 5%. The dashboard will often show early "wins" that aren't real.

  • The test runs for the pre-set duration, in full weeks only.
  • The dashboard can be watched for bugs and sample-ratio mismatch. Nobody makes ship decisions from it mid-test.
  • Only reason to stop early: a guardrail failure. That means a broken experience, a sample-ratio mismatch (p < 0.001), or order conversion down by more than 1.5 points.

Decision rules

The confidence interval (CI) below is the interval around the difference in order conversion between the variant and control.

Result on order conversionDecision
Up by a significant amountShip the variant
Down by a significant amountDon't ship. Look into other ways to reduce the support contacts
Not significant, and the CI excludes −0.5 pointsShip. We lose no meaningful number of orders and gain trust and fewer support contacts
Not significant, and the CI includes −0.5 pointsTreat as inconclusive. Chloe decides whether to extend the test or drop the change

Sign-off

  • ☐ Ravi: build plus assignment logging
  • ☐ Ines: baseline check and final duration
  • ☐ Chloe: approval

Grades and run details

Decision model 83 · LLM judge 9 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly7%
  • passAddresses the actual decision73%
  • partialRespects explicit constraints16%
  • passIdentifies material uncertainty90%
  • passAvoids unsupported claims45%
  • passProduces the required deliverable59%
  • passTests one change at a time100%
  • passFixes the duration up front98%
  • passAn unambiguous primary metric39%
  • passDecision rule written before the test94%
  • passSized from the real traffic55%
  • partialGuardrails with thresholds84%
Run
Run
#1
Time to output
31 s
Submitted
30 Sept 2026

Needs more than a quick edit10 of 12 checks passed

Guardrail thresholds are vague ('materially worse', 'drops meaningfully'), so a PM would need to define them before the test can run.

Output

Experiment spec: delivery fee in basket

Owner: PM, Checkout | Readers: Ravi (build), Ines (analysis), Chloe (approval)

Hypothesis

Showing the delivery fee in the basket reduces surprise at checkout (12% of last quarter's support contacts) without reducing orders. Chloe's concern is that an early fee deters people before they start checkout, so the test measures that directly.

Changes from the team draft

DraftProblemFix
Primary metric: checkout completionIt only counts people who already started checkout. If the early fee puts people off, they never enter the denominator, so the harm would be hidden.Primary metric is orders per basket viewer.
Split by sessionReturning customers (about half of weekly viewers) would see both versions.Randomise by customer.
Check daily, stop when significantRepeated peeking inflates false positives.Fixed sample size and duration, one readout.
Add the free-delivery banner to the variantWe couldn't tell whether the fee or the banner caused any change.Banner excluded. Test it separately afterwards.

Design

  • Control: fee first shown at final checkout step (today).
  • Variant: basket shows the delivery line: £3.99, or "Free" at £60+. Nothing else changes. The basket and checkout amounts must always match.
  • Unit: customer ID, with a 50/50 split. Assignment is sticky for the whole test. Logged-out users are bucketed by a persistent device ID (Ravi to confirm how many there are).
  • Enrolment: at first basket view during the test. Log an exposure event at that moment.

Metrics

Primary: share of enrolled customers who place at least one order within 7 days of first basket view. Baseline is about 8%.

Guardrails: - Checkout start rate (baseline about 12.9%, i.e. 8% ÷ 62%). - Revenue per enrolled customer (baseline about £3.76, i.e. 8% × £47). - Delivery-fee support contacts per 1,000 customers. This is the benefit metric, but volumes are low, so read it directionally.

Diagnostics (not decision metrics): - Checkout completion (orders ÷ checkout starts). Read it only alongside the primary metric, since it's conditional on starts. - New vs returning customers, and baskets under vs over £60. Both are exploratory.

Sample size and duration

  • Smallest effect worth detecting: 0.5 percentage points on orders (8.0% → 8.5%), per Chloe.
  • Requirement: 80% power, 5% two-sided significance. This needs about 46,000 customers per arm, roughly 92,000 in total.
  • Time to reach it: Week 1 brings about 40,000 unique customers. After that, only about 20,000 new customers enrol per week, because returning customers are already assigned. That reaches about 100,000 by the end of week 4.
  • Duration: enrol for 4 full weeks, plus 7 days of follow-up for the last cohort. Readout is at about 5 weeks, with no early stopping. Ines can monitor data quality but not results.
  • Limit: the test can't reliably detect effects smaller than 0.5pp. A "no significant difference" result means "no effect larger than about 0.5pp", not "no effect".

Data quality checks (before readout)

  • Sample ratio check: arms should be within normal variation of 50/50.
  • Exposure is logged for every enrolled customer.
  • Fee shown in basket equals fee charged, in all cases.

Decision rule (Chloe to confirm)

  • Ship if the 95% confidence interval for the change in orders rules out a drop of 0.5pp or more, no guardrail is materially worse, and delivery-fee contacts do not rise.
  • Don't ship if orders fall by 0.5pp or more, or revenue per customer drops meaningfully.
  • Otherwise, treat as inconclusive and discuss. Do not extend the test to "find" significance.

Next

If the variant ships or is neutral, run a follow-up test of the free-delivery-over-£60 banner, where the fee is already visible. Its effect is likely concentrated in baskets just under £60.

Open items

  • Ravi: logged-out identity handling; a flag to switch the variant off quickly.
  • Chloe: confirm the decision rule and that the 0.5pp threshold means absolute percentage points on orders per basket viewer.
  • Ines: confirm the 7-day order attribution and the support-contact tagging for fee-related contacts.

Grades and run details

Decision model 79 · LLM judge 11 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly4%
  • passAddresses the actual decision61%
  • partialRespects explicit constraints23%
  • passIdentifies material uncertainty80%
  • passAvoids unsupported claims35%
  • passProduces the required deliverable67%
  • passTests one change at a time99%
  • passFixes the duration up front100%
  • passAn unambiguous primary metric96%
  • passDecision rule written before the test95%
  • partialSized from the real traffic31%
  • partialGuardrails with thresholds94%
Run
Run
#1
API response time
29 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 10

Uses the supplied evidence correctlyWrongMixedMixed
Gemini 3.5 Flash-Lite · Gemini

The claim that 2 weeks yields ~80,000 unique users and adequate power is not supported by the brief; the brief's traffic and returning-user dynamics do not support that sample size or power.

Opus 5.5 · Claude

All claims about the current situation are directly from the brief or derived by correct arithmetic.

Sonnet 5.5 · API

All statements about the current situation are taken directly from the brief or derived by correct arithmetic.

Addresses the actual decisionMixedRightRight
Gemini 3.5 Flash-Lite · Gemini

The output commits to a specific experiment design and asks Chloe for approval, which is the deliverable requested.

Opus 5.5 · Claude

The spec commits to a clear decision framework with rules for shipping, not shipping, and inconclusive results, framed for Chloe.

Sonnet 5.5 · API

The spec commits to a clear decision rule with ship/don't ship/inconclusive actions, framed for Chloe to confirm.

Respects explicit constraintsWrongWrongRight
Gemini 3.5 Flash-Lite · Gemini

The brief says 'Fix what needs fixing' and the draft bundles the free-delivery banner; the output keeps it bundled, failing to fix that problem.

Opus 5.5 · Claude

The output is approximately 780 words, exceeding the 'under 700 words' limit.

Sonnet 5.5 · API

The output is under 700 words, addresses the named readers, and is a complete experiment spec.

Identifies material uncertaintyWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

The output does not name unknowns that could change the decision, nor does it specify what would happen if the result is inconclusive.

Opus 5.5 · Claude

It identifies the uncertain per-customer baseline over 4 weeks and the inconclusive outcome, with a plan to resolve the baseline.

Sonnet 5.5 · API

It names unknowns like logged-out identity, decision rule confirmation, and low support-contact volumes, and says how to resolve them.

Avoids unsupported claimsWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

The output presents the 2-week duration as yielding adequate power as fact, without labelling it as an assumption, and the evidence does not support it.

Opus 5.5 · Claude

Hypotheses and forecasts are clearly labelled, and no confident claims go beyond the supplied evidence.

Sonnet 5.5 · API

Hypotheses and forecasts are clearly labelled, and no interpretations are presented as established fact.

Produces the required deliverableWrongMixedRight
Gemini 3.5 Flash-Lite · Gemini

The spec is missing a decision rule, guardrail thresholds, proper sample-size justification, and a sample ratio check, so it is not usable as is.

Opus 5.5 · Claude

The spec is over the 700-word limit, so it does not fully meet the requested form.

Sonnet 5.5 · API

The spec is complete, in the right form, and usable by Ravi, Ines, and Chloe with light edits.

Tests one change at a timeWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

The spec bundles the free-delivery banner into the variant, arguing it is conceptually tied, instead of testing the fee change alone.

Opus 5.5 · Claude

The free-delivery banner is removed from the variant and the spec explains why bundling would make results uninterpretable.

Sonnet 5.5 · API

The free-delivery banner is excluded from the variant and a separate follow-up test is proposed, with the attribution problem explained.

An unambiguous primary metricWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

The spec names one primary metric but does not include a trust check such as a sample ratio check.

Opus 5.5 · Claude

Order conversion is the single primary metric with a clear rationale, and a sample-ratio mismatch check is included.

Sonnet 5.5 · API

The primary metric is orders per enrolled customer within 7 days, with a clear rationale and a planned sample ratio check.

Decision rule written before the testWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

No decision rule is stated; the spec does not map outcomes (including inconclusive) to actions with thresholds.

Opus 5.5 · Claude

All outcomes (significant up, significant down, non-significant with CI excluding -0.5pp, non-significant with CI including -0.5pp) map to stated actions.

Sonnet 5.5 · API

Every outcome (ship, don't ship, inconclusive) is mapped to an action with thresholds separating them.

Sized from the real trafficWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

The 2-week duration is not worked out from the traffic and MDE; the required sample size is much larger given the 8% baseline and 0.5pp effect.

Opus 5.5 · Claude

Sample size is calculated from the 8% baseline and 0.5pp effect, duration follows from traffic accumulation, and it runs full weeks.

Sonnet 5.5 · API

Sample size is correctly calculated from the 8% baseline and 0.5pp effect, and duration follows from the traffic pattern over whole weeks.

All got wrong 1

Guardrails with thresholdsWrongWrongWrong
Gemini 3.5 Flash-Lite · Gemini

Guardrail metrics are listed but no thresholds are set that would block a rollout.

Opus 5.5 · Claude

Guardrail metrics are listed but no thresholds are given for average order value, revenue, or support contacts to block rollout; only order conversion has a stop threshold.

Sonnet 5.5 · API

Guardrails are named but lack specific numeric thresholds; 'materially worse' and 'drops meaningfully' are too vague to block rollout unambiguously.

All got right 1

Fixes the duration up frontRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The spec replaces 'stop as soon as significant' with a fixed 2-week duration and explains that early peeking yields false positives.

Opus 5.5 · Claude

It sets a fixed 4-week duration and explains that daily peeking inflates false positives.

Sonnet 5.5 · API

Duration is fixed at 4 weeks enrollment plus 7 days follow-up, with no early stopping, and the false-positive risk of peeking is explained.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review