Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 9 graded outputs by 5 models. 56% were usable with at most a quick edit.

Reliably right

  1. A realistic plan that beats the freeze100% pass
    It converts 740 windows to about 9 days implicitly, explains that correlation and weekly cycles require a longer test (20 days, ~3 weeks), and provides dates that fit before 6 November with Porto notice given on 30 September.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Addresses the actual decision94% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  3. Identifies material uncertainty92% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. An unambiguous primary metric53% pass
    The spec names courier cost per order as the primary metric with a rationale, but it does not include a planned trust check such as a sample ratio check or verification that the arms receive the scheduled share of windows and similar order volumes.
    Sonnet 5.5 · API · Batching deliveries before peak season
  2. Guardrails with thresholds58% pass
    Guardrails are named but lack specific numeric thresholds; 'materially worse' and 'drops meaningfully' are too vague to block rollout unambiguously.
    Sonnet 5.5 · API · Showing the delivery fee up front
  3. Sized from the real traffic58% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for checkout at Basketful. We want to test showing the delivery fee in the basket, instead of only at the last step of checkout. Write the experiment spec for Ravi, the growth engineer who'll build it, Ines, the analyst who'll read it out, and Chloe, our Head of Growth, who approves it. Keep it under 700 words. The team's draft plan is below. Fix what needs fixing.

About BasketfulAn online grocery service. Delivery costs £3.99, or is free on orders over £60. Today the fee first appears on the final checkout step.
Why we're testing it12% of support contacts last quarter were about delivery fees the customer hadn't expected. Design believes showing the fee early will build trust; Chloe worries it will put people off before they start checkout.
Traffic and baselinesAbout 40,000 different customers view their basket each week. About half of each week's basket viewers are new that week; the rest came back from earlier weeks. 8% of basket viewers place an order. Of customers who start checkout, 62% complete it. Average order value: £47.
Smallest change worth acting onChloe: “Anything smaller than half a percentage point on orders isn't worth arguing about.”
The team's draft planPrimary metric: checkout completion rate (orders ÷ checkout starts). Split: 50/50 by session. Duration: we'll check the dashboard every day and stop as soon as it's significant. Chloe's addition: “While we're at it, let's put the free-delivery-over-£60 banner in the variant too.”
What a strong answer does

A spec that tests one change and fixes the draft's three problems. The primary metric is orders per basket-viewing customer, not checkout completion rate: showing the fee earlier changes who starts checkout, so completion can rise while orders fall. Customers, not sessions, are randomised, so returning customers see one version. The free-delivery banner stays out of the variant (or gets its own test), or the result can't be attributed. The duration is fixed in advance, with no daily stopping: detecting 0.5pp on an 8% base needs roughly 47,000 customers per arm (about 95,000 in total), and because half of each week's viewers are returning customers, new customers accumulate at about 40,000, 60,000, 80,000 and 100,000 after four weeks, so it runs about four full weeks, not two and a half. Guardrails with thresholds: average order value, fee-related support contacts, and the checkout-start rate as a diagnostic. A sample ratio check before reading results. A decision rule written up front: ship if orders are no worse than 0.5pp down and fee contacts fall; don't ship if orders drop by 0.5pp or more; with a stated plan for an inconclusive result.

Critical failures (cap the score)
  • Uses checkout completion rate (orders ÷ checkout starts) as the primary metric
  • Bundles the free-delivery banner into the variant
Case

v1.0 · synthetic · B2C, ecommerce, pricing, metric choice

Needs more than a quick edit11 of 12 checks passed

Guardrail thresholds are vague ('materially worse', 'drops meaningfully'), so a PM would need to define them before the test can run.

Output

Experiment spec: delivery fee in basket

Owner: PM, Checkout | Readers: Ravi (build), Ines (analysis), Chloe (approval)

Hypothesis

Showing the delivery fee in the basket reduces surprise at checkout (12% of last quarter's support contacts) without reducing orders. Chloe's concern is that an early fee deters people before they start checkout, so the test measures that directly.

Changes from the team draft

DraftProblemFix
Primary metric: checkout completionIt only counts people who already started checkout. If the early fee puts people off, they never enter the denominator, so the harm would be hidden.Primary metric is orders per basket viewer.
Split by sessionReturning customers (about half of weekly viewers) would see both versions.Randomise by customer.
Check daily, stop when significantRepeated peeking inflates false positives.Fixed sample size and duration, one readout.
Add the free-delivery banner to the variantWe couldn't tell whether the fee or the banner caused any change.Banner excluded. Test it separately afterwards.

Design

  • Control: fee first shown at final checkout step (today).
  • Variant: basket shows the delivery line: £3.99, or "Free" at £60+. Nothing else changes. The basket and checkout amounts must always match.
  • Unit: customer ID, with a 50/50 split. Assignment is sticky for the whole test. Logged-out users are bucketed by a persistent device ID (Ravi to confirm how many there are).
  • Enrolment: at first basket view during the test. Log an exposure event at that moment.

Metrics

Primary: share of enrolled customers who place at least one order within 7 days of first basket view. Baseline is about 8%.

Guardrails: - Checkout start rate (baseline about 12.9%, i.e. 8% ÷ 62%). - Revenue per enrolled customer (baseline about £3.76, i.e. 8% × £47). - Delivery-fee support contacts per 1,000 customers. This is the benefit metric, but volumes are low, so read it directionally.

Diagnostics (not decision metrics): - Checkout completion (orders ÷ checkout starts). Read it only alongside the primary metric, since it's conditional on starts. - New vs returning customers, and baskets under vs over £60. Both are exploratory.

Sample size and duration

  • Smallest effect worth detecting: 0.5 percentage points on orders (8.0% → 8.5%), per Chloe.
  • Requirement: 80% power, 5% two-sided significance. This needs about 46,000 customers per arm, roughly 92,000 in total.
  • Time to reach it: Week 1 brings about 40,000 unique customers. After that, only about 20,000 new customers enrol per week, because returning customers are already assigned. That reaches about 100,000 by the end of week 4.
  • Duration: enrol for 4 full weeks, plus 7 days of follow-up for the last cohort. Readout is at about 5 weeks, with no early stopping. Ines can monitor data quality but not results.
  • Limit: the test can't reliably detect effects smaller than 0.5pp. A "no significant difference" result means "no effect larger than about 0.5pp", not "no effect".

Data quality checks (before readout)

  • Sample ratio check: arms should be within normal variation of 50/50.
  • Exposure is logged for every enrolled customer.
  • Fee shown in basket equals fee charged, in all cases.

Decision rule (Chloe to confirm)

  • Ship if the 95% confidence interval for the change in orders rules out a drop of 0.5pp or more, no guardrail is materially worse, and delivery-fee contacts do not rise.
  • Don't ship if orders fall by 0.5pp or more, or revenue per customer drops meaningfully.
  • Otherwise, treat as inconclusive and discuss. Do not extend the test to "find" significance.

Next

If the variant ships or is neutral, run a follow-up test of the free-delivery-over-£60 banner, where the fee is already visible. Its effect is likely concentrated in baskets just under £60.

Open items

  • Ravi: logged-out identity handling; a flag to switch the variant off quickly.
  • Chloe: confirm the decision rule and that the 0.5pp threshold means absolute percentage points on orders per basket viewer.
  • Ines: confirm the 7-day order attribution and the support-contact tagging for fee-related contacts.

Check by check

Got wrong · 1
  • Guardrails with thresholdsGuardrails are named but lack specific numeric thresholds; 'materially worse' and 'drops meaningfully' are too vague to block rollout unambiguously.
Got right · 11
  • Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the brief or derived by correct arithmetic.
  • Addresses the actual decisionThe spec commits to a clear decision rule with ship/don't ship/inconclusive actions, framed for Chloe to confirm.
  • Respects explicit constraintsThe output is under 700 words, addresses the named readers, and is a complete experiment spec.
  • Identifies material uncertaintyIt names unknowns like logged-out identity, decision rule confirmation, and low support-contact volumes, and says how to resolve them.
  • Avoids unsupported claimsHypotheses and forecasts are clearly labelled, and no interpretations are presented as established fact.
  • Produces the required deliverableThe spec is complete, in the right form, and usable by Ravi, Ines, and Chloe with light edits.
  • Tests one change at a timeThe free-delivery banner is excluded from the variant and a separate follow-up test is proposed, with the attribution problem explained.
  • Fixes the duration up frontDuration is fixed at 4 weeks enrollment plus 7 days follow-up, with no early stopping, and the false-positive risk of peeking is explained.
  • An unambiguous primary metricThe primary metric is orders per enrolled customer within 7 days, with a clear rationale and a planned sample ratio check.
  • Decision rule written before the testEvery outcome (ship, don't ship, inconclusive) is mapped to an action with thresholds separating them.
  • Sized from the real trafficSample size is correctly calculated from the 8% baseline and 0.5pp effect, and duration follows from the traffic pattern over whole weeks.

Grades and run details

Decision model 92 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly5%
  • passAddresses the actual decision62%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty81%
  • passAvoids unsupported claims34%
  • passProduces the required deliverable72%
  • passTests one change at a time98%
  • passFixes the duration up front100%
  • passAn unambiguous primary metric82%
  • passDecision rule written before the test91%
  • partialSized from the real traffic48%
  • partialGuardrails with thresholds95%
Run
Run
#1
API response time
29 s
Submitted
30 Sept 2026

Usable with a quick edit9 of 12 checks passed

The core reasoning and evidence are correct; a quick edit to add a sample ratio check, guardrail thresholds and a one-line explanation for no peeking would make it fully ready.

Output

Basket delivery-fee experiment

For: Ravi (build), Ines (analysis), Chloe (approval) Decision requested: Approve a test of showing the existing delivery fee in the basket.

Hypothesis and scope

Showing the delivery cost before checkout will reduce surprise and fee-related support contacts, without materially reducing orders.

  • Control: Current experience; delivery fee first appears on the final checkout step.
  • Variant: Show the applicable fee in the basket: “Delivery £3.99” below the free-delivery threshold, or “Free delivery” when the basket qualifies. Keep the existing “free on orders over £60” rule and wording.
  • Change nothing else. Do not add the free-delivery banner: that would test a second intervention and make the result hard to interpret.

Build and assignment — Ravi

Randomise 50/50 by customer, not session, and keep each customer in the same arm across visits and devices where identity is available. Stratify assignment by new versus returning basket viewers. Exclude staff and test accounts. Log assignment, basket value/fee shown, basket views, checkout starts, orders, order value, and relevant support contacts.

Measurement — Ines

Primary metric: Basket-viewer conversion: percentage of assigned customers who place at least one order within 7 days of their first eligible basket view. Count each customer once, in their assigned arm (intention-to-treat). This captures any deterrence before checkout; checkout completion alone would miss it.

Guardrails / secondary measures: Checkout starts per basket viewer; checkout completion (orders ÷ checkout starts); average order value and revenue per assigned customer; delivery-fee-related support contacts per assigned customer. Support contacts may be too sparse for a definitive test, so treat them as directional, not a substitute for the primary outcome.

Sample, duration, and analysis

Baseline basket-viewer conversion is 8%. Power the test to detect a 0.5 percentage-point absolute difference (Chloe’s smallest worthwhile order effect), with 80% power and a two-sided 5% significance level. This requires approximately 95,000 unique customers total (about 47,500 per arm), assuming independent customer-level outcomes.

There are about 40,000 basket viewers weekly, but many return across weeks. Run for at least four full weeks and until the sample target is reached; then allow 7 days for outcomes to mature before reading results. If the target is not reached, extend in full weeks. Ines should confirm the final sample-size calculation and report actual enrolment.

Do not stop early for significance or check daily results to make a stopping decision. Monitor only for implementation or customer-harm issues. Analyse the primary metric at the customer level, with a 95% confidence interval and the pre-specified two-sided test.

Decision

Recommend rollout if the variant improves basket-viewer conversion by at least 0.5 percentage points, the result is statistically significant, and guardrails show no material harm. Do not roll out if there is significant harm. If results are inconclusive, report that clearly; do not interpret “not significant” as proof of no effect.

Check by check

Got wrong · 2
  • An unambiguous primary metricIt names one primary metric with a rationale but omits a planned trust check such as a sample ratio check.
  • Guardrails with thresholdsGuardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
Mixed · 1
  • Fixes the duration up frontIt sets a fixed duration but does not explain why daily peeking inflates false positives, only instructs not to do it.The two graders disagreed on this one.
Got right · 9
  • Uses the supplied evidence correctlyEvery statement about the current situation is taken directly from the supplied context or follows by arithmetic.
  • Addresses the actual decisionThe output clearly requests approval of the test and provides a complete spec for Chloe to decide, with a post-test decision rule.
  • Respects explicit constraintsThe output fixes all three problems in the draft, stays under 700 words, and is addressed to the three named readers.
  • Identifies material uncertaintyIt notes support contacts may be too sparse for a definitive test and that inconclusive results are possible, with a plan for each.
  • Avoids unsupported claimsThe hypothesis is clearly labelled, and no interpretations are presented as established fact.
  • Produces the required deliverableThe spec is complete, under 700 words, and directly usable by Ravi, Ines and Chloe.
  • Tests one change at a timeIt explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
  • Decision rule written before the testIt maps rollout, no rollout and inconclusive results to actions, with thresholds of 0.5pp improvement and statistical significance.
  • Sized from the real trafficSample size is calculated from the 8% baseline and 0.5pp effect, and the four-week duration accounts for returning visitors and whole-week cycles.

Grades and run details

Decision model 79 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly12%
  • passAddresses the actual decision62%
  • partialRespects explicit constraints28%
  • passIdentifies material uncertainty35%
  • passAvoids unsupported claims66%
  • passProduces the required deliverable62%
  • passTests one change at a time95%
  • passFixes the duration up front56%
  • partialAn unambiguous primary metric83%
  • partialDecision rule written before the test28%
  • partialSized from the real traffic24%
  • partialGuardrails with thresholds98%
Run
Run
#1
API response time
22 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI85.488.52None
2Sonnet 5.5withAPI87.588.52None
3Opus 5.5withClaude79.276.92None
4GPT-6 LunawithAPI81.373.12None
5Gemini 3.5 Flash-LitewithGemini33.315.411 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review