Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 9 graded outputs by 5 models. 56% were usable with at most a quick edit.

Reliably right

  1. A realistic plan that beats the freeze100% pass
    It converts 740 windows to about 9 days implicitly, explains that correlation and weekly cycles require a longer test (20 days, ~3 weeks), and provides dates that fit before 6 November with Porto notice given on 30 September.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Addresses the actual decision94% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  3. Identifies material uncertainty92% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. An unambiguous primary metric53% pass
    The spec names courier cost per order as the primary metric with a rationale, but it does not include a planned trust check such as a sample ratio check or verification that the arms receive the scheduled share of windows and similar order volumes.
    Sonnet 5.5 · API · Batching deliveries before peak season
  2. Guardrails with thresholds58% pass
    Guardrails are named but lack specific numeric thresholds; 'materially worse' and 'drops meaningfully' are too vague to block rollout unambiguously.
    Sonnet 5.5 · API · Showing the delivery fee up front
  3. Sized from the real traffic58% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for checkout at Basketful. We want to test showing the delivery fee in the basket, instead of only at the last step of checkout. Write the experiment spec for Ravi, the growth engineer who'll build it, Ines, the analyst who'll read it out, and Chloe, our Head of Growth, who approves it. Keep it under 700 words. The team's draft plan is below. Fix what needs fixing.

About BasketfulAn online grocery service. Delivery costs £3.99, or is free on orders over £60. Today the fee first appears on the final checkout step.
Why we're testing it12% of support contacts last quarter were about delivery fees the customer hadn't expected. Design believes showing the fee early will build trust; Chloe worries it will put people off before they start checkout.
Traffic and baselinesAbout 40,000 different customers view their basket each week. About half of each week's basket viewers are new that week; the rest came back from earlier weeks. 8% of basket viewers place an order. Of customers who start checkout, 62% complete it. Average order value: £47.
Smallest change worth acting onChloe: “Anything smaller than half a percentage point on orders isn't worth arguing about.”
The team's draft planPrimary metric: checkout completion rate (orders ÷ checkout starts). Split: 50/50 by session. Duration: we'll check the dashboard every day and stop as soon as it's significant. Chloe's addition: “While we're at it, let's put the free-delivery-over-£60 banner in the variant too.”
What a strong answer does

A spec that tests one change and fixes the draft's three problems. The primary metric is orders per basket-viewing customer, not checkout completion rate: showing the fee earlier changes who starts checkout, so completion can rise while orders fall. Customers, not sessions, are randomised, so returning customers see one version. The free-delivery banner stays out of the variant (or gets its own test), or the result can't be attributed. The duration is fixed in advance, with no daily stopping: detecting 0.5pp on an 8% base needs roughly 47,000 customers per arm (about 95,000 in total), and because half of each week's viewers are returning customers, new customers accumulate at about 40,000, 60,000, 80,000 and 100,000 after four weeks, so it runs about four full weeks, not two and a half. Guardrails with thresholds: average order value, fee-related support contacts, and the checkout-start rate as a diagnostic. A sample ratio check before reading results. A decision rule written up front: ship if orders are no worse than 0.5pp down and fee contacts fall; don't ship if orders drop by 0.5pp or more; with a stated plan for an inconclusive result.

Critical failures (cap the score)
  • Uses checkout completion rate (orders ÷ checkout starts) as the primary metric
  • Bundles the free-delivery banner into the variant
Case

v1.0 · synthetic · B2C, ecommerce, pricing, metric choice

Usable with a quick edit8 of 12 checks passed

The spec is well-structured and fixes the draft's main issues; a PM could quickly add guardrail thresholds and lock the duration to full weeks.

Output

Experiment spec: delivery-fee visibility in basket

Owner: Checkout PM Build: Ravi · Analysis: Ines · Approval: Chloe

Objective and hypothesis

Test whether showing delivery costs before checkout increases orders by reducing surprises—or decreases orders by discouraging customers earlier.

Last quarter, 12% of support contacts concerned unexpected delivery fees. This motivates the test but is not the percentage of customers affected.

Experience

Control: Current basket; delivery fee first appears at the final checkout step.

Variant: Show a clearly labelled delivery-fee line and fee-inclusive total in the basket. Update both when the basket changes. Use the production pricing logic: £3.99 delivery, free for orders over £60. Ravi must verify boundary behaviour at exactly £60 and match existing checkout pricing.

Keep the final checkout fee display unchanged. Everything else stays identical.

Do not add the free-delivery banner. It could independently encourage larger baskets, making it impossible to attribute results to earlier fee visibility. Test it separately.

Eligibility and assignment

Include customers viewing an orderable basket in supported markets. Apply existing employee, test-account and bot exclusions identically, before assignment.

Randomise 50/50 by customer, not session, at their first eligible basket view. Persist assignment across visits and devices where identity is available. Use a stable anonymous identifier for guests and a documented account-linking rule; Ravi and Ines must validate identity handling before launch.

Each customer enters the experiment once. Analyse everyone assigned, regardless of whether they start checkout or successfully see the treatment.

Metrics

Primary: Seven-day order conversion: customers placing at least one order within seven days of assignment ÷ all assigned eligible basket viewers.

This captures Chloe’s concern about discouraging checkout entry. Orders ÷ checkout starts is unsuitable as primary: the treatment can change who starts checkout.

Secondary/diagnostic: - Checkout-start rate per assigned customer. - Checkout completion among starters, interpreted cautiously because starter populations may differ. - Revenue per assigned customer, including non-buyers. - AOV and proportion of orders qualifying for free delivery. - Unexpected-fee support contacts per assigned customer within seven days, using consistent contact classification.

Operational guardrails: Basket/checkout errors, incorrect fee calculations and page performance. These can trigger a safety rollback, not an early declaration of success.

Sample size and duration

Interpret Chloe’s minimum worthwhile change as 0.5 percentage points absolute: 8.0% to 8.5%, not a 0.5% relative lift.

With a matching 8% baseline, two-sided 5% significance and 80% power, plan for approximately 95,000 unique customers total, evenly split. Ines must verify the historical baseline uses the same seven-day definition and finalise the calculation before launch.

Do not count 40,000 weekly viewers as entirely new experiment entrants. Estimated cumulative reach is 40,000 in week one, then roughly 20,000 additional customers weekly: about four weeks to enrol the target, plus seven days for outcomes to mature.

Lock the enrolment target before launch; stop enrolment when reached, then wait seven days. Extend only if unique-customer accrual is slower—not based on results.

Analysis and decision

Ines reports the intention-to-treat conversion difference, 95% confidence interval and absolute rates. Check allocation imbalance, logging, identity consistency and treatment exposure before interpreting results.

No daily significance-based stopping. Daily dashboards are for operational safety only.

Recommend rollout if the primary improvement is statistically significant, the point estimate is at least +0.5 percentage points, and there are no unresolved operational or economic concerns.

Otherwise, retain control pending review. Distinguish evidence of harm, evidence ruling out a worthwhile gain, and an inconclusive result; “not significant” does not mean “no effect.”

Launch approval

Ravi signs off implementation and QA; Ines signs off measurement and the locked analysis plan; Chloe approves the isolated treatment, decision rule and any economic guardrail thresholds before launch.

Check by check

Got wrong · 2
  • Sized from the real trafficSample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
  • Guardrails with thresholdsGuardrail metrics are listed but no explicit thresholds (e.g., for support contacts or AOV) are set to block rollout.
Mixed · 2
  • Uses the supplied evidence correctlyAll claims about the current situation are taken directly from the brief or follow logically from it.The two graders disagreed on this one.
  • Fixes the duration up frontIt does not set a fixed calendar duration in full weeks, and does not explain that daily stopping inflates false positives.The two graders disagreed on this one.
Got right · 8
  • Addresses the actual decisionThe spec includes a clear decision rule for rollout based on results, framed for Chloe.
  • Respects explicit constraintsThe output is under 700 words, addresses the named readers, and fixes the draft's problems.
  • Identifies material uncertaintyIt acknowledges inconclusive results and says to distinguish harm, no worthwhile gain, and uncertainty.
  • Avoids unsupported claimsNo interpretations or forecasts are presented as established facts; the 12% clarification is directly supported.
  • Produces the required deliverableThe spec is complete, in the right form, and usable by Ravi, Ines, and Chloe with light edits.
  • Tests one change at a timeIt explicitly keeps the free-delivery banner out of the variant and explains why bundling would confound results.
  • An unambiguous primary metricIt names one primary metric (seven-day order conversion), explains its direction, and includes a sample ratio check.
  • Decision rule written before the testIt maps significant positive result, harm, and inconclusive outcomes to actions (rollout or retain control) before the test.

Grades and run details

Decision model 79 · LLM judge 10 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly3%
  • passAddresses the actual decision68%
  • passRespects explicit constraints26%
  • passIdentifies material uncertainty53%
  • passAvoids unsupported claims64%
  • passProduces the required deliverable64%
  • passTests one change at a time100%
  • passFixes the duration up front75%
  • passAn unambiguous primary metric58%
  • partialDecision rule written before the test32%
  • partialSized from the real traffic56%
  • partialGuardrails with thresholds98%
Run
Run
#1
API response time
36 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI85.488.52None
2Sonnet 5.5withAPI87.588.52None
3Opus 5.5withClaude79.276.92None
4GPT-6 LunawithAPI81.373.12None
5Gemini 3.5 Flash-LitewithGemini33.315.411 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review