Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for checkout at Basketful. We want to test showing the delivery fee in the basket, instead of only at the last step of checkout. Write the experiment spec for Ravi, the growth engineer who'll build it, Ines, the analyst who'll read it out, and Chloe, our Head of Growth, who approves it. Keep it under 700 words. The team's draft plan is below. Fix what needs fixing.

What the model was given5 items: About Basketful, Why we're testing it, Traffic and baselines, Smallest change worth acting on, The team's draft plan
About BasketfulAn online grocery service. Delivery costs £3.99, or is free on orders over £60. Today the fee first appears on the final checkout step.
Why we're testing it12% of support contacts last quarter were about delivery fees the customer hadn't expected. Design believes showing the fee early will build trust; Chloe worries it will put people off before they start checkout.
Traffic and baselinesAbout 40,000 different customers view their basket each week. About half of each week's basket viewers are new that week; the rest came back from earlier weeks. 8% of basket viewers place an order. Of customers who start checkout, 62% complete it. Average order value: £47.
Smallest change worth acting onChloe: “Anything smaller than half a percentage point on orders isn't worth arguing about.”
The team's draft planPrimary metric: checkout completion rate (orders ÷ checkout starts). Split: 50/50 by session. Duration: we'll check the dashboard every day and stop as soon as it's significant. Chloe's addition: “While we're at it, let's put the free-delivery-over-£60 banner in the variant too.”
What a strong answer doesThe answer key the graders mark against

A spec that tests one change and fixes the draft's three problems. The primary metric is orders per basket-viewing customer, not checkout completion rate: showing the fee earlier changes who starts checkout, so completion can rise while orders fall. Customers, not sessions, are randomised, so returning customers see one version. The free-delivery banner stays out of the variant (or gets its own test), or the result can't be attributed. The duration is fixed in advance, with no daily stopping: detecting 0.5pp on an 8% base needs roughly 47,000 customers per arm (about 95,000 in total), and because half of each week's viewers are returning customers, new customers accumulate at about 40,000, 60,000, 80,000 and 100,000 after four weeks, so it runs about four full weeks, not two and a half. Guardrails with thresholds: average order value, fee-related support contacts, and the checkout-start rate as a diagnostic. A sample ratio check before reading results. A decision rule written up front: ship if orders are no worse than 0.5pp down and fee contacts fall; don't ship if orders drop by 0.5pp or more; with a stated plan for an inconclusive result.

Critical failures (cap the score at 40)
  • Uses checkout completion rate (orders ÷ checkout starts) as the primary metric
  • Bundles the free-delivery banner into the variant
Case

Mid-level · v1.1 · synthetic · B2C, ecommerce, pricing, metric choice

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit8 of 12 checks passed

The spec exceeds the 700-word limit and lacks guardrail thresholds, so a PM would need to trim it and add those thresholds before use.

Output

Experiment Spec: Delivery Fee Shown in Basket

Owner: Checkout PM · Build: Ravi · Analysis: Ines · Approver: Chloe

Why we're testing this

Today the £3.99 delivery fee (free over £60) first appears on the final checkout step. Last quarter, 12% of support contacts were about unexpected delivery fees. Design expects that showing the fee earlier will build trust. Chloe's concern is that it will put people off before they start checkout. This test tells us which effect is bigger.

Hypothesis

Showing the delivery fee in the basket will change the share of basket viewers who place an order. We don't assume a direction. The test is two-sided because either outcome is plausible.

Variants

  • Control: Current experience. The fee appears only at the final checkout step.
  • Variant: The basket shows a "Delivery: £3.99" line, or "Free" when the basket is over £60. The same fee still appears at checkout.

Change from draft: The free-delivery-over-£60 banner is removed from this test. If both changes ship together, we can't tell which one caused any result. The banner could also raise orders while the fee line lowers them, so the two effects would cancel and we'd learn nothing. We'll test the banner as a follow-up once this result is in.

Primary metric

Order conversion: the share of customers who view their basket during the test and go on to place at least one order.

Change from draft: The draft used checkout completion (orders ÷ checkout starts). That metric would be misleading here. If the variant puts off fee-sensitive customers before checkout, the people who remain in checkout are the ones already willing to pay. Completion would then rise even if total orders fell. Order conversion counts everyone who saw the basket, so it catches losses at either stage.

Secondary and guardrail metrics

  • Checkout start rate and checkout completion rate. These are diagnostic only and show where any change happens.
  • Average order value. Seeing the fee may push some customers over £60.
  • Revenue per basket viewer. This includes delivery fee income.
  • Fee-related support contacts per 1,000 orders. This is the problem we set out to fix.

Randomisation

  • Unit: customer (logged-in account ID). The draft split by session, which needs to change. About half of each week's basket viewers are returning customers. With a session split, they would see the fee in the basket on one visit and not on the next. That would contaminate both groups and confuse customers.
  • Split: 50/50, sticky for the whole test.
  • Ravi: Please log the assignment when a customer first views their basket, and confirm that logged-out traffic is either excluded or bucketed consistently.

Sample size and duration

  • Assumptions: baseline 8% order conversion, two-sided test at α = 0.05, 80% power.
  • Minimum effect we want to detect: 0.5 percentage points (8.0% → 8.5%). We're reading Chloe's threshold as absolute points on order conversion. That is roughly a 6% relative change.
  • Customers needed: about 47,500 per arm, or about 95,000 total.

How long that takes: About 40,000 customers view their basket in week 1, then about 20,000 new customers each week after. That gives roughly 40k, 60k, 80k and 100k cumulative customers. The test therefore runs for 4 full weeks.

Caveat for Ines: The 8% baseline is per basket viewer per week. Over a 4-week window, each customer has more chances to order, so the per-customer baseline will be higher. At 12%, for example, we'd need about 135,000 customers, which means about 6 weeks. Before launch, Ines will calculate the 4-week per-customer baseline from historical data. The final duration will be fixed from that number and recorded here.

Stopping rule

Change from draft: We won't stop the test as soon as the dashboard shows significance. If we check every day and stop on the first significant result, the false-positive rate climbs well above 5%. The dashboard will often show early "wins" that aren't real.

  • The test runs for the pre-set duration, in full weeks only.
  • The dashboard can be watched for bugs and sample-ratio mismatch. Nobody makes ship decisions from it mid-test.
  • Only reason to stop early: a guardrail failure. That means a broken experience, a sample-ratio mismatch (p < 0.001), or order conversion down by more than 1.5 points.

Decision rules

The confidence interval (CI) below is the interval around the difference in order conversion between the variant and control.

Result on order conversionDecision
Up by a significant amountShip the variant
Down by a significant amountDon't ship. Look into other ways to reduce the support contacts
Not significant, and the CI excludes −0.5 pointsShip. We lose no meaningful number of orders and gain trust and fewer support contacts
Not significant, and the CI includes −0.5 pointsTreat as inconclusive. Chloe decides whether to extend the test or drop the change

Sign-off

  • ☐ Ravi: build plus assignment logging
  • ☐ Ines: baseline check and final duration
  • ☐ Chloe: approval

Grades and run details

Decision model 83 · LLM judge 9 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly7%
  • passAddresses the actual decision73%
  • partialRespects explicit constraints16%
  • passIdentifies material uncertainty90%
  • passAvoids unsupported claims45%
  • passProduces the required deliverable59%
  • passTests one change at a time100%
  • passFixes the duration up front98%
  • passAn unambiguous primary metric39%
  • passDecision rule written before the test94%
  • passSized from the real traffic55%
  • partialGuardrails with thresholds84%
Run
Run
#1
Time to output
31 s
Submitted
30 Sept 2026

Needs more than a quick edit3 of 12 checks passed

A PM would need to redo the sizing and duration, remove or redesign the early kill rule, add a sample-ratio check, and complete guardrail thresholds. These are substantive reasoning changes, not quick wording edits.

Output

Experiment Spec: Delivery Fee Transparency in Basket

To: Chloe (Approval), Ravi (Build), Ines (Analysis) From: Checkout PM

---

1. Objective & Hypothesis

  • Context: 12% of support contacts complain about surprise delivery fees at final checkout.
  • Hypothesis: Displaying the delivery fee (£3.99, or "Free" if >£60) upfront in the basket will set accurate expectations, reduce support contacts, and build customer trust without reducing net orders.
  • Risk (Chloe’s concern): Upfront cost shock could cause drop-off before checkout begins.

---

2. Corrections to the Draft Plan

1. Primary Metric Changed: Draft used Checkout Completion Rate (Orders ÷ Starts). Correction: Use Basket-to-Order Conversion Rate (Orders ÷ Basket Viewers). Measuring only checkout starts introduces severe selection bias—if fee visibility deters low-intent users at the basket, checkout completion would artificially increase while overall orders plummet. 2. Unit of Randomization Changed: Draft used session-level. Correction: Randomize by User ID (falling back to persistent anonymous visitor ID). 50% of weekly viewers are repeat shoppers; session-level randomization would expose returning users to conflicting experiences. 3. No Peeking / Fixed Duration: Draft suggested stopping once "significant." Correction: Continuous monitoring with early stopping inflates false-positive rates to >30%. We will fix sample size upfront and run for complete weekly cycles. 4. Scope Control (Chloe’s Banner): Cut the promotional free-delivery banner from this test. Adding a promotional banner introduces a confounding variable, making it impossible to isolate whether fee transparency helped or hurt. We will test the banner in a follow-up experiment.

---

3. Variants (For Ravi)

  • Allocation: 50/50 split at User/Persistent Device ID level. Triggered on basket view.
  • Control (A): Current basket. No delivery fee shown (fee appears on final checkout step as today).
  • Variant (B): Add a delivery fee line item to the Basket Order Summary:
  • Basket < £60: `Delivery: £3.99`
  • Basket ≥ £60: `Delivery: FREE`
  • Include delivery fee in the displayed Basket Total.
  • Note: Do not add promotional progress bars or marketing banners.

---

4. Metrics Framework (For Ines)

  • Primary Metric: Basket-to-Order Conversion Rate (Total Unique Ordering Users ÷ Total Unique Basket Viewers).
  • Guardrail Metrics:
  • Average Order Value (AOV baseline: £47).
  • Basket-to-Checkout Start Rate (tracks where abandonment occurs).
  • Total Net Revenue (Orders × AOV).
  • Secondary / Diagnostic Metrics:
  • Delivery-fee-related customer support ticket rate per 1,000 orders.
  • Checkout Completion Rate (Checkout Starts → Orders).

---

5. Sample Size & Run Time (For Ines & Chloe)

  • Baseline Conversion: 8.0% of basket viewers place an order.
  • Minimum Detectable Effect (MDE): Absolute ±0.5 percentage points (7.5% or 8.5%, a 6.25% relative change), per Chloe's threshold.
  • Statistical Parameters: $\alpha = 0.05$ (two-tailed), Power ($1 - \beta$) = 80%.
  • Required Sample: ~47,000 unique users per variant (~94,000 total unique basket viewers).
  • Duration: 3 full weeks (21 days).
  • With 40,000 viewers/week (~20k new, ~20k repeat), 3 weeks yields ~80,000–90,000 unique basket viewers across three full grocery-shopping cycles.
  • Ines will run the read-out strictly after the 21-day window closes.

---

6. Decision & Rollout Criteria (For Chloe)

  • Ship Variant: Primary conversion does not drop by more than Chloe’s threshold of -0.5pp (non-inferiority) AND delivery-related support contacts decrease significantly.
  • Iterate / Re-test: Basket-to-order conversion drops by >0.5pp. In this scenario, early fee disclosure hurts conversion, justifying Chloe’s concern; we then test the £60 threshold banner/progress bar to offset drop-off.
  • Kill Immediately: Overall revenue drops >5% statistically significantly at the end of Week 1 (safety check only; not an early win stop).

Claims the judge couldn’t find in the brief

  • Three weeks yields about 80,000–90,000 unique basket viewers.
  • Early stopping inflates false-positive rates to more than 30%.
  • The variant's free-delivery threshold is £60 or more.

Grades and run details

Decision model 67 · LLM judge 3 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly55%
  • passAddresses the actual decision63%
  • partialRespects explicit constraints20%
  • passIdentifies material uncertainty65%
  • partialAvoids unsupported claims49%
  • partialProduces the required deliverable32%
  • passTests one change at a time100%
  • passFixes the duration up front66%
  • passAn unambiguous primary metric28%
  • partialDecision rule written before the test79%
  • partialSized from the real traffic94%
  • partialGuardrails with thresholds90%
Run
Run
#1
API response time
19 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 12 checks passed

The core reasoning and evidence are correct; a quick edit to add a sample ratio check, guardrail thresholds and a one-line explanation for no peeking would make it fully ready.

Output

Basket delivery-fee experiment

For: Ravi (build), Ines (analysis), Chloe (approval) Decision requested: Approve a test of showing the existing delivery fee in the basket.

Hypothesis and scope

Showing the delivery cost before checkout will reduce surprise and fee-related support contacts, without materially reducing orders.

  • Control: Current experience; delivery fee first appears on the final checkout step.
  • Variant: Show the applicable fee in the basket: “Delivery £3.99” below the free-delivery threshold, or “Free delivery” when the basket qualifies. Keep the existing “free on orders over £60” rule and wording.
  • Change nothing else. Do not add the free-delivery banner: that would test a second intervention and make the result hard to interpret.

Build and assignment — Ravi

Randomise 50/50 by customer, not session, and keep each customer in the same arm across visits and devices where identity is available. Stratify assignment by new versus returning basket viewers. Exclude staff and test accounts. Log assignment, basket value/fee shown, basket views, checkout starts, orders, order value, and relevant support contacts.

Measurement — Ines

Primary metric: Basket-viewer conversion: percentage of assigned customers who place at least one order within 7 days of their first eligible basket view. Count each customer once, in their assigned arm (intention-to-treat). This captures any deterrence before checkout; checkout completion alone would miss it.

Guardrails / secondary measures: Checkout starts per basket viewer; checkout completion (orders ÷ checkout starts); average order value and revenue per assigned customer; delivery-fee-related support contacts per assigned customer. Support contacts may be too sparse for a definitive test, so treat them as directional, not a substitute for the primary outcome.

Sample, duration, and analysis

Baseline basket-viewer conversion is 8%. Power the test to detect a 0.5 percentage-point absolute difference (Chloe’s smallest worthwhile order effect), with 80% power and a two-sided 5% significance level. This requires approximately 95,000 unique customers total (about 47,500 per arm), assuming independent customer-level outcomes.

There are about 40,000 basket viewers weekly, but many return across weeks. Run for at least four full weeks and until the sample target is reached; then allow 7 days for outcomes to mature before reading results. If the target is not reached, extend in full weeks. Ines should confirm the final sample-size calculation and report actual enrolment.

Do not stop early for significance or check daily results to make a stopping decision. Monitor only for implementation or customer-harm issues. Analyse the primary metric at the customer level, with a 95% confidence interval and the pre-specified two-sided test.

Decision

Recommend rollout if the variant improves basket-viewer conversion by at least 0.5 percentage points, the result is statistically significant, and guardrails show no material harm. Do not roll out if there is significant harm. If results are inconclusive, report that clearly; do not interpret “not significant” as proof of no effect.

Grades and run details

Decision model 88 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly14%
  • passAddresses the actual decision54%
  • partialRespects explicit constraints17%
  • passIdentifies material uncertainty35%
  • passAvoids unsupported claims64%
  • passProduces the required deliverable50%
  • passTests one change at a time94%
  • passFixes the duration up front62%
  • passAn unambiguous primary metric40%
  • partialDecision rule written before the test27%
  • passSized from the real traffic39%
  • partialGuardrails with thresholds98%
Run
Run
#1
API response time
22 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 8

Uses the supplied evidence correctlyMixedWrongRight
Opus 5.5 · Claude

All claims about the current situation are directly from the brief or derived by correct arithmetic.

Gemini 3.8 Flash · API

It misstates the free-delivery threshold as ≥£60, overstates unique viewers over three weeks, and adds an unsupported >30% false-positive claim.

GPT-6 Luna · API

Every statement about the current situation is taken directly from the supplied context or follows by arithmetic.

Respects explicit constraintsWrongWrongRight
Opus 5.5 · Claude

The output is approximately 780 words, exceeding the 'under 700 words' limit.

Gemini 3.8 Flash · API

It is under 700 words and addressed to the named readers, but it proposes a 21-day test that cannot reach the stated sample and includes a week-1 significance-based kill rule.

GPT-6 Luna · API

The output fixes all three problems in the draft, stays under 700 words, and is addressed to the three named readers.

Avoids unsupported claimsRightWrongRight
Opus 5.5 · Claude

Hypotheses and forecasts are clearly labelled, and no confident claims go beyond the supplied evidence.

Gemini 3.8 Flash · API

It presents the >30% false-positive rate and the 80,000–90,000 unique-viewer range as established facts without support.

GPT-6 Luna · API

The hypothesis is clearly labelled, and no interpretations are presented as established fact.

Produces the required deliverableMixedWrongRight
Opus 5.5 · Claude

The spec is over the 700-word limit, so it does not fully meet the requested form.

Gemini 3.8 Flash · API

The spec is usable in form but has major gaps: the duration is underpowered, the sample-ratio check is missing, and the guardrail thresholds are incomplete.

GPT-6 Luna · API

The spec is complete, under 700 words, and directly usable by Ravi, Ines and Chloe.

Fixes the duration up frontRightMixedMixed
Opus 5.5 · Claude

It sets a fixed 4-week duration and explains that daily peeking inflates false positives.

Gemini 3.8 Flash · API

It fixes a duration and explains peeking, but the duration is not derived from the required sample and it still allows a week-1 significance-based stop.

GPT-6 Luna · API

It sets a fixed duration but does not explain why daily peeking inflates false positives, only instructs not to do it.

An unambiguous primary metricRightMixedMixed
Opus 5.5 · Claude

Order conversion is the single primary metric with a clear rationale, and a sample-ratio mismatch check is included.

Gemini 3.8 Flash · API

It names one primary metric and rationale, but does not plan a sample-ratio check or other trustworthiness check.

GPT-6 Luna · API

It names one primary metric with a rationale but omits a planned trust check such as a sample ratio check.

Decision rule written before the testRightWrongRight
Opus 5.5 · Claude

All outcomes (significant up, significant down, non-significant with CI excluding -0.5pp, non-significant with CI including -0.5pp) map to stated actions.

Gemini 3.8 Flash · API

It does not define the action for an inconclusive result between the non-inferiority and harm thresholds.

GPT-6 Luna · API

It maps rollout, no rollout and inconclusive results to actions, with thresholds of 0.5pp improvement and statistical significance.

Sized from the real trafficRightWrongRight
Opus 5.5 · Claude

Sample size is calculated from the 8% baseline and 0.5pp effect, duration follows from traffic accumulation, and it runs full weeks.

Gemini 3.8 Flash · API

It calculates about 94,000 unique users but sets only three weeks, which yields about 80,000 unique users, not four weeks.

GPT-6 Luna · API

Sample size is calculated from the 8% baseline and 0.5pp effect, and the four-week duration accounts for returning visitors and whole-week cycles.

All got wrong 1

Guardrails with thresholdsWrongWrongWrong
Opus 5.5 · Claude

Guardrail metrics are listed but no thresholds are given for average order value, revenue, or support contacts to block rollout; only order conversion has a stop threshold.

Gemini 3.8 Flash · API

It names guardrails but gives thresholds only for revenue, not for AOV, checkout-start rate, or support contacts.

GPT-6 Luna · API

Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.

All got right 3

Addresses the actual decisionRightRightRight
Opus 5.5 · Claude

The spec commits to a clear decision framework with rules for shipping, not shipping, and inconclusive results, framed for Chloe.

Gemini 3.8 Flash · API

It commits to testing fee transparency alone, with ship, iterate, and kill rules tied to named thresholds.

GPT-6 Luna · API

The output clearly requests approval of the test and provides a complete spec for Chloe to decide, with a post-test decision rule.

Identifies material uncertaintyRightRightRight
Opus 5.5 · Claude

It identifies the uncertain per-customer baseline over 4 weeks and the inconclusive outcome, with a plan to resolve the baseline.

Gemini 3.8 Flash · API

It names the main uncertainty—whether early fee disclosure reduces orders—and maps outcomes to actions.

GPT-6 Luna · API

It notes support contacts may be too sparse for a definitive test and that inconclusive results are possible, with a plan for each.

Tests one change at a timeRightRightRight
Opus 5.5 · Claude

The free-delivery banner is removed from the variant and the spec explains why bundling would make results uninterpretable.

Gemini 3.8 Flash · API

It keeps the promotional free-delivery banner out of the variant and explains that bundling it would confound attribution.

GPT-6 Luna · API

It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review