Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for checkout at Basketful. We want to test showing the delivery fee in the basket, instead of only at the last step of checkout. Write the experiment spec for Ravi, the growth engineer who'll build it, Ines, the analyst who'll read it out, and Chloe, our Head of Growth, who approves it. Keep it under 700 words. The team's draft plan is below. Fix what needs fixing.

What the model was given5 items: About Basketful, Why we're testing it, Traffic and baselines, Smallest change worth acting on, The team's draft plan
About BasketfulAn online grocery service. Delivery costs £3.99, or is free on orders over £60. Today the fee first appears on the final checkout step.
Why we're testing it12% of support contacts last quarter were about delivery fees the customer hadn't expected. Design believes showing the fee early will build trust; Chloe worries it will put people off before they start checkout.
Traffic and baselinesAbout 40,000 different customers view their basket each week. About half of each week's basket viewers are new that week; the rest came back from earlier weeks. 8% of basket viewers place an order. Of customers who start checkout, 62% complete it. Average order value: £47.
Smallest change worth acting onChloe: “Anything smaller than half a percentage point on orders isn't worth arguing about.”
The team's draft planPrimary metric: checkout completion rate (orders ÷ checkout starts). Split: 50/50 by session. Duration: we'll check the dashboard every day and stop as soon as it's significant. Chloe's addition: “While we're at it, let's put the free-delivery-over-£60 banner in the variant too.”
What a strong answer doesThe answer key the graders mark against

A spec that tests one change and fixes the draft's three problems. The primary metric is orders per basket-viewing customer, not checkout completion rate: showing the fee earlier changes who starts checkout, so completion can rise while orders fall. Customers, not sessions, are randomised, so returning customers see one version. The free-delivery banner stays out of the variant (or gets its own test), or the result can't be attributed. The duration is fixed in advance, with no daily stopping: detecting 0.5pp on an 8% base needs roughly 47,000 customers per arm (about 95,000 in total), and because half of each week's viewers are returning customers, new customers accumulate at about 40,000, 60,000, 80,000 and 100,000 after four weeks, so it runs about four full weeks, not two and a half. Guardrails with thresholds: average order value, fee-related support contacts, and the checkout-start rate as a diagnostic. A sample ratio check before reading results. A decision rule written up front: ship if orders are no worse than 0.5pp down and fee contacts fall; don't ship if orders drop by 0.5pp or more; with a stated plan for an inconclusive result.

Critical failures (cap the score at 40)
  • Uses checkout completion rate (orders ÷ checkout starts) as the primary metric
  • Bundles the free-delivery banner into the variant
Case

Mid-level · v1.1 · synthetic · B2C, ecommerce, pricing, metric choice

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit8 of 12 checks passed

The spec is strong but needs substantive edits to add guardrail thresholds, a complete decision rule including inconclusive and no-worse-than-0.5pp outcomes, and to remove unsupported current-system assumptions.

Output

Experiment: show delivery fees in the basket

Owner/build: Ravi · Analysis: Ines · Approval: Chloe Decision: Should Basketful show delivery charges before checkout?

Hypothesis and scope

Showing the fee in the basket will reduce fee surprises, but could discourage ordering. Unexpected delivery fees accounted for 12% of support contacts last quarter; this is motivation, not a baseline customer complaint rate.

Test fee visibility only. Do not add the proposed free-delivery promotional banner: it introduces a second change and would prevent us attributing results to earlier disclosure.

Experience

  • Control: Current experience; delivery fee first appears at the final checkout step.
  • Treatment: Show a clearly labelled delivery line beside the basket total: £3.99 delivery, or Free delivery when eligible. Update the fee and total as the basket changes.
  • Use the existing pricing service and eligibility rules. Confirm behaviour at exactly £60 before launch; “over £60” must not silently become “£60 or more.”
  • Leave checkout, pricing and other messaging unchanged.

Eligibility and assignment

Include customers viewing a non-empty, orderable basket. Exclude staff, bots and test accounts using rules fixed before launch.

Randomise 50/50 by customer, not session, at their first eligible basket view. Persist assignment across visits and devices where identity is available. For signed-out customers, use a stable first-party identifier; Ravi must document identity merging and cross-device limitations before launch.

Analyse everyone assigned, whether or not they start checkout or successfully see the fee: intention to treat.

Metrics

Primary: Customer order conversion — percentage of assigned customers placing at least one order within seven days of assignment. Report treatment minus control in percentage points, with a two-sided 95% confidence interval.

Checkout completion is diagnostic only: treatment may change who starts checkout, so orders ÷ checkout starts could improve while overall ordering falls.

Secondary/diagnostic: - Checkout-start rate and checkout completion. - Delivery-fee-related support contacts per assigned customer within seven days, using consistent tagging. - Orders and revenue per assigned customer; average order value and free-delivery eligibility as diagnostics.

Guardrails: Contribution margin per assigned customer, if reliably available; otherwise net revenue per assigned customer as an explicitly limited proxy. Also monitor basket/checkout errors. Chloe and Ines must approve numerical economic and reliability tolerances before launch.

Sample size and schedule

Chloe’s smallest worthwhile order-conversion change is 0.5 percentage points: 8.0% to 8.5%, not a 0.5% relative lift.

At 5% two-sided significance and 80% power, detecting that increase requires approximately 48,000 customers per arm, or 96,000 total.

Ines must first verify that the 8% baseline matches the eligibility rules and seven-day outcome window; recalculate and freeze the sample target if necessary.

Do not count returning customers again. Expected recruitment is approximately:

  • Week 1: 40,000 unique customers.
  • Each subsequent week: 20,000 additional customers.
  • Four weeks: approximately 100,000 unique customers.

Enroll for at least four complete weeks and until the frozen sample target is reached, ending at a full-week boundary. Then wait seven days for the final cohort’s outcomes: roughly five weeks to readout.

Monitoring and decision

Ravi validates assignment persistence, fee accuracy, exposure/order linkage and sample-ratio balance before ramp-up.

Daily checks are for instrumentation and predefined safety breaches—not efficacy stopping. No “stop when significant,” sample extensions based on results, or repeated significance decisions.

Ines delivers one final primary analysis. Chloe’s default rollout rule is: - Estimated conversion gain at least +0.5 points; - 95% confidence interval excludes zero; - Guardrails pass.

Otherwise, do not roll out on this result alone. A non-significant result does not prove no harm. Secondary trust improvements inform follow-up, not a post-hoc replacement for the primary success rule.

Claims the judge couldn’t find in the brief

  • The current experience includes a pricing service and eligibility rules that can be used to show the fee.
  • Signed-out customers can be assigned using a stable first-party identifier, with identity merging and cross-device limitations to document.
  • Delivery-fee-related support contacts can be measured per assigned customer using consistent tagging.
  • Contribution margin or net revenue per assigned customer is available as a guardrail metric.
  • Staff, bots, test accounts, non-empty orderable baskets, and basket/checkout errors can be identified and excluded or monitored.

Grades and run details

Decision model 88 · LLM judge 8 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly7%
  • passAddresses the actual decision52%
  • partialRespects explicit constraints24%
  • passIdentifies material uncertainty73%
  • passAvoids unsupported claims69%
  • passProduces the required deliverable47%
  • passTests one change at a time97%
  • passFixes the duration up front82%
  • passAn unambiguous primary metric76%
  • partialDecision rule written before the test34%
  • passSized from the real traffic59%
  • partialGuardrails with thresholds99%
Run
Run
#1
Time to output
47 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 12 checks passed

The core reasoning and evidence are correct; a quick edit to add a sample ratio check, guardrail thresholds and a one-line explanation for no peeking would make it fully ready.

Output

Basket delivery-fee experiment

For: Ravi (build), Ines (analysis), Chloe (approval) Decision requested: Approve a test of showing the existing delivery fee in the basket.

Hypothesis and scope

Showing the delivery cost before checkout will reduce surprise and fee-related support contacts, without materially reducing orders.

  • Control: Current experience; delivery fee first appears on the final checkout step.
  • Variant: Show the applicable fee in the basket: “Delivery £3.99” below the free-delivery threshold, or “Free delivery” when the basket qualifies. Keep the existing “free on orders over £60” rule and wording.
  • Change nothing else. Do not add the free-delivery banner: that would test a second intervention and make the result hard to interpret.

Build and assignment — Ravi

Randomise 50/50 by customer, not session, and keep each customer in the same arm across visits and devices where identity is available. Stratify assignment by new versus returning basket viewers. Exclude staff and test accounts. Log assignment, basket value/fee shown, basket views, checkout starts, orders, order value, and relevant support contacts.

Measurement — Ines

Primary metric: Basket-viewer conversion: percentage of assigned customers who place at least one order within 7 days of their first eligible basket view. Count each customer once, in their assigned arm (intention-to-treat). This captures any deterrence before checkout; checkout completion alone would miss it.

Guardrails / secondary measures: Checkout starts per basket viewer; checkout completion (orders ÷ checkout starts); average order value and revenue per assigned customer; delivery-fee-related support contacts per assigned customer. Support contacts may be too sparse for a definitive test, so treat them as directional, not a substitute for the primary outcome.

Sample, duration, and analysis

Baseline basket-viewer conversion is 8%. Power the test to detect a 0.5 percentage-point absolute difference (Chloe’s smallest worthwhile order effect), with 80% power and a two-sided 5% significance level. This requires approximately 95,000 unique customers total (about 47,500 per arm), assuming independent customer-level outcomes.

There are about 40,000 basket viewers weekly, but many return across weeks. Run for at least four full weeks and until the sample target is reached; then allow 7 days for outcomes to mature before reading results. If the target is not reached, extend in full weeks. Ines should confirm the final sample-size calculation and report actual enrolment.

Do not stop early for significance or check daily results to make a stopping decision. Monitor only for implementation or customer-harm issues. Analyse the primary metric at the customer level, with a 95% confidence interval and the pre-specified two-sided test.

Decision

Recommend rollout if the variant improves basket-viewer conversion by at least 0.5 percentage points, the result is statistically significant, and guardrails show no material harm. Do not roll out if there is significant harm. If results are inconclusive, report that clearly; do not interpret “not significant” as proof of no effect.

Grades and run details

Decision model 88 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly14%
  • passAddresses the actual decision54%
  • partialRespects explicit constraints17%
  • passIdentifies material uncertainty35%
  • passAvoids unsupported claims64%
  • passProduces the required deliverable50%
  • passTests one change at a time94%
  • passFixes the duration up front62%
  • passAn unambiguous primary metric40%
  • partialDecision rule written before the test27%
  • passSized from the real traffic39%
  • partialGuardrails with thresholds98%
Run
Run
#1
API response time
22 s
Submitted
30 Sept 2026
Critical failure: Bundles the free-delivery banner into the variant. Score capped at 40.

Needs more than a quick edit1 of 12 checks passed

The spec bundles the banner, uses an inadequate duration, and lacks a decision rule and guardrail thresholds, so a PM would need to redo core parts.

Output

To: Chloe (Head of Growth), Ravi (Growth Engineer), Ines (Data Analyst) From: PM, Checkout Subject: Experiment Spec: Early Delivery Fee Display in Basket

Team, here is the polished spec for our upcoming checkout experiment. I’ve incorporated Chloe’s threshold, fixed a few methodological traps in the draft, and scoped this for a clean, reliable read.

---

1. Objective & Hypothesis

  • Current State: Delivery fee (£3.99, free over £60) is only revealed on the final checkout step. 12% of last quarter’s support contacts were fee-related surprises.
  • Hypothesis: Showing the delivery fee (and progress to free delivery) directly in the basket increases customer trust and sets clear expectations upfront.
  • Expected Impact: While Chloe rightly worries this might cause some upfront drop-off among sub-£60 baskets, we believe the resulting higher-intent checkout starts will lift overall conversion and reduce cart abandonment due to "sticker shock."

---

2. Variant Definition

  • Control (A): Existing basket experience. Delivery fee hidden until final checkout step.
  • Variant (B):
  • Displays the delivery fee clearly in the basket summary.
  • Note on scope: Per Chloe’s suggestion, we are bundling the "free delivery over £60" progress banner into Variant B. Since the fee and the threshold are conceptually tied, showing one without the context of the other would skew the test.

---

3. Traffic Allocation & Unit of Randomization

  • Unit of Randomization: User-level (logged-in state or persistent cookie). Correction from draft: Session-level splitting would cause a jarring experience if a user refreshes or returns and sees fluctuating fee visibility.
  • Split: 50/50 (Control / Variant).
  • Audience: All customers viewing the basket (~40,000 unique users/week, split roughly 50/50 between new and returning).

---

4. Metrics

  • Primary Metric (Guardrail for Business Value):
  • Order Conversion Rate: Completed Orders ÷ Total Basket Views.
  • (Correction from draft: The draft proposed "Checkout Completion Rate (Orders ÷ Checkout Starts)." That misses top-of-funnel drop-off. If Variant B scares people away before checkout, measuring only checkout-starts will falsely make the variant look better. We must measure from the basket view).
  • Secondary Metrics:
  • Checkout Start Rate: Checkout Starts ÷ Basket Views.
  • Checkout Completion Rate: Orders ÷ Checkout Starts (to isolate downstream behavior).
  • Average Order Value (AOV): To monitor if the £60 threshold banner successfully drives basket-building behavior (baseline: £47).
  • Support Metric:
  • Volume of delivery-fee-related support tickets (Ines, let’s pull this baseline post-launch).

---

5. Success Criteria & Statistical Guardrails

  • Minimum Detectable Effect (MDE): Per Chloe, we care about a ≥ +0.5 percentage point absolute change in the Primary Metric (Order Conversion Rate from basket view, baseline ~8%).
  • Duration & Peeking:
  • Correction from draft: Do not stop the test the moment p < 0.05. Early peeking yields false positives.
  • We will run the test for a fixed duration of 2 full weeks to capture day-of-week and new vs. returning user cycles (~80,000 total unique users), yielding adequate statistical power for our MDE.

---

Next Steps:

  • Ravi: Let’s sync on tracking implementation for the basket-level fee and banner. Target build completion: [Insert Date].
  • Ines: Please set up the tracking dashboard, ensuring we segment by New vs. Returning users and Basket Value (<£60 vs. >£60).
  • Chloe: Good to approve?

Claims the judge couldn’t find in the brief

  • 2 full weeks yields ~80,000 total unique users and adequate statistical power for the MDE.

Grades and run details

Decision model 33 · LLM judge 2 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly10%
  • failAddresses the actual decision65%
  • failRespects explicit constraints16%
  • partialIdentifies material uncertainty73%
  • partialAvoids unsupported claims30%
  • partialProduces the required deliverable75%
  • failTests one change at a time99%
  • partialFixes the duration up front76%
  • partialAn unambiguous primary metric69%
  • partialDecision rule written before the test47%
  • partialSized from the real traffic69%
  • partialGuardrails with thresholds45%
Run
Run
#1
Time to output
5 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 11

Uses the supplied evidence correctlyMixedRightWrong
GPT-6 Astra · ChatGPT

It uses the supplied numbers correctly but invents current systems and data capabilities such as an existing pricing service, stable first-party identifiers, identity merging, support-contact tagging, and margin/revenue guardrail data.

GPT-6 Luna · API

Every statement about the current situation is taken directly from the supplied context or follows by arithmetic.

Gemini 3.5 Flash-Lite · Gemini

The claim that 2 weeks yields ~80,000 unique users and adequate power is not supported by the brief; the brief's traffic and returning-user dynamics do not support that sample size or power.

Addresses the actual decisionRightRightMixed
GPT-6 Astra · ChatGPT

It gives Chloe a clear pre-registered rollout rule and states the result conditions that would change the decision.

GPT-6 Luna · API

The output clearly requests approval of the test and provides a complete spec for Chloe to decide, with a post-test decision rule.

Gemini 3.5 Flash-Lite · Gemini

The output commits to a specific experiment design and asks Chloe for approval, which is the deliverable requested.

Respects explicit constraintsRightRightWrong
GPT-6 Astra · ChatGPT

It is under 700 words, addressed to Ravi, Ines, and Chloe, and its proposed customer randomisation, fixed duration, and no-peeking checks would enforce the key constraints.

GPT-6 Luna · API

The output fixes all three problems in the draft, stays under 700 words, and is addressed to the three named readers.

Gemini 3.5 Flash-Lite · Gemini

The brief says 'Fix what needs fixing' and the draft bundles the free-delivery banner; the output keeps it bundled, failing to fix that problem.

Identifies material uncertaintyRightRightWrong
GPT-6 Astra · ChatGPT

It names material uncertainties such as baseline eligibility, identity merging, cross-device limits, and guardrail tolerances, and assigns pre-launch verification or approval.

GPT-6 Luna · API

It notes support contacts may be too sparse for a definitive test and that inconclusive results are possible, with a plan for each.

Gemini 3.5 Flash-Lite · Gemini

The output does not name unknowns that could change the decision, nor does it specify what would happen if the result is inconclusive.

Avoids unsupported claimsMixedRightWrong
GPT-6 Astra · ChatGPT

It presents several current capabilities and data sources as available without support, and its rollout rule is stricter than the supplied evidence requires.

GPT-6 Luna · API

The hypothesis is clearly labelled, and no interpretations are presented as established fact.

Gemini 3.5 Flash-Lite · Gemini

The output presents the 2-week duration as yielding adequate power as fact, without labelling it as an assumption, and the evidence does not support it.

Produces the required deliverableRightRightWrong
GPT-6 Astra · ChatGPT

It is a usable experiment spec for the named readers, with scope, assignment, metrics, sizing, monitoring, and decision rule, needing only light edits.

GPT-6 Luna · API

The spec is complete, under 700 words, and directly usable by Ravi, Ines and Chloe.

Gemini 3.5 Flash-Lite · Gemini

The spec is missing a decision rule, guardrail thresholds, proper sample-size justification, and a sample ratio check, so it is not usable as is.

Tests one change at a timeRightRightWrong
GPT-6 Astra · ChatGPT

It keeps the free-delivery banner out of the variant and explains that bundling it would prevent attribution.

GPT-6 Luna · API

It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.

Gemini 3.5 Flash-Lite · Gemini

The spec bundles the free-delivery banner into the variant, arguing it is conceptually tied, instead of testing the fee change alone.

Fixes the duration up frontRightMixedRight
GPT-6 Astra · ChatGPT

It replaces daily significance stopping with a fixed four-week enrollment plus seven-day outcome window and explains that daily efficacy stopping is not allowed.

GPT-6 Luna · API

It sets a fixed duration but does not explain why daily peeking inflates false positives, only instructs not to do it.

Gemini 3.5 Flash-Lite · Gemini

The spec replaces 'stop as soon as significant' with a fixed 2-week duration and explains that early peeking yields false positives.

An unambiguous primary metricRightMixedWrong
GPT-6 Astra · ChatGPT

It names one primary metric, explains why checkout completion is misleading, and includes a sample-ratio balance check.

GPT-6 Luna · API

It names one primary metric with a rationale but omits a planned trust check such as a sample ratio check.

Gemini 3.5 Flash-Lite · Gemini

The spec names one primary metric but does not include a trust check such as a sample ratio check.

Decision rule written before the testWrongRightWrong
GPT-6 Astra · ChatGPT

It does not map every outcome to an action, especially a flat or small positive result, and omits the supplied evidence-based rule to ship if orders are no worse than 0.5pp down while fee contacts fall.

GPT-6 Luna · API

It maps rollout, no rollout and inconclusive results to actions, with thresholds of 0.5pp improvement and statistical significance.

Gemini 3.5 Flash-Lite · Gemini

No decision rule is stated; the spec does not map outcomes (including inconclusive) to actions with thresholds.

Sized from the real trafficRightRightWrong
GPT-6 Astra · ChatGPT

The sample size follows from the 8% baseline and 0.5pp MDE, and the duration follows from weekly unique-customer accumulation in whole weeks.

GPT-6 Luna · API

Sample size is calculated from the 8% baseline and 0.5pp effect, and the four-week duration accounts for returning visitors and whole-week cycles.

Gemini 3.5 Flash-Lite · Gemini

The 2-week duration is not worked out from the traffic and MDE; the required sample size is much larger given the 8% baseline and 0.5pp effect.

All got wrong 1

Guardrails with thresholdsWrongWrongWrong
GPT-6 Astra · ChatGPT

Guardrails are named but lack numerical thresholds that would block rollout, deferring them to later approval.

GPT-6 Luna · API

Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.

Gemini 3.5 Flash-Lite · Gemini

Guardrail metrics are listed but no thresholds are set that would block a rollout.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review