Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for checkout at Basketful. We want to test showing the delivery fee in the basket, instead of only at the last step of checkout. Write the experiment spec for Ravi, the growth engineer who'll build it, Ines, the analyst who'll read it out, and Chloe, our Head of Growth, who approves it. Keep it under 700 words. The team's draft plan is below. Fix what needs fixing.

What the model was given5 items: About Basketful, Why we're testing it, Traffic and baselines, Smallest change worth acting on, The team's draft plan
About BasketfulAn online grocery service. Delivery costs £3.99, or is free on orders over £60. Today the fee first appears on the final checkout step.
Why we're testing it12% of support contacts last quarter were about delivery fees the customer hadn't expected. Design believes showing the fee early will build trust; Chloe worries it will put people off before they start checkout.
Traffic and baselinesAbout 40,000 different customers view their basket each week. About half of each week's basket viewers are new that week; the rest came back from earlier weeks. 8% of basket viewers place an order. Of customers who start checkout, 62% complete it. Average order value: £47.
Smallest change worth acting onChloe: “Anything smaller than half a percentage point on orders isn't worth arguing about.”
The team's draft planPrimary metric: checkout completion rate (orders ÷ checkout starts). Split: 50/50 by session. Duration: we'll check the dashboard every day and stop as soon as it's significant. Chloe's addition: “While we're at it, let's put the free-delivery-over-£60 banner in the variant too.”
What a strong answer doesThe answer key the graders mark against

A spec that tests one change and fixes the draft's three problems. The primary metric is orders per basket-viewing customer, not checkout completion rate: showing the fee earlier changes who starts checkout, so completion can rise while orders fall. Customers, not sessions, are randomised, so returning customers see one version. The free-delivery banner stays out of the variant (or gets its own test), or the result can't be attributed. The duration is fixed in advance, with no daily stopping: detecting 0.5pp on an 8% base needs roughly 47,000 customers per arm (about 95,000 in total), and because half of each week's viewers are returning customers, new customers accumulate at about 40,000, 60,000, 80,000 and 100,000 after four weeks, so it runs about four full weeks, not two and a half. Guardrails with thresholds: average order value, fee-related support contacts, and the checkout-start rate as a diagnostic. A sample ratio check before reading results. A decision rule written up front: ship if orders are no worse than 0.5pp down and fee contacts fall; don't ship if orders drop by 0.5pp or more; with a stated plan for an inconclusive result.

Critical failures (cap the score at 40)
  • Uses checkout completion rate (orders ÷ checkout starts) as the primary metric
  • Bundles the free-delivery banner into the variant
Case

Mid-level · v1.1 · synthetic · B2C, ecommerce, pricing, metric choice

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit10 of 12 checks passed

Guardrail thresholds are vague ('materially worse', 'drops meaningfully'), so a PM would need to define them before the test can run.

Output

Experiment spec: delivery fee in basket

Owner: PM, Checkout | Readers: Ravi (build), Ines (analysis), Chloe (approval)

Hypothesis

Showing the delivery fee in the basket reduces surprise at checkout (12% of last quarter's support contacts) without reducing orders. Chloe's concern is that an early fee deters people before they start checkout, so the test measures that directly.

Changes from the team draft

DraftProblemFix
Primary metric: checkout completionIt only counts people who already started checkout. If the early fee puts people off, they never enter the denominator, so the harm would be hidden.Primary metric is orders per basket viewer.
Split by sessionReturning customers (about half of weekly viewers) would see both versions.Randomise by customer.
Check daily, stop when significantRepeated peeking inflates false positives.Fixed sample size and duration, one readout.
Add the free-delivery banner to the variantWe couldn't tell whether the fee or the banner caused any change.Banner excluded. Test it separately afterwards.

Design

  • Control: fee first shown at final checkout step (today).
  • Variant: basket shows the delivery line: £3.99, or "Free" at £60+. Nothing else changes. The basket and checkout amounts must always match.
  • Unit: customer ID, with a 50/50 split. Assignment is sticky for the whole test. Logged-out users are bucketed by a persistent device ID (Ravi to confirm how many there are).
  • Enrolment: at first basket view during the test. Log an exposure event at that moment.

Metrics

Primary: share of enrolled customers who place at least one order within 7 days of first basket view. Baseline is about 8%.

Guardrails: - Checkout start rate (baseline about 12.9%, i.e. 8% ÷ 62%). - Revenue per enrolled customer (baseline about £3.76, i.e. 8% × £47). - Delivery-fee support contacts per 1,000 customers. This is the benefit metric, but volumes are low, so read it directionally.

Diagnostics (not decision metrics): - Checkout completion (orders ÷ checkout starts). Read it only alongside the primary metric, since it's conditional on starts. - New vs returning customers, and baskets under vs over £60. Both are exploratory.

Sample size and duration

  • Smallest effect worth detecting: 0.5 percentage points on orders (8.0% → 8.5%), per Chloe.
  • Requirement: 80% power, 5% two-sided significance. This needs about 46,000 customers per arm, roughly 92,000 in total.
  • Time to reach it: Week 1 brings about 40,000 unique customers. After that, only about 20,000 new customers enrol per week, because returning customers are already assigned. That reaches about 100,000 by the end of week 4.
  • Duration: enrol for 4 full weeks, plus 7 days of follow-up for the last cohort. Readout is at about 5 weeks, with no early stopping. Ines can monitor data quality but not results.
  • Limit: the test can't reliably detect effects smaller than 0.5pp. A "no significant difference" result means "no effect larger than about 0.5pp", not "no effect".

Data quality checks (before readout)

  • Sample ratio check: arms should be within normal variation of 50/50.
  • Exposure is logged for every enrolled customer.
  • Fee shown in basket equals fee charged, in all cases.

Decision rule (Chloe to confirm)

  • Ship if the 95% confidence interval for the change in orders rules out a drop of 0.5pp or more, no guardrail is materially worse, and delivery-fee contacts do not rise.
  • Don't ship if orders fall by 0.5pp or more, or revenue per customer drops meaningfully.
  • Otherwise, treat as inconclusive and discuss. Do not extend the test to "find" significance.

Next

If the variant ships or is neutral, run a follow-up test of the free-delivery-over-£60 banner, where the fee is already visible. Its effect is likely concentrated in baskets just under £60.

Open items

  • Ravi: logged-out identity handling; a flag to switch the variant off quickly.
  • Chloe: confirm the decision rule and that the 0.5pp threshold means absolute percentage points on orders per basket viewer.
  • Ines: confirm the 7-day order attribution and the support-contact tagging for fee-related contacts.

Grades and run details

Decision model 79 · LLM judge 11 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly4%
  • passAddresses the actual decision61%
  • partialRespects explicit constraints23%
  • passIdentifies material uncertainty80%
  • passAvoids unsupported claims35%
  • passProduces the required deliverable67%
  • passTests one change at a time99%
  • passFixes the duration up front100%
  • passAn unambiguous primary metric96%
  • passDecision rule written before the test95%
  • partialSized from the real traffic31%
  • partialGuardrails with thresholds94%
Run
Run
#1
API response time
29 s
Submitted
30 Sept 2026

Needs more than a quick edit3 of 12 checks passed

A PM would need to redo the sizing and duration, remove or redesign the early kill rule, add a sample-ratio check, and complete guardrail thresholds. These are substantive reasoning changes, not quick wording edits.

Output

Experiment Spec: Delivery Fee Transparency in Basket

To: Chloe (Approval), Ravi (Build), Ines (Analysis) From: Checkout PM

---

1. Objective & Hypothesis

  • Context: 12% of support contacts complain about surprise delivery fees at final checkout.
  • Hypothesis: Displaying the delivery fee (£3.99, or "Free" if >£60) upfront in the basket will set accurate expectations, reduce support contacts, and build customer trust without reducing net orders.
  • Risk (Chloe’s concern): Upfront cost shock could cause drop-off before checkout begins.

---

2. Corrections to the Draft Plan

1. Primary Metric Changed: Draft used Checkout Completion Rate (Orders ÷ Starts). Correction: Use Basket-to-Order Conversion Rate (Orders ÷ Basket Viewers). Measuring only checkout starts introduces severe selection bias—if fee visibility deters low-intent users at the basket, checkout completion would artificially increase while overall orders plummet. 2. Unit of Randomization Changed: Draft used session-level. Correction: Randomize by User ID (falling back to persistent anonymous visitor ID). 50% of weekly viewers are repeat shoppers; session-level randomization would expose returning users to conflicting experiences. 3. No Peeking / Fixed Duration: Draft suggested stopping once "significant." Correction: Continuous monitoring with early stopping inflates false-positive rates to >30%. We will fix sample size upfront and run for complete weekly cycles. 4. Scope Control (Chloe’s Banner): Cut the promotional free-delivery banner from this test. Adding a promotional banner introduces a confounding variable, making it impossible to isolate whether fee transparency helped or hurt. We will test the banner in a follow-up experiment.

---

3. Variants (For Ravi)

  • Allocation: 50/50 split at User/Persistent Device ID level. Triggered on basket view.
  • Control (A): Current basket. No delivery fee shown (fee appears on final checkout step as today).
  • Variant (B): Add a delivery fee line item to the Basket Order Summary:
  • Basket < £60: `Delivery: £3.99`
  • Basket ≥ £60: `Delivery: FREE`
  • Include delivery fee in the displayed Basket Total.
  • Note: Do not add promotional progress bars or marketing banners.

---

4. Metrics Framework (For Ines)

  • Primary Metric: Basket-to-Order Conversion Rate (Total Unique Ordering Users ÷ Total Unique Basket Viewers).
  • Guardrail Metrics:
  • Average Order Value (AOV baseline: £47).
  • Basket-to-Checkout Start Rate (tracks where abandonment occurs).
  • Total Net Revenue (Orders × AOV).
  • Secondary / Diagnostic Metrics:
  • Delivery-fee-related customer support ticket rate per 1,000 orders.
  • Checkout Completion Rate (Checkout Starts → Orders).

---

5. Sample Size & Run Time (For Ines & Chloe)

  • Baseline Conversion: 8.0% of basket viewers place an order.
  • Minimum Detectable Effect (MDE): Absolute ±0.5 percentage points (7.5% or 8.5%, a 6.25% relative change), per Chloe's threshold.
  • Statistical Parameters: $\alpha = 0.05$ (two-tailed), Power ($1 - \beta$) = 80%.
  • Required Sample: ~47,000 unique users per variant (~94,000 total unique basket viewers).
  • Duration: 3 full weeks (21 days).
  • With 40,000 viewers/week (~20k new, ~20k repeat), 3 weeks yields ~80,000–90,000 unique basket viewers across three full grocery-shopping cycles.
  • Ines will run the read-out strictly after the 21-day window closes.

---

6. Decision & Rollout Criteria (For Chloe)

  • Ship Variant: Primary conversion does not drop by more than Chloe’s threshold of -0.5pp (non-inferiority) AND delivery-related support contacts decrease significantly.
  • Iterate / Re-test: Basket-to-order conversion drops by >0.5pp. In this scenario, early fee disclosure hurts conversion, justifying Chloe’s concern; we then test the £60 threshold banner/progress bar to offset drop-off.
  • Kill Immediately: Overall revenue drops >5% statistically significantly at the end of Week 1 (safety check only; not an early win stop).

Claims the judge couldn’t find in the brief

  • Three weeks yields about 80,000–90,000 unique basket viewers.
  • Early stopping inflates false-positive rates to more than 30%.
  • The variant's free-delivery threshold is £60 or more.

Grades and run details

Decision model 67 · LLM judge 3 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly55%
  • passAddresses the actual decision63%
  • partialRespects explicit constraints20%
  • passIdentifies material uncertainty65%
  • partialAvoids unsupported claims49%
  • partialProduces the required deliverable32%
  • passTests one change at a time100%
  • passFixes the duration up front66%
  • passAn unambiguous primary metric28%
  • partialDecision rule written before the test79%
  • partialSized from the real traffic94%
  • partialGuardrails with thresholds90%
Run
Run
#1
API response time
19 s
Submitted
1 Oct 2026

Needs more than a quick edit8 of 12 checks passed

The spec is strong but needs substantive edits to add guardrail thresholds, a complete decision rule including inconclusive and no-worse-than-0.5pp outcomes, and to remove unsupported current-system assumptions.

Output

Experiment: show delivery fees in the basket

Owner/build: Ravi · Analysis: Ines · Approval: Chloe Decision: Should Basketful show delivery charges before checkout?

Hypothesis and scope

Showing the fee in the basket will reduce fee surprises, but could discourage ordering. Unexpected delivery fees accounted for 12% of support contacts last quarter; this is motivation, not a baseline customer complaint rate.

Test fee visibility only. Do not add the proposed free-delivery promotional banner: it introduces a second change and would prevent us attributing results to earlier disclosure.

Experience

  • Control: Current experience; delivery fee first appears at the final checkout step.
  • Treatment: Show a clearly labelled delivery line beside the basket total: £3.99 delivery, or Free delivery when eligible. Update the fee and total as the basket changes.
  • Use the existing pricing service and eligibility rules. Confirm behaviour at exactly £60 before launch; “over £60” must not silently become “£60 or more.”
  • Leave checkout, pricing and other messaging unchanged.

Eligibility and assignment

Include customers viewing a non-empty, orderable basket. Exclude staff, bots and test accounts using rules fixed before launch.

Randomise 50/50 by customer, not session, at their first eligible basket view. Persist assignment across visits and devices where identity is available. For signed-out customers, use a stable first-party identifier; Ravi must document identity merging and cross-device limitations before launch.

Analyse everyone assigned, whether or not they start checkout or successfully see the fee: intention to treat.

Metrics

Primary: Customer order conversion — percentage of assigned customers placing at least one order within seven days of assignment. Report treatment minus control in percentage points, with a two-sided 95% confidence interval.

Checkout completion is diagnostic only: treatment may change who starts checkout, so orders ÷ checkout starts could improve while overall ordering falls.

Secondary/diagnostic: - Checkout-start rate and checkout completion. - Delivery-fee-related support contacts per assigned customer within seven days, using consistent tagging. - Orders and revenue per assigned customer; average order value and free-delivery eligibility as diagnostics.

Guardrails: Contribution margin per assigned customer, if reliably available; otherwise net revenue per assigned customer as an explicitly limited proxy. Also monitor basket/checkout errors. Chloe and Ines must approve numerical economic and reliability tolerances before launch.

Sample size and schedule

Chloe’s smallest worthwhile order-conversion change is 0.5 percentage points: 8.0% to 8.5%, not a 0.5% relative lift.

At 5% two-sided significance and 80% power, detecting that increase requires approximately 48,000 customers per arm, or 96,000 total.

Ines must first verify that the 8% baseline matches the eligibility rules and seven-day outcome window; recalculate and freeze the sample target if necessary.

Do not count returning customers again. Expected recruitment is approximately:

  • Week 1: 40,000 unique customers.
  • Each subsequent week: 20,000 additional customers.
  • Four weeks: approximately 100,000 unique customers.

Enroll for at least four complete weeks and until the frozen sample target is reached, ending at a full-week boundary. Then wait seven days for the final cohort’s outcomes: roughly five weeks to readout.

Monitoring and decision

Ravi validates assignment persistence, fee accuracy, exposure/order linkage and sample-ratio balance before ramp-up.

Daily checks are for instrumentation and predefined safety breaches—not efficacy stopping. No “stop when significant,” sample extensions based on results, or repeated significance decisions.

Ines delivers one final primary analysis. Chloe’s default rollout rule is: - Estimated conversion gain at least +0.5 points; - 95% confidence interval excludes zero; - Guardrails pass.

Otherwise, do not roll out on this result alone. A non-significant result does not prove no harm. Secondary trust improvements inform follow-up, not a post-hoc replacement for the primary success rule.

Claims the judge couldn’t find in the brief

  • The current experience includes a pricing service and eligibility rules that can be used to show the fee.
  • Signed-out customers can be assigned using a stable first-party identifier, with identity merging and cross-device limitations to document.
  • Delivery-fee-related support contacts can be measured per assigned customer using consistent tagging.
  • Contribution margin or net revenue per assigned customer is available as a guardrail metric.
  • Staff, bots, test accounts, non-empty orderable baskets, and basket/checkout errors can be identified and excluded or monitored.

Grades and run details

Decision model 88 · LLM judge 8 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly7%
  • passAddresses the actual decision52%
  • partialRespects explicit constraints24%
  • passIdentifies material uncertainty73%
  • passAvoids unsupported claims69%
  • passProduces the required deliverable47%
  • passTests one change at a time97%
  • passFixes the duration up front82%
  • passAn unambiguous primary metric76%
  • partialDecision rule written before the test34%
  • passSized from the real traffic59%
  • partialGuardrails with thresholds99%
Run
Run
#1
Time to output
47 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 8

Uses the supplied evidence correctlyMixedWrongMixed
Sonnet 5.5 · API

All statements about the current situation are taken directly from the brief or derived by correct arithmetic.

Gemini 3.8 Flash · API

It misstates the free-delivery threshold as ≥£60, overstates unique viewers over three weeks, and adds an unsupported >30% false-positive claim.

GPT-6 Astra · ChatGPT

It uses the supplied numbers correctly but invents current systems and data capabilities such as an existing pricing service, stable first-party identifiers, identity merging, support-contact tagging, and margin/revenue guardrail data.

Respects explicit constraintsRightWrongRight
Sonnet 5.5 · API

The output is under 700 words, addresses the named readers, and is a complete experiment spec.

Gemini 3.8 Flash · API

It is under 700 words and addressed to the named readers, but it proposes a 21-day test that cannot reach the stated sample and includes a week-1 significance-based kill rule.

GPT-6 Astra · ChatGPT

It is under 700 words, addressed to Ravi, Ines, and Chloe, and its proposed customer randomisation, fixed duration, and no-peeking checks would enforce the key constraints.

Avoids unsupported claimsRightWrongMixed
Sonnet 5.5 · API

Hypotheses and forecasts are clearly labelled, and no interpretations are presented as established fact.

Gemini 3.8 Flash · API

It presents the >30% false-positive rate and the 80,000–90,000 unique-viewer range as established facts without support.

GPT-6 Astra · ChatGPT

It presents several current capabilities and data sources as available without support, and its rollout rule is stricter than the supplied evidence requires.

Produces the required deliverableRightWrongRight
Sonnet 5.5 · API

The spec is complete, in the right form, and usable by Ravi, Ines, and Chloe with light edits.

Gemini 3.8 Flash · API

The spec is usable in form but has major gaps: the duration is underpowered, the sample-ratio check is missing, and the guardrail thresholds are incomplete.

GPT-6 Astra · ChatGPT

It is a usable experiment spec for the named readers, with scope, assignment, metrics, sizing, monitoring, and decision rule, needing only light edits.

Fixes the duration up frontRightMixedRight
Sonnet 5.5 · API

Duration is fixed at 4 weeks enrollment plus 7 days follow-up, with no early stopping, and the false-positive risk of peeking is explained.

Gemini 3.8 Flash · API

It fixes a duration and explains peeking, but the duration is not derived from the required sample and it still allows a week-1 significance-based stop.

GPT-6 Astra · ChatGPT

It replaces daily significance stopping with a fixed four-week enrollment plus seven-day outcome window and explains that daily efficacy stopping is not allowed.

An unambiguous primary metricRightMixedRight
Sonnet 5.5 · API

The primary metric is orders per enrolled customer within 7 days, with a clear rationale and a planned sample ratio check.

Gemini 3.8 Flash · API

It names one primary metric and rationale, but does not plan a sample-ratio check or other trustworthiness check.

GPT-6 Astra · ChatGPT

It names one primary metric, explains why checkout completion is misleading, and includes a sample-ratio balance check.

Decision rule written before the testRightWrongWrong
Sonnet 5.5 · API

Every outcome (ship, don't ship, inconclusive) is mapped to an action with thresholds separating them.

Gemini 3.8 Flash · API

It does not define the action for an inconclusive result between the non-inferiority and harm thresholds.

GPT-6 Astra · ChatGPT

It does not map every outcome to an action, especially a flat or small positive result, and omits the supplied evidence-based rule to ship if orders are no worse than 0.5pp down while fee contacts fall.

Sized from the real trafficRightWrongRight
Sonnet 5.5 · API

Sample size is correctly calculated from the 8% baseline and 0.5pp effect, and duration follows from the traffic pattern over whole weeks.

Gemini 3.8 Flash · API

It calculates about 94,000 unique users but sets only three weeks, which yields about 80,000 unique users, not four weeks.

GPT-6 Astra · ChatGPT

The sample size follows from the 8% baseline and 0.5pp MDE, and the duration follows from weekly unique-customer accumulation in whole weeks.

All got wrong 1

Guardrails with thresholdsWrongWrongWrong
Sonnet 5.5 · API

Guardrails are named but lack specific numeric thresholds; 'materially worse' and 'drops meaningfully' are too vague to block rollout unambiguously.

Gemini 3.8 Flash · API

It names guardrails but gives thresholds only for revenue, not for AOV, checkout-start rate, or support contacts.

GPT-6 Astra · ChatGPT

Guardrails are named but lack numerical thresholds that would block rollout, deferring them to later approval.

All got right 3

Addresses the actual decisionRightRightRight
Sonnet 5.5 · API

The spec commits to a clear decision rule with ship/don't ship/inconclusive actions, framed for Chloe to confirm.

Gemini 3.8 Flash · API

It commits to testing fee transparency alone, with ship, iterate, and kill rules tied to named thresholds.

GPT-6 Astra · ChatGPT

It gives Chloe a clear pre-registered rollout rule and states the result conditions that would change the decision.

Identifies material uncertaintyRightRightRight
Sonnet 5.5 · API

It names unknowns like logged-out identity, decision rule confirmation, and low support-contact volumes, and says how to resolve them.

Gemini 3.8 Flash · API

It names the main uncertainty—whether early fee disclosure reduces orders—and maps outcomes to actions.

GPT-6 Astra · ChatGPT

It names material uncertainties such as baseline eligibility, identity merging, cross-device limits, and guardrail tolerances, and assigns pre-launch verification or approval.

Tests one change at a timeRightRightRight
Sonnet 5.5 · API

The free-delivery banner is excluded from the variant and a separate follow-up test is proposed, with the attribution problem explained.

Gemini 3.8 Flash · API

It keeps the promotional free-delivery banner out of the variant and explains that bundling it would confound attribution.

GPT-6 Astra · ChatGPT

It keeps the free-delivery banner out of the variant and explains that bundling it would prevent attribution.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review