Tasks / Experiment

Experiment specification

Can the model design a test that could actually change the decision?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision91% pass
    The output commits to a clear decision rule: launch only if all gates pass, otherwise do not launch, and an inconclusive result is not a pass.
    GPT-6.1 Sol · API · Batching deliveries before peak season
  2. Tests one change at a time86% pass
    It explicitly keeps the free-delivery banner out of the variant and explains that adding it would make the result uninterpretable.
    GPT-6 Luna · API · Showing the delivery fee up front
  3. Identifies material uncertainty84% pass
    It identifies unknowns like power under correlation, carryover, and city differences, and states that if power is insufficient the rollout will not proceed, resolving the uncertainty.
    GPT-6.1 Sol · API · Batching deliveries before peak season

Where it slips

  1. Respects explicit constraints55% pass
    The output is approximately 780 words, exceeding the 'under 700 words' limit.
    Opus 5.5 · Claude · Showing the delivery fee up front
  2. Sized from the real traffic57% pass
    Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.
    GPT-6.1 Sol · API · Showing the delivery fee up front
  3. Guardrails with thresholds59% pass
    Guardrail metrics are listed but no specific thresholds (e.g., maximum acceptable drop in AOV) are given to block rollout.
    GPT-6 Luna · API · Showing the delivery fee up front

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for checkout at Basketful. We want to test showing the delivery fee in the basket, instead of only at the last step of checkout. Write the experiment spec for Ravi, the growth engineer who'll build it, Ines, the analyst who'll read it out, and Chloe, our Head of Growth, who approves it. Keep it under 700 words. The team's draft plan is below. Fix what needs fixing.

What the model was given5 items: About Basketful, Why we're testing it, Traffic and baselines, Smallest change worth acting on, The team's draft plan
About BasketfulAn online grocery service. Delivery costs £3.99, or is free on orders over £60. Today the fee first appears on the final checkout step.
Why we're testing it12% of support contacts last quarter were about delivery fees the customer hadn't expected. Design believes showing the fee early will build trust; Chloe worries it will put people off before they start checkout.
Traffic and baselinesAbout 40,000 different customers view their basket each week. About half of each week's basket viewers are new that week; the rest came back from earlier weeks. 8% of basket viewers place an order. Of customers who start checkout, 62% complete it. Average order value: £47.
Smallest change worth acting onChloe: “Anything smaller than half a percentage point on orders isn't worth arguing about.”
The team's draft planPrimary metric: checkout completion rate (orders ÷ checkout starts). Split: 50/50 by session. Duration: we'll check the dashboard every day and stop as soon as it's significant. Chloe's addition: “While we're at it, let's put the free-delivery-over-£60 banner in the variant too.”
What a strong answer doesThe answer key the graders mark against

A spec that tests one change and fixes the draft's three problems. The primary metric is orders per basket-viewing customer, not checkout completion rate: showing the fee earlier changes who starts checkout, so completion can rise while orders fall. Customers, not sessions, are randomised, so returning customers see one version. The free-delivery banner stays out of the variant (or gets its own test), or the result can't be attributed. The duration is fixed in advance, with no daily stopping: detecting 0.5pp on an 8% base needs roughly 47,000 customers per arm (about 95,000 in total), and because half of each week's viewers are returning customers, new customers accumulate at about 40,000, 60,000, 80,000 and 100,000 after four weeks, so it runs about four full weeks, not two and a half. Guardrails with thresholds: average order value, fee-related support contacts, and the checkout-start rate as a diagnostic. A sample ratio check before reading results. A decision rule written up front: ship if orders are no worse than 0.5pp down and fee contacts fall; don't ship if orders drop by 0.5pp or more; with a stated plan for an inconclusive result.

Critical failures (cap the score at 40)
  • Uses checkout completion rate (orders ÷ checkout starts) as the primary metric
  • Bundles the free-delivery banner into the variant
Case

Mid-level · v1.1 · synthetic · B2C, ecommerce, pricing, metric choice

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit8 of 12 checks passed

The spec exceeds the 700-word limit and lacks guardrail thresholds, so a PM would need to trim it and add those thresholds before use.

Output

Experiment Spec: Delivery Fee Shown in Basket

Owner: Checkout PM · Build: Ravi · Analysis: Ines · Approver: Chloe

Why we're testing this

Today the £3.99 delivery fee (free over £60) first appears on the final checkout step. Last quarter, 12% of support contacts were about unexpected delivery fees. Design expects that showing the fee earlier will build trust. Chloe's concern is that it will put people off before they start checkout. This test tells us which effect is bigger.

Hypothesis

Showing the delivery fee in the basket will change the share of basket viewers who place an order. We don't assume a direction. The test is two-sided because either outcome is plausible.

Variants

  • Control: Current experience. The fee appears only at the final checkout step.
  • Variant: The basket shows a "Delivery: £3.99" line, or "Free" when the basket is over £60. The same fee still appears at checkout.

Change from draft: The free-delivery-over-£60 banner is removed from this test. If both changes ship together, we can't tell which one caused any result. The banner could also raise orders while the fee line lowers them, so the two effects would cancel and we'd learn nothing. We'll test the banner as a follow-up once this result is in.

Primary metric

Order conversion: the share of customers who view their basket during the test and go on to place at least one order.

Change from draft: The draft used checkout completion (orders ÷ checkout starts). That metric would be misleading here. If the variant puts off fee-sensitive customers before checkout, the people who remain in checkout are the ones already willing to pay. Completion would then rise even if total orders fell. Order conversion counts everyone who saw the basket, so it catches losses at either stage.

Secondary and guardrail metrics

  • Checkout start rate and checkout completion rate. These are diagnostic only and show where any change happens.
  • Average order value. Seeing the fee may push some customers over £60.
  • Revenue per basket viewer. This includes delivery fee income.
  • Fee-related support contacts per 1,000 orders. This is the problem we set out to fix.

Randomisation

  • Unit: customer (logged-in account ID). The draft split by session, which needs to change. About half of each week's basket viewers are returning customers. With a session split, they would see the fee in the basket on one visit and not on the next. That would contaminate both groups and confuse customers.
  • Split: 50/50, sticky for the whole test.
  • Ravi: Please log the assignment when a customer first views their basket, and confirm that logged-out traffic is either excluded or bucketed consistently.

Sample size and duration

  • Assumptions: baseline 8% order conversion, two-sided test at α = 0.05, 80% power.
  • Minimum effect we want to detect: 0.5 percentage points (8.0% → 8.5%). We're reading Chloe's threshold as absolute points on order conversion. That is roughly a 6% relative change.
  • Customers needed: about 47,500 per arm, or about 95,000 total.

How long that takes: About 40,000 customers view their basket in week 1, then about 20,000 new customers each week after. That gives roughly 40k, 60k, 80k and 100k cumulative customers. The test therefore runs for 4 full weeks.

Caveat for Ines: The 8% baseline is per basket viewer per week. Over a 4-week window, each customer has more chances to order, so the per-customer baseline will be higher. At 12%, for example, we'd need about 135,000 customers, which means about 6 weeks. Before launch, Ines will calculate the 4-week per-customer baseline from historical data. The final duration will be fixed from that number and recorded here.

Stopping rule

Change from draft: We won't stop the test as soon as the dashboard shows significance. If we check every day and stop on the first significant result, the false-positive rate climbs well above 5%. The dashboard will often show early "wins" that aren't real.

  • The test runs for the pre-set duration, in full weeks only.
  • The dashboard can be watched for bugs and sample-ratio mismatch. Nobody makes ship decisions from it mid-test.
  • Only reason to stop early: a guardrail failure. That means a broken experience, a sample-ratio mismatch (p < 0.001), or order conversion down by more than 1.5 points.

Decision rules

The confidence interval (CI) below is the interval around the difference in order conversion between the variant and control.

Result on order conversionDecision
Up by a significant amountShip the variant
Down by a significant amountDon't ship. Look into other ways to reduce the support contacts
Not significant, and the CI excludes −0.5 pointsShip. We lose no meaningful number of orders and gain trust and fewer support contacts
Not significant, and the CI includes −0.5 pointsTreat as inconclusive. Chloe decides whether to extend the test or drop the change

Sign-off

  • ☐ Ravi: build plus assignment logging
  • ☐ Ines: baseline check and final duration
  • ☐ Chloe: approval

Grades and run details

Decision model 83 · LLM judge 9 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly7%
  • passAddresses the actual decision73%
  • partialRespects explicit constraints16%
  • passIdentifies material uncertainty90%
  • passAvoids unsupported claims45%
  • passProduces the required deliverable59%
  • passTests one change at a time100%
  • passFixes the duration up front98%
  • passAn unambiguous primary metric39%
  • passDecision rule written before the test94%
  • passSized from the real traffic55%
  • partialGuardrails with thresholds84%
Run
Run
#1
Time to output
31 s
Submitted
30 Sept 2026

Needs more than a quick edit8 of 12 checks passed

The spec is strong but needs substantive edits to add guardrail thresholds, a complete decision rule including inconclusive and no-worse-than-0.5pp outcomes, and to remove unsupported current-system assumptions.

Output

Experiment: show delivery fees in the basket

Owner/build: Ravi · Analysis: Ines · Approval: Chloe Decision: Should Basketful show delivery charges before checkout?

Hypothesis and scope

Showing the fee in the basket will reduce fee surprises, but could discourage ordering. Unexpected delivery fees accounted for 12% of support contacts last quarter; this is motivation, not a baseline customer complaint rate.

Test fee visibility only. Do not add the proposed free-delivery promotional banner: it introduces a second change and would prevent us attributing results to earlier disclosure.

Experience

  • Control: Current experience; delivery fee first appears at the final checkout step.
  • Treatment: Show a clearly labelled delivery line beside the basket total: £3.99 delivery, or Free delivery when eligible. Update the fee and total as the basket changes.
  • Use the existing pricing service and eligibility rules. Confirm behaviour at exactly £60 before launch; “over £60” must not silently become “£60 or more.”
  • Leave checkout, pricing and other messaging unchanged.

Eligibility and assignment

Include customers viewing a non-empty, orderable basket. Exclude staff, bots and test accounts using rules fixed before launch.

Randomise 50/50 by customer, not session, at their first eligible basket view. Persist assignment across visits and devices where identity is available. For signed-out customers, use a stable first-party identifier; Ravi must document identity merging and cross-device limitations before launch.

Analyse everyone assigned, whether or not they start checkout or successfully see the fee: intention to treat.

Metrics

Primary: Customer order conversion — percentage of assigned customers placing at least one order within seven days of assignment. Report treatment minus control in percentage points, with a two-sided 95% confidence interval.

Checkout completion is diagnostic only: treatment may change who starts checkout, so orders ÷ checkout starts could improve while overall ordering falls.

Secondary/diagnostic: - Checkout-start rate and checkout completion. - Delivery-fee-related support contacts per assigned customer within seven days, using consistent tagging. - Orders and revenue per assigned customer; average order value and free-delivery eligibility as diagnostics.

Guardrails: Contribution margin per assigned customer, if reliably available; otherwise net revenue per assigned customer as an explicitly limited proxy. Also monitor basket/checkout errors. Chloe and Ines must approve numerical economic and reliability tolerances before launch.

Sample size and schedule

Chloe’s smallest worthwhile order-conversion change is 0.5 percentage points: 8.0% to 8.5%, not a 0.5% relative lift.

At 5% two-sided significance and 80% power, detecting that increase requires approximately 48,000 customers per arm, or 96,000 total.

Ines must first verify that the 8% baseline matches the eligibility rules and seven-day outcome window; recalculate and freeze the sample target if necessary.

Do not count returning customers again. Expected recruitment is approximately:

  • Week 1: 40,000 unique customers.
  • Each subsequent week: 20,000 additional customers.
  • Four weeks: approximately 100,000 unique customers.

Enroll for at least four complete weeks and until the frozen sample target is reached, ending at a full-week boundary. Then wait seven days for the final cohort’s outcomes: roughly five weeks to readout.

Monitoring and decision

Ravi validates assignment persistence, fee accuracy, exposure/order linkage and sample-ratio balance before ramp-up.

Daily checks are for instrumentation and predefined safety breaches—not efficacy stopping. No “stop when significant,” sample extensions based on results, or repeated significance decisions.

Ines delivers one final primary analysis. Chloe’s default rollout rule is: - Estimated conversion gain at least +0.5 points; - 95% confidence interval excludes zero; - Guardrails pass.

Otherwise, do not roll out on this result alone. A non-significant result does not prove no harm. Secondary trust improvements inform follow-up, not a post-hoc replacement for the primary success rule.

Claims the judge couldn’t find in the brief

  • The current experience includes a pricing service and eligibility rules that can be used to show the fee.
  • Signed-out customers can be assigned using a stable first-party identifier, with identity merging and cross-device limitations to document.
  • Delivery-fee-related support contacts can be measured per assigned customer using consistent tagging.
  • Contribution margin or net revenue per assigned customer is available as a guardrail metric.
  • Staff, bots, test accounts, non-empty orderable baskets, and basket/checkout errors can be identified and excluded or monitored.

Grades and run details

Decision model 88 · LLM judge 8 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly7%
  • passAddresses the actual decision52%
  • partialRespects explicit constraints24%
  • passIdentifies material uncertainty73%
  • passAvoids unsupported claims69%
  • passProduces the required deliverable47%
  • passTests one change at a time97%
  • passFixes the duration up front82%
  • passAn unambiguous primary metric76%
  • partialDecision rule written before the test34%
  • passSized from the real traffic59%
  • partialGuardrails with thresholds99%
Run
Run
#1
Time to output
47 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 12 checks passed

The spec is well-structured and fixes the draft's main issues; a PM could quickly add guardrail thresholds and lock the duration to full weeks.

Output

Experiment spec: delivery-fee visibility in basket

Owner: Checkout PM Build: Ravi · Analysis: Ines · Approval: Chloe

Objective and hypothesis

Test whether showing delivery costs before checkout increases orders by reducing surprises—or decreases orders by discouraging customers earlier.

Last quarter, 12% of support contacts concerned unexpected delivery fees. This motivates the test but is not the percentage of customers affected.

Experience

Control: Current basket; delivery fee first appears at the final checkout step.

Variant: Show a clearly labelled delivery-fee line and fee-inclusive total in the basket. Update both when the basket changes. Use the production pricing logic: £3.99 delivery, free for orders over £60. Ravi must verify boundary behaviour at exactly £60 and match existing checkout pricing.

Keep the final checkout fee display unchanged. Everything else stays identical.

Do not add the free-delivery banner. It could independently encourage larger baskets, making it impossible to attribute results to earlier fee visibility. Test it separately.

Eligibility and assignment

Include customers viewing an orderable basket in supported markets. Apply existing employee, test-account and bot exclusions identically, before assignment.

Randomise 50/50 by customer, not session, at their first eligible basket view. Persist assignment across visits and devices where identity is available. Use a stable anonymous identifier for guests and a documented account-linking rule; Ravi and Ines must validate identity handling before launch.

Each customer enters the experiment once. Analyse everyone assigned, regardless of whether they start checkout or successfully see the treatment.

Metrics

Primary: Seven-day order conversion: customers placing at least one order within seven days of assignment ÷ all assigned eligible basket viewers.

This captures Chloe’s concern about discouraging checkout entry. Orders ÷ checkout starts is unsuitable as primary: the treatment can change who starts checkout.

Secondary/diagnostic: - Checkout-start rate per assigned customer. - Checkout completion among starters, interpreted cautiously because starter populations may differ. - Revenue per assigned customer, including non-buyers. - AOV and proportion of orders qualifying for free delivery. - Unexpected-fee support contacts per assigned customer within seven days, using consistent contact classification.

Operational guardrails: Basket/checkout errors, incorrect fee calculations and page performance. These can trigger a safety rollback, not an early declaration of success.

Sample size and duration

Interpret Chloe’s minimum worthwhile change as 0.5 percentage points absolute: 8.0% to 8.5%, not a 0.5% relative lift.

With a matching 8% baseline, two-sided 5% significance and 80% power, plan for approximately 95,000 unique customers total, evenly split. Ines must verify the historical baseline uses the same seven-day definition and finalise the calculation before launch.

Do not count 40,000 weekly viewers as entirely new experiment entrants. Estimated cumulative reach is 40,000 in week one, then roughly 20,000 additional customers weekly: about four weeks to enrol the target, plus seven days for outcomes to mature.

Lock the enrolment target before launch; stop enrolment when reached, then wait seven days. Extend only if unique-customer accrual is slower—not based on results.

Analysis and decision

Ines reports the intention-to-treat conversion difference, 95% confidence interval and absolute rates. Check allocation imbalance, logging, identity consistency and treatment exposure before interpreting results.

No daily significance-based stopping. Daily dashboards are for operational safety only.

Recommend rollout if the primary improvement is statistically significant, the point estimate is at least +0.5 percentage points, and there are no unresolved operational or economic concerns.

Otherwise, retain control pending review. Distinguish evidence of harm, evidence ruling out a worthwhile gain, and an inconclusive result; “not significant” does not mean “no effect.”

Launch approval

Ravi signs off implementation and QA; Ines signs off measurement and the locked analysis plan; Chloe approves the isolated treatment, decision rule and any economic guardrail thresholds before launch.

Grades and run details

Decision model 83 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly7%
  • passAddresses the actual decision65%
  • partialRespects explicit constraints21%
  • passIdentifies material uncertainty45%
  • passAvoids unsupported claims66%
  • passProduces the required deliverable59%
  • passTests one change at a time100%
  • passFixes the duration up front77%
  • passAn unambiguous primary metric84%
  • partialDecision rule written before the test33%
  • partialSized from the real traffic37%
  • partialGuardrails with thresholds98%
Run
Run
#1
API response time
36 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyMixedMixedRight
Opus 5.5 · Claude

All claims about the current situation are directly from the brief or derived by correct arithmetic.

GPT-6 Astra · ChatGPT

It uses the supplied numbers correctly but invents current systems and data capabilities such as an existing pricing service, stable first-party identifiers, identity merging, support-contact tagging, and margin/revenue guardrail data.

GPT-6.1 Sol · API

All claims about the current situation are taken directly from the brief or follow logically from it.

Respects explicit constraintsWrongRightRight
Opus 5.5 · Claude

The output is approximately 780 words, exceeding the 'under 700 words' limit.

GPT-6 Astra · ChatGPT

It is under 700 words, addressed to Ravi, Ines, and Chloe, and its proposed customer randomisation, fixed duration, and no-peeking checks would enforce the key constraints.

GPT-6.1 Sol · API

The output is under 700 words, addresses the named readers, and fixes the draft's problems.

Avoids unsupported claimsRightMixedRight
Opus 5.5 · Claude

Hypotheses and forecasts are clearly labelled, and no confident claims go beyond the supplied evidence.

GPT-6 Astra · ChatGPT

It presents several current capabilities and data sources as available without support, and its rollout rule is stricter than the supplied evidence requires.

GPT-6.1 Sol · API

No interpretations or forecasts are presented as established facts; the 12% clarification is directly supported.

Produces the required deliverableMixedRightRight
Opus 5.5 · Claude

The spec is over the 700-word limit, so it does not fully meet the requested form.

GPT-6 Astra · ChatGPT

It is a usable experiment spec for the named readers, with scope, assignment, metrics, sizing, monitoring, and decision rule, needing only light edits.

GPT-6.1 Sol · API

The spec is complete, in the right form, and usable by Ravi, Ines, and Chloe with light edits.

Fixes the duration up frontRightRightMixed
Opus 5.5 · Claude

It sets a fixed 4-week duration and explains that daily peeking inflates false positives.

GPT-6 Astra · ChatGPT

It replaces daily significance stopping with a fixed four-week enrollment plus seven-day outcome window and explains that daily efficacy stopping is not allowed.

GPT-6.1 Sol · API

It does not set a fixed calendar duration in full weeks, and does not explain that daily stopping inflates false positives.

Decision rule written before the testRightWrongRight
Opus 5.5 · Claude

All outcomes (significant up, significant down, non-significant with CI excluding -0.5pp, non-significant with CI including -0.5pp) map to stated actions.

GPT-6 Astra · ChatGPT

It does not map every outcome to an action, especially a flat or small positive result, and omits the supplied evidence-based rule to ship if orders are no worse than 0.5pp down while fee contacts fall.

GPT-6.1 Sol · API

It maps significant positive result, harm, and inconclusive outcomes to actions (rollout or retain control) before the test.

Sized from the real trafficRightRightWrong
Opus 5.5 · Claude

Sample size is calculated from the 8% baseline and 0.5pp effect, duration follows from traffic accumulation, and it runs full weeks.

GPT-6 Astra · ChatGPT

The sample size follows from the 8% baseline and 0.5pp MDE, and the duration follows from weekly unique-customer accumulation in whole weeks.

GPT-6.1 Sol · API

Sample size is correct, but the test stops when the enrolment target is reached rather than running fixed whole weeks.

All got wrong 1

Guardrails with thresholdsWrongWrongWrong
Opus 5.5 · Claude

Guardrail metrics are listed but no thresholds are given for average order value, revenue, or support contacts to block rollout; only order conversion has a stop threshold.

GPT-6 Astra · ChatGPT

Guardrails are named but lack numerical thresholds that would block rollout, deferring them to later approval.

GPT-6.1 Sol · API

Guardrail metrics are listed but no explicit thresholds (e.g., for support contacts or AOV) are set to block rollout.

All got right 4

Addresses the actual decisionRightRightRight
Opus 5.5 · Claude

The spec commits to a clear decision framework with rules for shipping, not shipping, and inconclusive results, framed for Chloe.

GPT-6 Astra · ChatGPT

It gives Chloe a clear pre-registered rollout rule and states the result conditions that would change the decision.

GPT-6.1 Sol · API

The spec includes a clear decision rule for rollout based on results, framed for Chloe.

Identifies material uncertaintyRightRightRight
Opus 5.5 · Claude

It identifies the uncertain per-customer baseline over 4 weeks and the inconclusive outcome, with a plan to resolve the baseline.

GPT-6 Astra · ChatGPT

It names material uncertainties such as baseline eligibility, identity merging, cross-device limits, and guardrail tolerances, and assigns pre-launch verification or approval.

GPT-6.1 Sol · API

It acknowledges inconclusive results and says to distinguish harm, no worthwhile gain, and uncertainty.

Tests one change at a timeRightRightRight
Opus 5.5 · Claude

The free-delivery banner is removed from the variant and the spec explains why bundling would make results uninterpretable.

GPT-6 Astra · ChatGPT

It keeps the free-delivery banner out of the variant and explains that bundling it would prevent attribution.

GPT-6.1 Sol · API

It explicitly keeps the free-delivery banner out of the variant and explains why bundling would confound results.

An unambiguous primary metricRightRightRight
Opus 5.5 · Claude

Order conversion is the single primary metric with a clear rationale, and a sample-ratio mismatch check is included.

GPT-6 Astra · ChatGPT

It names one primary metric, explains why checkout completion is misleading, and includes a sample-ratio balance check.

GPT-6.1 Sol · API

It names one primary metric (seven-day order conversion), explains its direction, and includes a sample ratio check.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 90% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI87.588.52None
2Sonnet 5.5withAPI81.388.52None
3Opus 5.5withClaude85.476.92None
4GPT-6 LunawithAPI85.473.12None
5GPT-6 AstrawithChatGPT87.565.42None
6Gemini 3.8 FlashwithAPI60.438.52None
7Gemini 3.5 Flash-LitewithGemini37.515.422 capped

About the task

The PM job

Specifying an A/B test before running it.

Why it matters

Most tests are underpowered, or measure a metric nobody agrees is good.

What good looks like

  • One primary metric everyone agrees on the direction of
  • Decision rule stated up front
  • Guardrails
  • Power considered

Deliberately not measured

    Capability tested

    Test design

    The failure we’re looking for

    A test with no decision attached

    Grading

    Decision model and LLM judge, calibrated against a blind PM review