Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable93% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims54% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly54% pass
    It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
    Sonnet 5.5 · API · Fitness app first week
  3. Identifies material uncertainty59% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review the first-week experience of our fitness app, described below, and recommend the changes you would make, in priority order, as a memo to the growth PM. Keep it under 500 words.

What the model was given4 items: Scenario, First week, step by step, Funnel, Team plan
ScenarioStride is a home-workout app. It costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.
First week, step by stepDay 0: nine onboarding questions (goal, age, height, weight, fitness level, equipment, injuries, preferred days, preferred time), then the paywall to start the trial, then a plan with the first workout scheduled for the next morning. Every workout is 30–45 minutes. Day 1: a 7am push notification, “Your workout is ready.” Days 2–7: one push a day at the preferred time. Missing a workout moves the plan on to the next one; there is no shorter or catch-up option. The streak counter first appears after seven days.
FunnelInstall → finished onboarding 71% → started trial 52% → first workout 48% → three workouts in week one 19%. Users who complete three workouts in week one are four times as likely to still be subscribed at day 60.
Team planThe design team wants to spend this quarter redesigning the onboarding questionnaire and the paywall screen.
What a strong answer doesThe answer key the graders mark against

Focus on getting new users to three workouts in week one (shorter first workouts, a same-day first session, catch-up options, earlier progress feedback) rather than onboarding and paywall polish.

Case

Warm-up · v1.8 · synthetic · consumer, mobile

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit6 of 9 checks passed

A PM would need to re-check and soften unsupported causal claims before acting. The core recommendation is usable, but the evidence framing is not safe as written.

Output

MEMO

TO: Growth PM FROM: Product & Growth Review SUBJECT: Week 1 Retention & Experience Optimization

Executive Summary

Push back on the design team’s plan to redesign onboarding and the paywall. Converting 71% through onboarding and 52% into a card-upfront trial is exceptionally strong.

Our growth bottleneck is Week 1 activation: while 48% log a first workout, only 19% reach three workouts—the critical threshold that drives 4x Day-60 retention. We are losing trialists during their first week because of delayed gratification, rigid scheduling, and high initial friction.

Here are the four prioritized changes to drive 3-workout completion.

---

Recommended Changes (Priority Order)

#### 1. Instant Gratification on Day 0 (“Start Now” Option) * Problem: Scheduling the first session for the next morning wastes peak user motivation right after purchase. * Fix: Offer an immediate, optional 5–10 minute "Quick Start / Assessment" on Day 0. Completing this counts as Workout #1, immediately advancing users toward the 3-workout milestone.

#### 2. Introduce "Micro-Workouts" & Flexible Session Lengths * Problem: A rigid 30–45 minute commitment is daunting for home workouts, leading to missed days. When a day is missed, the plan simply skips ahead without recourse, inducing guilt and churn. * Fix: Introduce 10- and 20-minute “time-crunched” alternatives when a user opens a session or is at risk of missing one. A 10-minute workout must count toward the 3-workout goal. Never skip a workout automatically; let users swap, reschedule, or condense.

#### 3. Surface Streaks & Milestones on Day 1 (Not Day 7) * Problem: Hiding the streak counter until Day 7 removes the most effective psychological hook during the make-or-break trial period. * Fix: * Show the streak counter immediately upon completing Workout #1. * Explicitly gamify the activation metric: show a progress bar toward the "Week 1 Foundation Badge" (1/3, 2/3, 3/3 workouts completed).

#### 4. Action-Oriented, Personalized Push Notifications * Problem: A generic 7:00 AM push ("Your workout is ready") is transactional and easily dismissed. * Fix: Leverage the rich onboarding data (goals, fitness level) to trigger personalized, high-context pushes: * Copy: “Crush your [Goal] today: 15 mins to hit 2/3 of your Week 1 target.” * Timing: Align strictly to their chosen time, with an evening "last chance" nudge offering an express 10-minute session if uncompleted.

---

Next Steps

Pause the onboarding/paywall redesign. Run a two-week sprint to implement immediate Day 0 activation, 15-minute workout variants, and Day 1 streak mechanics. Our target metric is moving Week 1 (3+ workouts) from 19% to 30%.

Claims the judge couldn’t find in the brief

  • The first week has high initial friction.

Grades and run details

Decision model 67 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly40%
  • passAddresses the actual decision98%
  • passRespects explicit constraints54%
  • failIdentifies material uncertainty88%
  • failAvoids unsupported claims33%
  • passProduces the required deliverable98%
  • passPrioritises by impact91%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right93%
Run
Run
#1
API response time
15 s
Submitted
1 Oct 2026

Needs more than a quick edit5 of 9 checks passed

The unsupported claim about refund/churn means a PM would need to verify and potentially correct a factual assumption before acting on the memo.

Output

Stride: fixing the first week

Sep 25, 2026 · @Martin Slaney

Summary

Spend this quarter on the first seven days, not the questionnaire and paywall. The biggest leak is between first workout (48%) and three workouts (19%), and three workouts is the behaviour tied to day-60 retention. Redesigning the top of the funnel pours more trial users into a week that currently loses most of them.

What the funnel says

StepShare of installsKept from previous step
Finished onboarding71%71%
Started trial52%73%
First workout48%92%
Three workouts in week one19%40%

Trial-to-first-workout is healthy. Six in ten people who work out once never reach three. Because the card is taken up front, those users are charged on day 7 and then churn or ask for refunds.1

The week-one design explains the drop:

  • Too long. Every session is 30–45 minutes, for everyone, from day one.
  • Missing one punishes you. The plan moves on; there is no shorter or catch-up option.
  • Day 0 intent is wasted. The first workout is the next morning, not now.
  • Day 1 push ignores the user. It fires at 7am even though we asked their preferred time.
  • Progress is hidden. The streak appears only after seven days, so it rewards people who have already succeeded.

Changes, in priority order

  1. Short sessions in week one. Default to 10–15 minute workouts for days 1–7, with the full session as an opt-in. Lowest effort, biggest lever on reaching three.
  2. Recover missed days instead of skipping. A missed workout offers a 10-minute version the same evening or next day, not the next plan step.
  3. First workout on day 0. Straight after the plan screen, offer a 5–10 minute starter. It counts toward the three.
  4. Fix notifications. Honour the preferred time from day 1. Replace "Your workout is ready" with the specific session and length ("12 min, no equipment"). Add one recovery nudge after a miss.
  5. Show progress to three from day 1. A "1 of 3 this week" tracker replaces the hidden streak and makes the target explicit.
  6. Trim onboarding, lightly. Keep only questions that change the first plan; move height, weight and preferred days to after the first workout. A sprint, not a quarter.
  7. Leave the paywall. 73% of finishers start the trial; it is not the constraint.2

Measurement and the ask

The 4× day-60 figure is a correlation; motivated users may simply do more workouts. Ship changes 1–5 as A/B tests so we learn whether reaching three causes retention.3

  • Primary metric: share of trial starters completing three workouts in days 0–7 (today roughly 37%).
  • Guardrails: trial-to-paid conversion, day-7 refund requests, day-60 retention.

Ask: redirect the design team's quarter from questionnaire and paywall to the week-one experience, with a light onboarding trim as a side task.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    Source: ScenarioIt costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.

    What we’d changeRemove it: the brief says nothing about what these users do at day 7.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeA 73% step rate doesn't show the paywall isn't a constraint: a quarter of people who finish onboarding stop there.

  3. 3Test or gate too weakTighten the testQuick edit

    What we’d changeA/B tests show whether these changes improve retention. They won't show that reaching three workouts is itself the cause.

Claims the judge couldn’t find in the brief

  • Because the card is taken up front, those users are charged on day 7 and then churn or ask for refunds.

Grades and run details

Decision model 67 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly39%
  • passAddresses the actual decision99%
  • partialRespects explicit constraints13%
  • passIdentifies material uncertainty72%
  • partialAvoids unsupported claims51%
  • passProduces the required deliverable92%
  • passPrioritises by impact98%
  • passActivation defined by what predicts retention100%
  • failGets the base of every number right19%
Run
Run
#1
Time to output
49 s
Submitted
25 Sept 2026

Needs more than a quick edit6 of 9 checks passed

It contains a miscomputed funnel rate and an unsupported causal claim, so an editor would need to re-check the evidence before using it.

Output

To: Growth PM Re: First-week experience: recommended changes, in priority order

Summary: The design team's plan targets the wrong part of the funnel. The biggest leak that matters is after the trial starts. Only 40% of users who do a first workout go on to do three (48% → 19%), and three workouts is the behavior linked to 4x day-60 retention. Onboarding and the paywall do lose people, but the week-one experience is where I'd invest first.

Priorities

  1. Let users work out on Day 0. Today the first workout is the next morning, so motivation peaks at the paywall and then goes unused. After trial start, offer "Start your first workout now" with a 10–15 minute option. Also, 4% of trial starters (52% → 48%) never do a first workout, and this fixes that too.
  1. Add short and catch-up workouts. Every session is 30–45 minutes, and a missed one is simply skipped. That punishes exactly the users at risk. Add 10–15 minute versions and a "catch up" option after a miss, so the plan flexes instead of moving on. This is likely the biggest lever on the 48% → 19% drop.
  1. Fix the notifications. Day 1 goes out at 7am, not the time users told us they prefer, which looks like a bug and wastes the most important nudge. Beyond fixing that, make the messages specific ("15 minutes, no equipment, ready when you are"). Add a gentle recovery message after a missed day instead of continuing the normal schedule.
  1. Show the streak from Day 1. It currently appears after seven days, when the habit has already formed or failed. Show it after the first workout. Make it forgiving (e.g. weekly goals or a freeze) so one missed day doesn't feel like failure.
  1. Trim onboarding, don't redesign it. Nine questions likely contribute to the 29% who drop before finishing. Keep what personalization needs (goal, level, equipment, injuries, preferred days and time) and defer height, weight and age, or make them optional. This is a small change and can ship alongside the items above.
  1. Test the paywall last. Only 27% of onboarded users drop there, and a card-up-front trial is already a strong filter. Once the trial experience is better, we can test the paywall with better evidence. Consider a trial-end reminder on day 5 to build trust.

Suggested redirect for design: Rather than redesigning the questionnaire and paywall, spend the quarter on the Day 0–7 experience: the immediate first workout, short sessions, catch-up flow, and streak.

Caveats and measurement - The 4x retention figure is correlational. Motivated users may simply do more workouts. Run these changes as A/B tests with day-60 retention as the outcome, not just workout counts. - Primary metric: share of trial starters completing three workouts in week one (currently about 37%, or 19/52). - Guardrails: trial start rate, and whether shorter workouts reduce perceived value or conversion to paid.

Sequencing: items 1, 3 and 5 are quick to ship, and items 2 and 4 need more design and content work. I'd begin experiments on 1 and 3 within the first few weeks.

Claims the judge couldn’t find in the brief

  • 4% of trial starters never do a first workout.
  • Motivation peaks at the paywall and then goes unused by the next morning.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly21%
  • passAddresses the actual decision100%
  • passRespects explicit constraints28%
  • passIdentifies material uncertainty76%
  • partialAvoids unsupported claims56%
  • passProduces the required deliverable94%
  • passPrioritises by impact94%
  • passActivation defined by what predicts retention100%
  • failGets the base of every number right75%
Run
Run
#1
API response time
19 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Addresses the actual decisionRightMixedRight
Gemini 3.8 Flash · API

It clearly commits to prioritising week-one activation over onboarding/paywall redesign and gives a ranked set of changes for the growth PM.

Opus 5.5 · Claude

The output commits to redirecting the design team but does not explicitly state what result or condition would change that decision, even though the brief requires it.

Sonnet 5.5 · API

It commits early to a clear ranked list for the Growth PM and notes A/B tests and guardrails that would change the call.

Identifies material uncertaintyWrongRightRight
Gemini 3.8 Flash · API

It does not name material unknowns or conditions that would change the call, such as whether shorter workouts preserve habit formation or whether the 4x relationship is causal.

Opus 5.5 · Claude

It names the key unknown (correlation vs causation for the 4× day‑60 figure) and proposes an A/B test to resolve it.

Sonnet 5.5 · API

It names the causal uncertainty in the 4x retention figure and the risk that shorter workouts reduce perceived value, and says how to resolve them with A/B tests and guardrails.

Gets the base of every number rightRightMixedWrong
Gemini 3.8 Flash · API

The percentages and 4x relationship are used against the correct funnel steps and denominators.

Opus 5.5 · Claude

All derived percentages (73 % trial start from onboarding, 60 % drop, 37 % three workouts among trial starters) use the correct denominators from the supplied funnel.

Sonnet 5.5 · API

The 52% to 48% drop is misstated as 4% of trial starters; the correct share of trial starters is about 7.7%, or 4 percentage points of installs.

All got wrong 2

Uses the supplied evidence correctlyWrongWrongWrong
Gemini 3.8 Flash · API

It states unsupported current-situation claims such as high initial friction, missed days, guilt, churn, and a psychological hook as facts rather than labelled hypotheses.

Opus 5.5 · Claude

The output claims that users charged on day 7 “churn or ask for refunds”, which is not in the supplied context and is not supported by the data.

Sonnet 5.5 · API

It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.

Avoids unsupported claimsWrongWrongWrong
Gemini 3.8 Flash · API

It presents several causes and psychological effects as established facts when the supplied evidence only supports correlations or funnel drop-offs.

Opus 5.5 · Claude

The statement that users “churn or ask for refunds” is presented as fact without labelling it as an assumption, and the supplied evidence does not support it.

Sonnet 5.5 · API

The motivation-peak and punishment interpretations are presented as fact rather than as hypotheses the evidence only suggests.

All got right 4

Respects explicit constraintsRightRightRight
Gemini 3.8 Flash · API

It is a memo to the growth PM, under 500 words, and recommends changes in priority order.

Opus 5.5 · Claude

The output is a memo to the growth PM, under 500 words, with changes in priority order, respecting all given constraints.

Sonnet 5.5 · API

It is a memo to the Growth PM, prioritized, and under the 500-word limit.

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

The requested memo is present, complete, addressed to the growth PM, and actionable within the length limit.

Opus 5.5 · Claude

The request deliverable (a priority‑ordered memo to the growth PM, under 500 words) is present and immediately usable.

Sonnet 5.5 · API

The requested memo is present, complete, actionable, and within the requested length.

Prioritises by impactRightRightRight
Gemini 3.8 Flash · API

The fixes are ranked in a plausible impact order centred on reaching three workouts in week one.

Opus 5.5 · Claude

Fixes are ranked by likely impact on reaching three workouts, with low‑effort, high‑lever actions first.

Sonnet 5.5 · API

Fixes are ranked by likely impact on reaching three workouts and retention.

Activation defined by what predicts retentionRightRightRight
Gemini 3.8 Flash · API

It defines activation as three workouts in week one, links that to day-60 retention, and ranks fixes by their effect on that activation event.

Opus 5.5 · Claude

The activation event (three workouts in week one) is defined by its link to day‑60 retention, and fixes are judged by their effect on that event.

Sonnet 5.5 · API

It defines activation as three workouts in week one and judges fixes by their effect on that retention-predictive behavior.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 82% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT94.4100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude80.665.02None
5Gemini 3.5 Flash-LitewithGemini69.455.02None
6Sonnet 5.5withAPI77.845.02None
7Gemini 3.8 FlashwithAPI69.450.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review