Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 58% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable92% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims60% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly63% pass
    The output claims that users charged on day 7 “churn or ask for refunds”, which is not in the supplied context and is not supported by the data.
    Opus 5.5 · Claude · Fitness app first week
  3. Identifies material uncertainty69% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review the first-week experience of our fitness app, described below, and recommend the changes you would make, in priority order, as a memo to the growth PM. Keep it under 500 words.

ScenarioStride is a home-workout app. It costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.
First week, step by stepDay 0: nine onboarding questions (goal, age, height, weight, fitness level, equipment, injuries, preferred days, preferred time), then the paywall to start the trial, then a plan with the first workout scheduled for the next morning. Every workout is 30–45 minutes. Day 1: a 7am push notification, “Your workout is ready.” Days 2–7: one push a day at the preferred time. Missing a workout moves the plan on to the next one; there is no shorter or catch-up option. The streak counter first appears after seven days.
FunnelInstall → finished onboarding 71% → started trial 52% → first workout 48% → three workouts in week one 19%. Users who complete three workouts in week one are four times as likely to still be subscribed at day 60.
Team planThe design team wants to spend this quarter redesigning the onboarding questionnaire and the paywall screen.
What a strong answer does

Focus on getting new users to three workouts in week one (shorter first workouts, a same-day first session, catch-up options, earlier progress feedback) rather than onboarding and paywall polish.

Case

v1.6 · synthetic · consumer, mobile

Usable with a quick edit9 of 9 checks passed

It makes the right activation-focused call, ranks concrete changes, and flags the key measurement caveat. A PM could act on it with only light wording edits.

Output

Memo

To: Growth PM Subject: Priorities for Stride’s first week

I would not make a full onboarding and paywall redesign the quarter’s main bet. The clearest retention signal is that people who complete three workouts in week one are much more likely to stay subscribed. Prioritize helping more new users reach that behavior, then address the earlier funnel friction.

  1. Make the first workout easy to start and finish. Users currently wait until the next morning and face a 30–45-minute session. Offer an optional workout immediately, plus a 10–15-minute beginner session. Let people choose when to start rather than defaulting to 7am; the Day 1 notification should respect their chosen time.
  1. Help users recover when they miss a session. A missed workout currently pushes the plan forward, with no shorter or catch-up option. Let users reschedule, swap in a quick workout, or resume without feeling behind. Build the week around three achievable sessions, not a rigid sequence.
  1. Show progress early. Don’t wait seven days to introduce the streak counter. Celebrate the first workout and show progress toward the second and third. Test milestone prompts that make the three-workout goal feel attainable without shaming users who miss a day.
  1. Reduce onboarding friction selectively. Nine questions—including weight and injury details—come before users experience the product. Test asking only what is needed to recommend a safe first workout, and collect the rest later or as optional. Preserve essential safety questions.
  1. Improve trial clarity and measurement. Keep the card-up-front model for now, but ensure the price, trial end date, and cancellation terms are unmistakable. Before committing to a paywall redesign, test specific changes. Also confirm whether funnel percentages are install-based or step-to-step, and instrument workout starts, completions, and reschedules.

Run focused experiments on these changes and judge them by first-workout completion and three workouts in week one—not just trial starts. The reported day-60 relationship is promising, but it is correlational; verify that lifting week-one completion improves retention.

Check by check

Got right · 9
  • Uses the supplied evidence correctlyAll current-situation statements in the output are drawn correctly from the supplied context.
  • Addresses the actual decisionThe memo commits early to not making onboarding/paywall redesign the main bet and states that lifting week-one completion would need to be verified as the condition that could change the call.
  • Respects explicit constraintsIt is a memo to the Growth PM, under 500 words, and respects the requested prioritised recommendation format.
  • Identifies material uncertaintyIt names specific unknowns—funnel percentage basis and whether lifting week-one completion improves retention—and says how to resolve them.
  • Avoids unsupported claimsThe one causal caveat is explicitly labelled as correlational, and confident claims are supported by the supplied evidence.
  • Produces the required deliverableThe required memo is complete, actionable, reader-appropriate, and within length.
  • Prioritises by impactChanges are ranked by likely impact on getting users to three workouts in week one, with onboarding/paywall changes deprioritised.
  • Activation defined by what predicts retentionIt defines three workouts in week one as the activation behaviour linked to day-60 retention and ranks fixes by their effect on that event.
  • Gets the base of every number rightNo percentages are derived incorrectly, and the memo flags ambiguous funnel bases rather than misusing them.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly40%
  • passAddresses the actual decision95%
  • passRespects explicit constraints72%
  • passIdentifies material uncertainty65%
  • passAvoids unsupported claims67%
  • passProduces the required deliverable93%
  • passPrioritises by impact96%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right71%
Run
Run
#1
API response time
13 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT97.2100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude86.165.02None
5Sonnet 5.5withAPI77.845.02None
6Gemini 3.5 Flash-LitewithGemini61.155.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review