Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 58% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable92% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims60% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly63% pass
    The output claims that users charged on day 7 “churn or ask for refunds”, which is not in the supplied context and is not supported by the data.
    Opus 5.5 · Claude · Fitness app first week
  3. Identifies material uncertainty69% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review the first-week experience of our fitness app, described below, and recommend the changes you would make, in priority order, as a memo to the growth PM. Keep it under 500 words.

ScenarioStride is a home-workout app. It costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.
First week, step by stepDay 0: nine onboarding questions (goal, age, height, weight, fitness level, equipment, injuries, preferred days, preferred time), then the paywall to start the trial, then a plan with the first workout scheduled for the next morning. Every workout is 30–45 minutes. Day 1: a 7am push notification, “Your workout is ready.” Days 2–7: one push a day at the preferred time. Missing a workout moves the plan on to the next one; there is no shorter or catch-up option. The streak counter first appears after seven days.
FunnelInstall → finished onboarding 71% → started trial 52% → first workout 48% → three workouts in week one 19%. Users who complete three workouts in week one are four times as likely to still be subscribed at day 60.
Team planThe design team wants to spend this quarter redesigning the onboarding questionnaire and the paywall screen.
What a strong answer does

Focus on getting new users to three workouts in week one (shorter first workouts, a same-day first session, catch-up options, earlier progress feedback) rather than onboarding and paywall polish.

Case

v1.6 · synthetic · consumer, mobile

Usable with a quick edit9 of 9 checks passed

The memo is clear, evidence-based, and actionable with only trivial copy edits. No re-checking of evidence or redoing of reasoning is needed.

Output

To: Growth PM Subject: Prioritise first-week activation over a full onboarding redesign

Stride asks users to commit before delivering value, then makes consistency difficult: the first workout is delayed, every session requires 30–45 minutes, and missed sessions have no recovery path. I would redirect most of this quarter’s effort from questionnaire/paywall redesign to helping users complete—and repeat—a workout.

Three week-one workouts is a promising activation signal, but the fourfold day-60 retention association is not proof that driving three completions will cause higher retention; motivated users may do both.

Recommended changes, in priority order

1. Make the first workout immediate and achievable. After trial activation, offer “Start now” with a 10–15-minute beginner-friendly session, alongside the option to schedule it. Keep longer sessions available, but do not make 30–45 minutes the only entry point. Preserve injury and equipment screening before recommending exercise.

2. Build a forgiving path to three sessions. Offer shorter alternatives throughout the week. After a missed workout, let users reschedule, repeat, or choose a lighter session rather than automatically advancing. Show progress toward a flexible weekly goal from day zero; do not wait seven days to introduce feedback. Avoid punitive streak mechanics that turn one missed day into perceived failure.

3. Make reminders support the user’s commitment. Replace the fixed day-one 7am push with the user’s selected time. Ask users to confirm their first session’s timing and send a specific prompt, such as “Your 12-minute starter workout is ready.” Adapt reminders after missed sessions and provide notification controls.

4. Streamline onboarding and clarify the trial—without a wholesale redesign. Keep questions necessary for safe, relevant recommendations; defer other questions until their answers improve the experience. Clearly state the £9.99 monthly charge, billing date, and cancellation process. Send a pre-charge reminder. Test focused changes before committing the quarter to visual redesign.

Measurement and rollout

First confirm whether funnel percentages are install-based or step-to-step, and define workout completion consistently. Instrument time to first workout, workout duration, missed-session recovery, and reasons for abandonment.

Run staged randomised tests, starting with immediate/short workouts, then recovery mechanics. Use first-workout completion and three-session activation as leading indicators; judge success by day-60 paid retention and revenue per assigned user—not activated users alone. Monitor trial starts, cancellations, refunds, notification opt-outs, and safety complaints.

Decision: prioritise activation and repeatability, while making low-cost onboarding and billing improvements now. Revisit a full redesign once evidence identifies which pre-trial friction materially limits retained-user growth.

Check by check

Got right · 9
  • Uses the supplied evidence correctlyAll current-state facts it relies on are taken correctly from the brief, with no invented figures or system details.
  • Addresses the actual decisionIt commits early to prioritising activation over a full onboarding redesign and says a redesign would be revisited only if evidence identifies meaningful pre-trial friction.
  • Respects explicit constraintsIt is a memo to the growth PM, stays under the 500-word limit, and its proposals would enforce the requested focus in practice.
  • Identifies material uncertaintyIt names causal uncertainty in the retention signal, asks to confirm funnel bases and workout-completion definitions, and states what would change the call.
  • Avoids unsupported claimsIt labels the retention link as an association rather than proof and does not present untested causes or forecasts as established fact.
  • Produces the required deliverableThe memo is complete, correctly addressed, within the required length, and actionable with at most light edits.
  • Prioritises by impactChanges are ranked 1–4 by likely impact on activation, starting with making the first workout immediate and building a forgiving path to three sessions.
  • Activation defined by what predicts retentionIt defines activation as three first-week workouts, uses the supplied day-60 retention link, and ranks fixes by their effect on that behaviour.
  • Gets the base of every number rightThe output does not compute step-to-step percentages, and the only cited statistic, the fourfold retention link, matches the supplied funnel note.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly31%
  • passAddresses the actual decision92%
  • passRespects explicit constraints26%
  • passIdentifies material uncertainty84%
  • passAvoids unsupported claims64%
  • passProduces the required deliverable89%
  • passPrioritises by impact98%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right80%
Run
Run
#1
API response time
16 s
Submitted
29 Sept 2026

Needs more than a quick edit6 of 9 checks passed

It contains a miscomputed funnel rate and an unsupported causal claim, so an editor would need to re-check the evidence before using it.

Output

To: Growth PM Re: First-week experience: recommended changes, in priority order

Summary: The design team's plan targets the wrong part of the funnel. The biggest leak that matters is after the trial starts. Only 40% of users who do a first workout go on to do three (48% → 19%), and three workouts is the behavior linked to 4x day-60 retention. Onboarding and the paywall do lose people, but the week-one experience is where I'd invest first.

Priorities

  1. Let users work out on Day 0. Today the first workout is the next morning, so motivation peaks at the paywall and then goes unused. After trial start, offer "Start your first workout now" with a 10–15 minute option. Also, 4% of trial starters (52% → 48%) never do a first workout, and this fixes that too.
  1. Add short and catch-up workouts. Every session is 30–45 minutes, and a missed one is simply skipped. That punishes exactly the users at risk. Add 10–15 minute versions and a "catch up" option after a miss, so the plan flexes instead of moving on. This is likely the biggest lever on the 48% → 19% drop.
  1. Fix the notifications. Day 1 goes out at 7am, not the time users told us they prefer, which looks like a bug and wastes the most important nudge. Beyond fixing that, make the messages specific ("15 minutes, no equipment, ready when you are"). Add a gentle recovery message after a missed day instead of continuing the normal schedule.
  1. Show the streak from Day 1. It currently appears after seven days, when the habit has already formed or failed. Show it after the first workout. Make it forgiving (e.g. weekly goals or a freeze) so one missed day doesn't feel like failure.
  1. Trim onboarding, don't redesign it. Nine questions likely contribute to the 29% who drop before finishing. Keep what personalization needs (goal, level, equipment, injuries, preferred days and time) and defer height, weight and age, or make them optional. This is a small change and can ship alongside the items above.
  1. Test the paywall last. Only 27% of onboarded users drop there, and a card-up-front trial is already a strong filter. Once the trial experience is better, we can test the paywall with better evidence. Consider a trial-end reminder on day 5 to build trust.

Suggested redirect for design: Rather than redesigning the questionnaire and paywall, spend the quarter on the Day 0–7 experience: the immediate first workout, short sessions, catch-up flow, and streak.

Caveats and measurement - The 4x retention figure is correlational. Motivated users may simply do more workouts. Run these changes as A/B tests with day-60 retention as the outcome, not just workout counts. - Primary metric: share of trial starters completing three workouts in week one (currently about 37%, or 19/52). - Guardrails: trial start rate, and whether shorter workouts reduce perceived value or conversion to paid.

Sequencing: items 1, 3 and 5 are quick to ship, and items 2 and 4 need more design and content work. I'd begin experiments on 1 and 3 within the first few weeks.

Check by check

Got wrong · 3
  • Uses the supplied evidence correctlyIt includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
  • Avoids unsupported claimsThe motivation-peak and punishment interpretations are presented as fact rather than as hypotheses the evidence only suggests.
  • Gets the base of every number rightThe 52% to 48% drop is misstated as 4% of trial starters; the correct share of trial starters is about 7.7%, or 4 percentage points of installs.
Got right · 6
  • Addresses the actual decisionIt commits early to a clear ranked list for the Growth PM and notes A/B tests and guardrails that would change the call.
  • Respects explicit constraintsIt is a memo to the Growth PM, prioritized, and under the 500-word limit.
  • Identifies material uncertaintyIt names the causal uncertainty in the 4x retention figure and the risk that shorter workouts reduce perceived value, and says how to resolve them with A/B tests and guardrails.
  • Produces the required deliverableThe requested memo is present, complete, actionable, and within the requested length.
  • Prioritises by impactFixes are ranked by likely impact on reaching three workouts and retention.
  • Activation defined by what predicts retentionIt defines activation as three workouts in week one and judges fixes by their effect on that retention-predictive behavior.

Claims the judge couldn’t find in the brief

  • 4% of trial starters never do a first workout.
  • Motivation peaks at the paywall and then goes unused by the next morning.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly22%
  • passAddresses the actual decision97%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty76%
  • partialAvoids unsupported claims56%
  • passProduces the required deliverable93%
  • passPrioritises by impact94%
  • passActivation defined by what predicts retention100%
  • failGets the base of every number right76%
Run
Run
#1
API response time
19 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT97.2100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude86.165.02None
5Sonnet 5.5withAPI77.845.02None
6Gemini 3.5 Flash-LitewithGemini61.155.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review