Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable93% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims54% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly54% pass
    It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
    Sonnet 5.5 · API · Fitness app first week
  3. Identifies material uncertainty59% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review the first-week experience of our fitness app, described below, and recommend the changes you would make, in priority order, as a memo to the growth PM. Keep it under 500 words.

What the model was given4 items: Scenario, First week, step by step, Funnel, Team plan
ScenarioStride is a home-workout app. It costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.
First week, step by stepDay 0: nine onboarding questions (goal, age, height, weight, fitness level, equipment, injuries, preferred days, preferred time), then the paywall to start the trial, then a plan with the first workout scheduled for the next morning. Every workout is 30–45 minutes. Day 1: a 7am push notification, “Your workout is ready.” Days 2–7: one push a day at the preferred time. Missing a workout moves the plan on to the next one; there is no shorter or catch-up option. The streak counter first appears after seven days.
FunnelInstall → finished onboarding 71% → started trial 52% → first workout 48% → three workouts in week one 19%. Users who complete three workouts in week one are four times as likely to still be subscribed at day 60.
Team planThe design team wants to spend this quarter redesigning the onboarding questionnaire and the paywall screen.
What a strong answer doesThe answer key the graders mark against

Focus on getting new users to three workouts in week one (shorter first workouts, a same-day first session, catch-up options, earlier progress feedback) rather than onboarding and paywall polish.

Case

Warm-up · v1.8 · synthetic · consumer, mobile

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit6 of 9 checks passed

A PM would need to re-check and soften unsupported causal claims before acting. The core recommendation is usable, but the evidence framing is not safe as written.

Output

MEMO

TO: Growth PM FROM: Product & Growth Review SUBJECT: Week 1 Retention & Experience Optimization

Executive Summary

Push back on the design team’s plan to redesign onboarding and the paywall. Converting 71% through onboarding and 52% into a card-upfront trial is exceptionally strong.

Our growth bottleneck is Week 1 activation: while 48% log a first workout, only 19% reach three workouts—the critical threshold that drives 4x Day-60 retention. We are losing trialists during their first week because of delayed gratification, rigid scheduling, and high initial friction.

Here are the four prioritized changes to drive 3-workout completion.

---

Recommended Changes (Priority Order)

#### 1. Instant Gratification on Day 0 (“Start Now” Option) * Problem: Scheduling the first session for the next morning wastes peak user motivation right after purchase. * Fix: Offer an immediate, optional 5–10 minute "Quick Start / Assessment" on Day 0. Completing this counts as Workout #1, immediately advancing users toward the 3-workout milestone.

#### 2. Introduce "Micro-Workouts" & Flexible Session Lengths * Problem: A rigid 30–45 minute commitment is daunting for home workouts, leading to missed days. When a day is missed, the plan simply skips ahead without recourse, inducing guilt and churn. * Fix: Introduce 10- and 20-minute “time-crunched” alternatives when a user opens a session or is at risk of missing one. A 10-minute workout must count toward the 3-workout goal. Never skip a workout automatically; let users swap, reschedule, or condense.

#### 3. Surface Streaks & Milestones on Day 1 (Not Day 7) * Problem: Hiding the streak counter until Day 7 removes the most effective psychological hook during the make-or-break trial period. * Fix: * Show the streak counter immediately upon completing Workout #1. * Explicitly gamify the activation metric: show a progress bar toward the "Week 1 Foundation Badge" (1/3, 2/3, 3/3 workouts completed).

#### 4. Action-Oriented, Personalized Push Notifications * Problem: A generic 7:00 AM push ("Your workout is ready") is transactional and easily dismissed. * Fix: Leverage the rich onboarding data (goals, fitness level) to trigger personalized, high-context pushes: * Copy: “Crush your [Goal] today: 15 mins to hit 2/3 of your Week 1 target.” * Timing: Align strictly to their chosen time, with an evening "last chance" nudge offering an express 10-minute session if uncompleted.

---

Next Steps

Pause the onboarding/paywall redesign. Run a two-week sprint to implement immediate Day 0 activation, 15-minute workout variants, and Day 1 streak mechanics. Our target metric is moving Week 1 (3+ workouts) from 19% to 30%.

Claims the judge couldn’t find in the brief

  • The first week has high initial friction.

Grades and run details

Decision model 67 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly40%
  • passAddresses the actual decision98%
  • passRespects explicit constraints54%
  • failIdentifies material uncertainty88%
  • failAvoids unsupported claims33%
  • passProduces the required deliverable98%
  • passPrioritises by impact91%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right93%
Run
Run
#1
API response time
15 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 9 checks passed

The memo is clear, actionable, and prioritised; only minor stylistic edits would be needed.

Output

To: Growth PM Subject: Prioritise repeat workouts over an onboarding redesign

We should redirect the quarter’s work towards helping people fit workouts into their lives and return after missing one. The questionnaire and paywall deserve attention, but a full redesign is premature.

Assuming the funnel percentages share an install denominator, 92% of trial starters reach a first workout (48/52), but only 40% of first-workout users reach three sessions (19/48). We also lose 29% before onboarding finishes.1 Both matter, but the current plan leaves the substantial repeat-workout drop untouched.

Three workouts is a useful activation hypothesis. The fourfold association with day-60 subscription does not prove causation: more motivated users may both exercise more and stay subscribed.

1. Make repeat workouts achievable, including after a missed day.

Offer a 10–15-minute option alongside full sessions, with suitable adjustments for fitness level, equipment and injuries. After a missed workout, let users reschedule or choose a shorter session rather than automatically advancing the plan.

Requiring 30–45 minutes every time creates a substantial commitment before a routine exists.2 Recovery should be part of the experience, not something users must work out themselves. Test whether these options increase repeat participation and total active minutes, rather than simply inflating completion counts.

2. Offer a first workout immediately and respect the user’s schedule.

After trial activation, offer “Start a short workout now” or let users choose their first session time. Replace the fixed day-one 7am push with their chosen time.

Scheduling everyone for tomorrow wastes initial motivation and consumes part of a seven-day trial before value is experienced. This ranks below repeat-workout flexibility because first-workout conversion among trial starters is already strong.

3. Show progress during the week and make reminders actionable.

Show the first completed session immediately and progress towards a user-chosen weekly goal. Avoid making consecutive-day streaks the default for people who selected fewer workout days.

Use reminders that offer a relevant next step: start, shorten or reschedule. Stop treating every day as another generic prompt. Measure notification opt-outs alongside workout returns.

4. Simplify onboarding through focused experiments.

Keep information needed for safe, relevant workouts upfront, particularly injuries, fitness level and equipment. Test deferring other questions unless their answers change the initial experience. Inspect question-level abandonment before redesigning all nine.

Keep the card requirement and £9.99 charge after seven days clear. Test whether showing a useful plan preview before the paywall improves informed trial starts.

How we decide

Validate funnel definitions and baseline volumes, then test these changes in priority order. Use week-one repeat workouts as an early signal, followed by trial-to-paid conversion and day-60 retention. Track cancellations, refunds and workout-related complaints. Ship changes that improve sustained participation and subscription outcomes, not just trial starts.

What a PM had to fix

  1. 1Contradiction missedSurface the contradictionTargeted repair

    Source: FunnelInstall → finished onboarding 71% → started trial 52%

    What we’d changeAccount for the onboarding-to-trial step: 71% finish onboarding but 52% start the trial, so about a quarter stop at the card-up-front paywall. That bears directly on the team's paywall plan.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent the reasons as hypotheses to test, as the memo already does for the 4x figure.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly33%
  • passAddresses the actual decision98%
  • passRespects explicit constraints49%
  • passIdentifies material uncertainty60%
  • passAvoids unsupported claims43%
  • passProduces the required deliverable91%
  • passPrioritises by impact93%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right63%
Run
Run
#1
Time to output
39 s
Submitted
25 Sept 2026

Usable with a quick edit9 of 9 checks passed

The memo is clear, evidence-based, and actionable with only trivial copy edits. No re-checking of evidence or redoing of reasoning is needed.

Output

To: Growth PM Subject: Prioritise first-week activation over a full onboarding redesign

Stride asks users to commit before delivering value, then makes consistency difficult: the first workout is delayed, every session requires 30–45 minutes, and missed sessions have no recovery path. I would redirect most of this quarter’s effort from questionnaire/paywall redesign to helping users complete—and repeat—a workout.

Three week-one workouts is a promising activation signal, but the fourfold day-60 retention association is not proof that driving three completions will cause higher retention; motivated users may do both.

Recommended changes, in priority order

1. Make the first workout immediate and achievable. After trial activation, offer “Start now” with a 10–15-minute beginner-friendly session, alongside the option to schedule it. Keep longer sessions available, but do not make 30–45 minutes the only entry point. Preserve injury and equipment screening before recommending exercise.

2. Build a forgiving path to three sessions. Offer shorter alternatives throughout the week. After a missed workout, let users reschedule, repeat, or choose a lighter session rather than automatically advancing. Show progress toward a flexible weekly goal from day zero; do not wait seven days to introduce feedback. Avoid punitive streak mechanics that turn one missed day into perceived failure.

3. Make reminders support the user’s commitment. Replace the fixed day-one 7am push with the user’s selected time. Ask users to confirm their first session’s timing and send a specific prompt, such as “Your 12-minute starter workout is ready.” Adapt reminders after missed sessions and provide notification controls.

4. Streamline onboarding and clarify the trial—without a wholesale redesign. Keep questions necessary for safe, relevant recommendations; defer other questions until their answers improve the experience. Clearly state the £9.99 monthly charge, billing date, and cancellation process. Send a pre-charge reminder. Test focused changes before committing the quarter to visual redesign.

Measurement and rollout

First confirm whether funnel percentages are install-based or step-to-step, and define workout completion consistently. Instrument time to first workout, workout duration, missed-session recovery, and reasons for abandonment.

Run staged randomised tests, starting with immediate/short workouts, then recovery mechanics. Use first-workout completion and three-session activation as leading indicators; judge success by day-60 paid retention and revenue per assigned user—not activated users alone. Monitor trial starts, cancellations, refunds, notification opt-outs, and safety complaints.

Decision: prioritise activation and repeatability, while making low-cost onboarding and billing improvements now. Revisit a full redesign once evidence identifies which pre-trial friction materially limits retained-user growth.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly29%
  • passAddresses the actual decision99%
  • passRespects explicit constraints24%
  • passIdentifies material uncertainty82%
  • passAvoids unsupported claims61%
  • passProduces the required deliverable87%
  • passPrioritises by impact97%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right74%
Run
Run
#1
API response time
16 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Uses the supplied evidence correctlyWrongRightRight
Gemini 3.8 Flash · API

It states unsupported current-situation claims such as high initial friction, missed days, guilt, churn, and a psychological hook as facts rather than labelled hypotheses.

GPT-6 Astra · ChatGPT

All statements about the current situation are taken directly from the brief or derived arithmetically with a clearly stated assumption.

GPT-6.1 Sol · API

All current-state facts it relies on are taken correctly from the brief, with no invented figures or system details.

Identifies material uncertaintyWrongRightRight
Gemini 3.8 Flash · API

It does not name material unknowns or conditions that would change the call, such as whether shorter workouts preserve habit formation or whether the 4x relationship is causal.

GPT-6 Astra · ChatGPT

Acknowledges the activation hypothesis may not be causal, flags need to validate funnel definitions and baseline volumes, and outlines metrics to inform the decision.

GPT-6.1 Sol · API

It names causal uncertainty in the retention signal, asks to confirm funnel bases and workout-completion definitions, and states what would change the call.

Avoids unsupported claimsWrongRightRight
Gemini 3.8 Flash · API

It presents several causes and psychological effects as established facts when the supplied evidence only supports correlations or funnel drop-offs.

GPT-6 Astra · ChatGPT

Labels causation hypothesis and interpretations as such, and does not present unsupported claims as fact.

GPT-6.1 Sol · API

It labels the retention link as an association rather than proof and does not present untested causes or forecasts as established fact.

All got right 6

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

It clearly commits to prioritising week-one activation over onboarding/paywall redesign and gives a ranked set of changes for the growth PM.

GPT-6 Astra · ChatGPT

Commits to redirecting quarter's work and ranks four changes in priority order, with a measurement plan to validate.

GPT-6.1 Sol · API

It commits early to prioritising activation over a full onboarding redesign and says a redesign would be revisited only if evidence identifies meaningful pre-trial friction.

Respects explicit constraintsRightRightRight
Gemini 3.8 Flash · API

It is a memo to the growth PM, under 500 words, and recommends changes in priority order.

GPT-6 Astra · ChatGPT

Delivers a memo to the Growth PM under 500 words, respecting form and length.

GPT-6.1 Sol · API

It is a memo to the growth PM, stays under the 500-word limit, and its proposals would enforce the requested focus in practice.

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

The requested memo is present, complete, addressed to the growth PM, and actionable within the length limit.

GPT-6 Astra · ChatGPT

Provides a complete, actionable memo that a growth PM could implement without major gaps.

GPT-6.1 Sol · API

The memo is complete, correctly addressed, within the required length, and actionable with at most light edits.

Prioritises by impactRightRightRight
Gemini 3.8 Flash · API

The fixes are ranked in a plausible impact order centred on reaching three workouts in week one.

GPT-6 Astra · ChatGPT

Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.

GPT-6.1 Sol · API

Changes are ranked 1–4 by likely impact on activation, starting with making the first workout immediate and building a forgiving path to three sessions.

Activation defined by what predicts retentionRightRightRight
Gemini 3.8 Flash · API

It defines activation as three workouts in week one, links that to day-60 retention, and ranks fixes by their effect on that activation event.

GPT-6 Astra · ChatGPT

Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.

GPT-6.1 Sol · API

It defines activation as three first-week workouts, uses the supplied day-60 retention link, and ranks fixes by their effect on that behaviour.

Gets the base of every number rightRightRightRight
Gemini 3.8 Flash · API

The percentages and 4x relationship are used against the correct funnel steps and denominators.

GPT-6 Astra · ChatGPT

Derived percentages correctly use trial starters (48/52) and first-workout users (19/48), with the base stated via the assumption.

GPT-6.1 Sol · API

The output does not compute step-to-step percentages, and the only cited statistic, the fourfold retention link, matches the supplied funnel note.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 82% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT94.4100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude80.665.02None
5Gemini 3.5 Flash-LitewithGemini69.455.02None
6Sonnet 5.5withAPI77.845.02None
7Gemini 3.8 FlashwithAPI69.450.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review