Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 58% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable92% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims60% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly63% pass
    The output claims that users charged on day 7 “churn or ask for refunds”, which is not in the supplied context and is not supported by the data.
    Opus 5.5 · Claude · Fitness app first week
  3. Identifies material uncertainty69% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review the first-week experience of our fitness app, described below, and recommend the changes you would make, in priority order, as a memo to the growth PM. Keep it under 500 words.

ScenarioStride is a home-workout app. It costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.
First week, step by stepDay 0: nine onboarding questions (goal, age, height, weight, fitness level, equipment, injuries, preferred days, preferred time), then the paywall to start the trial, then a plan with the first workout scheduled for the next morning. Every workout is 30–45 minutes. Day 1: a 7am push notification, “Your workout is ready.” Days 2–7: one push a day at the preferred time. Missing a workout moves the plan on to the next one; there is no shorter or catch-up option. The streak counter first appears after seven days.
FunnelInstall → finished onboarding 71% → started trial 52% → first workout 48% → three workouts in week one 19%. Users who complete three workouts in week one are four times as likely to still be subscribed at day 60.
Team planThe design team wants to spend this quarter redesigning the onboarding questionnaire and the paywall screen.
What a strong answer does

Focus on getting new users to three workouts in week one (shorter first workouts, a same-day first session, catch-up options, earlier progress feedback) rather than onboarding and paywall polish.

Case

v1.6 · synthetic · consumer, mobile

Usable with a quick edit9 of 9 checks passed

The memo is clear, actionable, and prioritised; only minor stylistic edits would be needed.

Output

To: Growth PM Subject: Prioritise repeat workouts over an onboarding redesign

We should redirect the quarter’s work towards helping people fit workouts into their lives and return after missing one. The questionnaire and paywall deserve attention, but a full redesign is premature.

Assuming the funnel percentages share an install denominator, 92% of trial starters reach a first workout (48/52), but only 40% of first-workout users reach three sessions (19/48). We also lose 29% before onboarding finishes.1 Both matter, but the current plan leaves the substantial repeat-workout drop untouched.

Three workouts is a useful activation hypothesis. The fourfold association with day-60 subscription does not prove causation: more motivated users may both exercise more and stay subscribed.

1. Make repeat workouts achievable, including after a missed day.

Offer a 10–15-minute option alongside full sessions, with suitable adjustments for fitness level, equipment and injuries. After a missed workout, let users reschedule or choose a shorter session rather than automatically advancing the plan.

Requiring 30–45 minutes every time creates a substantial commitment before a routine exists.2 Recovery should be part of the experience, not something users must work out themselves. Test whether these options increase repeat participation and total active minutes, rather than simply inflating completion counts.

2. Offer a first workout immediately and respect the user’s schedule.

After trial activation, offer “Start a short workout now” or let users choose their first session time. Replace the fixed day-one 7am push with their chosen time.

Scheduling everyone for tomorrow wastes initial motivation and consumes part of a seven-day trial before value is experienced. This ranks below repeat-workout flexibility because first-workout conversion among trial starters is already strong.

3. Show progress during the week and make reminders actionable.

Show the first completed session immediately and progress towards a user-chosen weekly goal. Avoid making consecutive-day streaks the default for people who selected fewer workout days.

Use reminders that offer a relevant next step: start, shorten or reschedule. Stop treating every day as another generic prompt. Measure notification opt-outs alongside workout returns.

4. Simplify onboarding through focused experiments.

Keep information needed for safe, relevant workouts upfront, particularly injuries, fitness level and equipment. Test deferring other questions unless their answers change the initial experience. Inspect question-level abandonment before redesigning all nine.

Keep the card requirement and £9.99 charge after seven days clear. Test whether showing a useful plan preview before the paywall improves informed trial starts.

How we decide

Validate funnel definitions and baseline volumes, then test these changes in priority order. Use week-one repeat workouts as an early signal, followed by trial-to-paid conversion and day-60 retention. Track cancellations, refunds and workout-related complaints. Ship changes that improve sustained participation and subscription outcomes, not just trial starts.

What a PM had to fix

  1. 1Contradiction missedSurface the contradictionTargeted repair

    Source: FunnelInstall → finished onboarding 71% → started trial 52%

    What we’d changeAccount for the onboarding-to-trial step: 71% finish onboarding but 52% start the trial, so about a quarter stop at the card-up-front paywall. That bears directly on the team's paywall plan.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent the reasons as hypotheses to test, as the memo already does for the 4x figure.

Check by check

Got right · 9
  • Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the brief or derived arithmetically with a clearly stated assumption.
  • Addresses the actual decisionCommits to redirecting quarter's work and ranks four changes in priority order, with a measurement plan to validate.
  • Respects explicit constraintsDelivers a memo to the Growth PM under 500 words, respecting form and length.
  • Identifies material uncertaintyAcknowledges the activation hypothesis may not be causal, flags need to validate funnel definitions and baseline volumes, and outlines metrics to inform the decision.
  • Avoids unsupported claimsLabels causation hypothesis and interpretations as such, and does not present unsupported claims as fact.
  • Produces the required deliverableProvides a complete, actionable memo that a growth PM could implement without major gaps.
  • Prioritises by impactRanks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
  • Activation defined by what predicts retentionDefines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
  • Gets the base of every number rightDerived percentages correctly use trial starters (48/52) and first-workout users (19/48), with the base stated via the assumption.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly25%
  • passAddresses the actual decision91%
  • passRespects explicit constraints51%
  • passIdentifies material uncertainty63%
  • passAvoids unsupported claims43%
  • passProduces the required deliverable91%
  • passPrioritises by impact95%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right59%
Run
Run
#1
Time to output
39 s
Submitted
25 Sept 2026

Needs more than a quick edit5 of 9 checks passed

Unsupported claim about user sentiment and missing conditions for the recommendation mean the audience would need to re-check evidence and add decision criteria before using it.

Output

MEMORANDUM

To: Growth Product Manager From: Growth & Product Strategy Date: September 25, 2026 Subject: First-Week Experience Review & Recommendations

---

Executive Summary

Our funnel reveals a critical bottleneck: while 52% start the trial, only 19% complete three workouts in week one1—our primary retention driver (users hitting this milestone are 4X more likely to subscribe at day 60). Furthermore, the design team's current focus on redesigning onboarding and the paywall is misaligned; onboarding completion is already healthy at 71%2, and trial conversion sits at 52%. Instead, optimization efforts must shift to activation, early habit formation, and flexibility.

Below are our recommended changes, ordered by priority.

---

Priority Recommendations

#### 1. Introduce Flexible "Catch-Up" & Micro-Workouts (High Impact) * The Problem: Currently, missing a scheduled 30–45 minute workout simply shuffles the plan forward without offering alternative options, leaving users overwhelmed and leading to immediate drop-off4 before reaching the vital three-workout milestone. * The Fix: Implement a "Catch-Up" mode offering shorter 10–15 minute micro-workouts or scaled options when a user misses a session. This lowers the barrier to entry and protects the Day 7 retention streak.

#### 2. Shift the Streak Counter and Early Milestone Rewards to Day 1 (High Impact) * The Problem: The streak counter currently remains hidden until Day 7. Users miss out on crucial early psychological reinforcement during their most vulnerable window. * The Fix: Surface the streak counter immediately on Day 1 upon completing the first workout. Introduce immediate celebratory micro-milestones (e.g., badges or progress indicators) after the very first completed session to build early momentum.

#### 3. Personalize and Optimize Push Notification Triggers (Medium Impact) * The Problem: Rigid 7 AM notifications or static reminders often feel intrusive or misaligned with user behavior3, contributing to friction during the trial week. * The Fix: Dynamically adapt push notifications based on user behavior. If a user misses their preferred time, send a gentle, encouraging nudge offering a shorter alternative workout rather than a rigid command.

#### 4. Deprioritize Onboarding and Paywall Redesigns (Strategic Pivot) * The Problem: The design team’s current plan to spend the quarter redesigning a 71% completion onboarding flow and a 52% trial-conversion paywall addresses healthy metrics. * The Fix: Reallocate design and engineering resources away from onboarding/paywalls and toward activation features (such as the micro-workout player, catch-up flows, and early streak tracking).

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticStart again

    Source: Funnelstarted trial 52% → first workout 48% → three workouts in week one 19%

    What we’d changeRead each step against the one before: 92% of trial starters do a first workout, and 40% of those reach three. That second drop is the leak. Add a test plan and success metrics.

  2. 2Invented evidenceVerify or remove the claimSubstantial rework

    What we’d changeRemove 'healthy': nothing in the brief benchmarks it. And note that a quarter of finishers still drop at the paywall.

  3. 3Invented evidenceVerify or remove the claimQuick edit

    Source: First week, step by stepDays 2–7: one push a day at the preferred time.

    What we’d changeOnly the day-1 push is at 7am. The rest already go at the user's preferred time. The fix is the day-1 push, not all reminders.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent it as a likely reason to test. The brief shows the drop-off, not how users felt.

Check by check

Got wrong · 4
  • Uses the supplied evidence correctlyStates that notifications often feel intrusive or misaligned, which is not in the supplied context.
  • Addresses the actual decisionDoes not indicate what result or condition would change the recommended priority or actions.
  • Identifies material uncertaintyDoes not name any unknowns that could change the decision or say how they would be resolved.
  • Avoids unsupported claimsPresents the claim that notifications feel intrusive as established fact without labelling it as a hypothesis.
Got right · 5
  • Respects explicit constraintsDelivered as a memo to the growth PM under 500 words.
  • Produces the required deliverableMemo is present, complete, within length, and gives actionable prioritized recommendations.
  • Prioritises by impactFixes are ranked by priority with impact levels (high/medium), targeting activation.
  • Activation defined by what predicts retentionActivation is defined as three workouts, linked to retention (4x), and ranking is based on effect on this event.
  • Gets the base of every number rightAll numbers are directly from the supplied funnel; no derived figures require base checks.

Claims the judge couldn’t find in the brief

  • Rigid 7 AM notifications or static reminders often feel intrusive or misaligned with user behavior, contributing to friction during the trial week

Grades and run details

Decision model 61 · LLM judge 5 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly45%
  • partialAddresses the actual decision10%
  • passRespects explicit constraints40%
  • failIdentifies material uncertainty99%
  • partialAvoids unsupported claims33%
  • passProduces the required deliverable83%
  • partialPrioritises by impact27%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right97%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT97.2100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude86.165.02None
5Sonnet 5.5withAPI77.845.02None
6Gemini 3.5 Flash-LitewithGemini61.155.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review