Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable93% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims54% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly54% pass
    It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
    Sonnet 5.5 · API · Fitness app first week
  3. Identifies material uncertainty59% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review the first-week experience of our fitness app, described below, and recommend the changes you would make, in priority order, as a memo to the growth PM. Keep it under 500 words.

What the model was given4 items: Scenario, First week, step by step, Funnel, Team plan
ScenarioStride is a home-workout app. It costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.
First week, step by stepDay 0: nine onboarding questions (goal, age, height, weight, fitness level, equipment, injuries, preferred days, preferred time), then the paywall to start the trial, then a plan with the first workout scheduled for the next morning. Every workout is 30–45 minutes. Day 1: a 7am push notification, “Your workout is ready.” Days 2–7: one push a day at the preferred time. Missing a workout moves the plan on to the next one; there is no shorter or catch-up option. The streak counter first appears after seven days.
FunnelInstall → finished onboarding 71% → started trial 52% → first workout 48% → three workouts in week one 19%. Users who complete three workouts in week one are four times as likely to still be subscribed at day 60.
Team planThe design team wants to spend this quarter redesigning the onboarding questionnaire and the paywall screen.
What a strong answer doesThe answer key the graders mark against

Focus on getting new users to three workouts in week one (shorter first workouts, a same-day first session, catch-up options, earlier progress feedback) rather than onboarding and paywall polish.

Case

Warm-up · v1.8 · synthetic · consumer, mobile

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit5 of 9 checks passed

Unsupported claim about user sentiment and missing conditions for the recommendation mean the audience would need to re-check evidence and add decision criteria before using it.

Output

MEMORANDUM

To: Growth Product Manager From: Growth & Product Strategy Date: September 25, 2026 Subject: First-Week Experience Review & Recommendations

---

Executive Summary

Our funnel reveals a critical bottleneck: while 52% start the trial, only 19% complete three workouts in week one1—our primary retention driver (users hitting this milestone are 4X more likely to subscribe at day 60). Furthermore, the design team's current focus on redesigning onboarding and the paywall is misaligned; onboarding completion is already healthy at 71%2, and trial conversion sits at 52%. Instead, optimization efforts must shift to activation, early habit formation, and flexibility.

Below are our recommended changes, ordered by priority.

---

Priority Recommendations

#### 1. Introduce Flexible "Catch-Up" & Micro-Workouts (High Impact) * The Problem: Currently, missing a scheduled 30–45 minute workout simply shuffles the plan forward without offering alternative options, leaving users overwhelmed and leading to immediate drop-off4 before reaching the vital three-workout milestone. * The Fix: Implement a "Catch-Up" mode offering shorter 10–15 minute micro-workouts or scaled options when a user misses a session. This lowers the barrier to entry and protects the Day 7 retention streak.

#### 2. Shift the Streak Counter and Early Milestone Rewards to Day 1 (High Impact) * The Problem: The streak counter currently remains hidden until Day 7. Users miss out on crucial early psychological reinforcement during their most vulnerable window. * The Fix: Surface the streak counter immediately on Day 1 upon completing the first workout. Introduce immediate celebratory micro-milestones (e.g., badges or progress indicators) after the very first completed session to build early momentum.

#### 3. Personalize and Optimize Push Notification Triggers (Medium Impact) * The Problem: Rigid 7 AM notifications or static reminders often feel intrusive or misaligned with user behavior3, contributing to friction during the trial week. * The Fix: Dynamically adapt push notifications based on user behavior. If a user misses their preferred time, send a gentle, encouraging nudge offering a shorter alternative workout rather than a rigid command.

#### 4. Deprioritize Onboarding and Paywall Redesigns (Strategic Pivot) * The Problem: The design team’s current plan to spend the quarter redesigning a 71% completion onboarding flow and a 52% trial-conversion paywall addresses healthy metrics. * The Fix: Reallocate design and engineering resources away from onboarding/paywalls and toward activation features (such as the micro-workout player, catch-up flows, and early streak tracking).

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticStart again

    Source: Funnelstarted trial 52% → first workout 48% → three workouts in week one 19%

    What we’d changeRead each step against the one before: 92% of trial starters do a first workout, and 40% of those reach three. That second drop is the leak. Add a test plan and success metrics.

  2. 2Invented evidenceVerify or remove the claimSubstantial rework

    What we’d changeRemove 'healthy': nothing in the brief benchmarks it. And note that a quarter of finishers still drop at the paywall.

  3. 3Invented evidenceVerify or remove the claimQuick edit

    Source: First week, step by stepDays 2–7: one push a day at the preferred time.

    What we’d changeOnly the day-1 push is at 7am. The rest already go at the user's preferred time. The fix is the day-1 push, not all reminders.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent it as a likely reason to test. The brief shows the drop-off, not how users felt.

Claims the judge couldn’t find in the brief

  • Rigid 7 AM notifications or static reminders often feel intrusive or misaligned with user behavior, contributing to friction during the trial week

Grades and run details

Decision model 67 · LLM judge 5 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly47%
  • passAddresses the actual decision93%
  • passRespects explicit constraints45%
  • failIdentifies material uncertainty99%
  • partialAvoids unsupported claims31%
  • passProduces the required deliverable81%
  • partialPrioritises by impact38%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right96%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Needs more than a quick edit5 of 9 checks passed

The unsupported claim about refund/churn means a PM would need to verify and potentially correct a factual assumption before acting on the memo.

Output

Stride: fixing the first week

Sep 25, 2026 · @Martin Slaney

Summary

Spend this quarter on the first seven days, not the questionnaire and paywall. The biggest leak is between first workout (48%) and three workouts (19%), and three workouts is the behaviour tied to day-60 retention. Redesigning the top of the funnel pours more trial users into a week that currently loses most of them.

What the funnel says

StepShare of installsKept from previous step
Finished onboarding71%71%
Started trial52%73%
First workout48%92%
Three workouts in week one19%40%

Trial-to-first-workout is healthy. Six in ten people who work out once never reach three. Because the card is taken up front, those users are charged on day 7 and then churn or ask for refunds.1

The week-one design explains the drop:

  • Too long. Every session is 30–45 minutes, for everyone, from day one.
  • Missing one punishes you. The plan moves on; there is no shorter or catch-up option.
  • Day 0 intent is wasted. The first workout is the next morning, not now.
  • Day 1 push ignores the user. It fires at 7am even though we asked their preferred time.
  • Progress is hidden. The streak appears only after seven days, so it rewards people who have already succeeded.

Changes, in priority order

  1. Short sessions in week one. Default to 10–15 minute workouts for days 1–7, with the full session as an opt-in. Lowest effort, biggest lever on reaching three.
  2. Recover missed days instead of skipping. A missed workout offers a 10-minute version the same evening or next day, not the next plan step.
  3. First workout on day 0. Straight after the plan screen, offer a 5–10 minute starter. It counts toward the three.
  4. Fix notifications. Honour the preferred time from day 1. Replace "Your workout is ready" with the specific session and length ("12 min, no equipment"). Add one recovery nudge after a miss.
  5. Show progress to three from day 1. A "1 of 3 this week" tracker replaces the hidden streak and makes the target explicit.
  6. Trim onboarding, lightly. Keep only questions that change the first plan; move height, weight and preferred days to after the first workout. A sprint, not a quarter.
  7. Leave the paywall. 73% of finishers start the trial; it is not the constraint.2

Measurement and the ask

The 4× day-60 figure is a correlation; motivated users may simply do more workouts. Ship changes 1–5 as A/B tests so we learn whether reaching three causes retention.3

  • Primary metric: share of trial starters completing three workouts in days 0–7 (today roughly 37%).
  • Guardrails: trial-to-paid conversion, day-7 refund requests, day-60 retention.

Ask: redirect the design team's quarter from questionnaire and paywall to the week-one experience, with a light onboarding trim as a side task.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    Source: ScenarioIt costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.

    What we’d changeRemove it: the brief says nothing about what these users do at day 7.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeA 73% step rate doesn't show the paywall isn't a constraint: a quarter of people who finish onboarding stop there.

  3. 3Test or gate too weakTighten the testQuick edit

    What we’d changeA/B tests show whether these changes improve retention. They won't show that reaching three workouts is itself the cause.

Claims the judge couldn’t find in the brief

  • Because the card is taken up front, those users are charged on day 7 and then churn or ask for refunds.

Grades and run details

Decision model 67 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly39%
  • passAddresses the actual decision99%
  • partialRespects explicit constraints13%
  • passIdentifies material uncertainty72%
  • partialAvoids unsupported claims51%
  • passProduces the required deliverable92%
  • passPrioritises by impact98%
  • passActivation defined by what predicts retention100%
  • failGets the base of every number right19%
Run
Run
#1
Time to output
49 s
Submitted
25 Sept 2026

Usable with a quick edit9 of 9 checks passed

The memo is clear, evidence-based, and actionable with only trivial copy edits. No re-checking of evidence or redoing of reasoning is needed.

Output

To: Growth PM Subject: Prioritise first-week activation over a full onboarding redesign

Stride asks users to commit before delivering value, then makes consistency difficult: the first workout is delayed, every session requires 30–45 minutes, and missed sessions have no recovery path. I would redirect most of this quarter’s effort from questionnaire/paywall redesign to helping users complete—and repeat—a workout.

Three week-one workouts is a promising activation signal, but the fourfold day-60 retention association is not proof that driving three completions will cause higher retention; motivated users may do both.

Recommended changes, in priority order

1. Make the first workout immediate and achievable. After trial activation, offer “Start now” with a 10–15-minute beginner-friendly session, alongside the option to schedule it. Keep longer sessions available, but do not make 30–45 minutes the only entry point. Preserve injury and equipment screening before recommending exercise.

2. Build a forgiving path to three sessions. Offer shorter alternatives throughout the week. After a missed workout, let users reschedule, repeat, or choose a lighter session rather than automatically advancing. Show progress toward a flexible weekly goal from day zero; do not wait seven days to introduce feedback. Avoid punitive streak mechanics that turn one missed day into perceived failure.

3. Make reminders support the user’s commitment. Replace the fixed day-one 7am push with the user’s selected time. Ask users to confirm their first session’s timing and send a specific prompt, such as “Your 12-minute starter workout is ready.” Adapt reminders after missed sessions and provide notification controls.

4. Streamline onboarding and clarify the trial—without a wholesale redesign. Keep questions necessary for safe, relevant recommendations; defer other questions until their answers improve the experience. Clearly state the £9.99 monthly charge, billing date, and cancellation process. Send a pre-charge reminder. Test focused changes before committing the quarter to visual redesign.

Measurement and rollout

First confirm whether funnel percentages are install-based or step-to-step, and define workout completion consistently. Instrument time to first workout, workout duration, missed-session recovery, and reasons for abandonment.

Run staged randomised tests, starting with immediate/short workouts, then recovery mechanics. Use first-workout completion and three-session activation as leading indicators; judge success by day-60 paid retention and revenue per assigned user—not activated users alone. Monitor trial starts, cancellations, refunds, notification opt-outs, and safety complaints.

Decision: prioritise activation and repeatability, while making low-cost onboarding and billing improvements now. Revisit a full redesign once evidence identifies which pre-trial friction materially limits retained-user growth.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly29%
  • passAddresses the actual decision99%
  • passRespects explicit constraints24%
  • passIdentifies material uncertainty82%
  • passAvoids unsupported claims61%
  • passProduces the required deliverable87%
  • passPrioritises by impact97%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right74%
Run
Run
#1
API response time
16 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyWrongWrongRight
Gemini 3.5 Flash-Lite · Gemini

States that notifications often feel intrusive or misaligned, which is not in the supplied context.

Opus 5.5 · Claude

The output claims that users charged on day 7 “churn or ask for refunds”, which is not in the supplied context and is not supported by the data.

GPT-6.1 Sol · API

All current-state facts it relies on are taken correctly from the brief, with no invented figures or system details.

Addresses the actual decisionMixedMixedRight
Gemini 3.5 Flash-Lite · Gemini

Does not indicate what result or condition would change the recommended priority or actions.

Opus 5.5 · Claude

The output commits to redirecting the design team but does not explicitly state what result or condition would change that decision, even though the brief requires it.

GPT-6.1 Sol · API

It commits early to prioritising activation over a full onboarding redesign and says a redesign would be revisited only if evidence identifies meaningful pre-trial friction.

Identifies material uncertaintyWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

Does not name any unknowns that could change the decision or say how they would be resolved.

Opus 5.5 · Claude

It names the key unknown (correlation vs causation for the 4× day‑60 figure) and proposes an A/B test to resolve it.

GPT-6.1 Sol · API

It names causal uncertainty in the retention signal, asks to confirm funnel bases and workout-completion definitions, and states what would change the call.

Avoids unsupported claimsWrongWrongRight
Gemini 3.5 Flash-Lite · Gemini

Presents the claim that notifications feel intrusive as established fact without labelling it as a hypothesis.

Opus 5.5 · Claude

The statement that users “churn or ask for refunds” is presented as fact without labelling it as an assumption, and the supplied evidence does not support it.

GPT-6.1 Sol · API

It labels the retention link as an association rather than proof and does not present untested causes or forecasts as established fact.

Gets the base of every number rightRightMixedRight
Gemini 3.5 Flash-Lite · Gemini

All numbers are directly from the supplied funnel; no derived figures require base checks.

Opus 5.5 · Claude

All derived percentages (73 % trial start from onboarding, 60 % drop, 37 % three workouts among trial starters) use the correct denominators from the supplied funnel.

GPT-6.1 Sol · API

The output does not compute step-to-step percentages, and the only cited statistic, the fourfold retention link, matches the supplied funnel note.

All got right 4

Respects explicit constraintsRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Delivered as a memo to the growth PM under 500 words.

Opus 5.5 · Claude

The output is a memo to the growth PM, under 500 words, with changes in priority order, respecting all given constraints.

GPT-6.1 Sol · API

It is a memo to the growth PM, stays under the 500-word limit, and its proposals would enforce the requested focus in practice.

Produces the required deliverableRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Memo is present, complete, within length, and gives actionable prioritized recommendations.

Opus 5.5 · Claude

The request deliverable (a priority‑ordered memo to the growth PM, under 500 words) is present and immediately usable.

GPT-6.1 Sol · API

The memo is complete, correctly addressed, within the required length, and actionable with at most light edits.

Prioritises by impactRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Fixes are ranked by priority with impact levels (high/medium), targeting activation.

Opus 5.5 · Claude

Fixes are ranked by likely impact on reaching three workouts, with low‑effort, high‑lever actions first.

GPT-6.1 Sol · API

Changes are ranked 1–4 by likely impact on activation, starting with making the first workout immediate and building a forgiving path to three sessions.

Activation defined by what predicts retentionRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Activation is defined as three workouts, linked to retention (4x), and ranking is based on effect on this event.

Opus 5.5 · Claude

The activation event (three workouts in week one) is defined by its link to day‑60 retention, and fixes are judged by their effect on that event.

GPT-6.1 Sol · API

It defines activation as three first-week workouts, uses the supplied day-60 retention link, and ranks fixes by their effect on that behaviour.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 82% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT94.4100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude80.665.02None
5Gemini 3.5 Flash-LitewithGemini69.455.02None
6Sonnet 5.5withAPI77.845.02None
7Gemini 3.8 FlashwithAPI69.450.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review