Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable93% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims54% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly54% pass
    It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
    Sonnet 5.5 · API · Fitness app first week
  3. Identifies material uncertainty59% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review the first-week experience of our fitness app, described below, and recommend the changes you would make, in priority order, as a memo to the growth PM. Keep it under 500 words.

What the model was given4 items: Scenario, First week, step by step, Funnel, Team plan
ScenarioStride is a home-workout app. It costs £9.99 a month after a seven-day free trial, and the trial needs a card up front.
First week, step by stepDay 0: nine onboarding questions (goal, age, height, weight, fitness level, equipment, injuries, preferred days, preferred time), then the paywall to start the trial, then a plan with the first workout scheduled for the next morning. Every workout is 30–45 minutes. Day 1: a 7am push notification, “Your workout is ready.” Days 2–7: one push a day at the preferred time. Missing a workout moves the plan on to the next one; there is no shorter or catch-up option. The streak counter first appears after seven days.
FunnelInstall → finished onboarding 71% → started trial 52% → first workout 48% → three workouts in week one 19%. Users who complete three workouts in week one are four times as likely to still be subscribed at day 60.
Team planThe design team wants to spend this quarter redesigning the onboarding questionnaire and the paywall screen.
What a strong answer doesThe answer key the graders mark against

Focus on getting new users to three workouts in week one (shorter first workouts, a same-day first session, catch-up options, earlier progress feedback) rather than onboarding and paywall polish.

Case

Warm-up · v1.8 · synthetic · consumer, mobile

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit6 of 9 checks passed

A PM would need to re-check and soften unsupported causal claims before acting. The core recommendation is usable, but the evidence framing is not safe as written.

Output

MEMO

TO: Growth PM FROM: Product & Growth Review SUBJECT: Week 1 Retention & Experience Optimization

Executive Summary

Push back on the design team’s plan to redesign onboarding and the paywall. Converting 71% through onboarding and 52% into a card-upfront trial is exceptionally strong.

Our growth bottleneck is Week 1 activation: while 48% log a first workout, only 19% reach three workouts—the critical threshold that drives 4x Day-60 retention. We are losing trialists during their first week because of delayed gratification, rigid scheduling, and high initial friction.

Here are the four prioritized changes to drive 3-workout completion.

---

Recommended Changes (Priority Order)

#### 1. Instant Gratification on Day 0 (“Start Now” Option) * Problem: Scheduling the first session for the next morning wastes peak user motivation right after purchase. * Fix: Offer an immediate, optional 5–10 minute "Quick Start / Assessment" on Day 0. Completing this counts as Workout #1, immediately advancing users toward the 3-workout milestone.

#### 2. Introduce "Micro-Workouts" & Flexible Session Lengths * Problem: A rigid 30–45 minute commitment is daunting for home workouts, leading to missed days. When a day is missed, the plan simply skips ahead without recourse, inducing guilt and churn. * Fix: Introduce 10- and 20-minute “time-crunched” alternatives when a user opens a session or is at risk of missing one. A 10-minute workout must count toward the 3-workout goal. Never skip a workout automatically; let users swap, reschedule, or condense.

#### 3. Surface Streaks & Milestones on Day 1 (Not Day 7) * Problem: Hiding the streak counter until Day 7 removes the most effective psychological hook during the make-or-break trial period. * Fix: * Show the streak counter immediately upon completing Workout #1. * Explicitly gamify the activation metric: show a progress bar toward the "Week 1 Foundation Badge" (1/3, 2/3, 3/3 workouts completed).

#### 4. Action-Oriented, Personalized Push Notifications * Problem: A generic 7:00 AM push ("Your workout is ready") is transactional and easily dismissed. * Fix: Leverage the rich onboarding data (goals, fitness level) to trigger personalized, high-context pushes: * Copy: “Crush your [Goal] today: 15 mins to hit 2/3 of your Week 1 target.” * Timing: Align strictly to their chosen time, with an evening "last chance" nudge offering an express 10-minute session if uncompleted.

---

Next Steps

Pause the onboarding/paywall redesign. Run a two-week sprint to implement immediate Day 0 activation, 15-minute workout variants, and Day 1 streak mechanics. Our target metric is moving Week 1 (3+ workouts) from 19% to 30%.

Claims the judge couldn’t find in the brief

  • The first week has high initial friction.

Grades and run details

Decision model 67 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly40%
  • passAddresses the actual decision98%
  • passRespects explicit constraints54%
  • failIdentifies material uncertainty88%
  • failAvoids unsupported claims33%
  • passProduces the required deliverable98%
  • passPrioritises by impact91%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right93%
Run
Run
#1
API response time
15 s
Submitted
1 Oct 2026

Needs more than a quick edit5 of 9 checks passed

Unsupported claim about user sentiment and missing conditions for the recommendation mean the audience would need to re-check evidence and add decision criteria before using it.

Output

MEMORANDUM

To: Growth Product Manager From: Growth & Product Strategy Date: September 25, 2026 Subject: First-Week Experience Review & Recommendations

---

Executive Summary

Our funnel reveals a critical bottleneck: while 52% start the trial, only 19% complete three workouts in week one1—our primary retention driver (users hitting this milestone are 4X more likely to subscribe at day 60). Furthermore, the design team's current focus on redesigning onboarding and the paywall is misaligned; onboarding completion is already healthy at 71%2, and trial conversion sits at 52%. Instead, optimization efforts must shift to activation, early habit formation, and flexibility.

Below are our recommended changes, ordered by priority.

---

Priority Recommendations

#### 1. Introduce Flexible "Catch-Up" & Micro-Workouts (High Impact) * The Problem: Currently, missing a scheduled 30–45 minute workout simply shuffles the plan forward without offering alternative options, leaving users overwhelmed and leading to immediate drop-off4 before reaching the vital three-workout milestone. * The Fix: Implement a "Catch-Up" mode offering shorter 10–15 minute micro-workouts or scaled options when a user misses a session. This lowers the barrier to entry and protects the Day 7 retention streak.

#### 2. Shift the Streak Counter and Early Milestone Rewards to Day 1 (High Impact) * The Problem: The streak counter currently remains hidden until Day 7. Users miss out on crucial early psychological reinforcement during their most vulnerable window. * The Fix: Surface the streak counter immediately on Day 1 upon completing the first workout. Introduce immediate celebratory micro-milestones (e.g., badges or progress indicators) after the very first completed session to build early momentum.

#### 3. Personalize and Optimize Push Notification Triggers (Medium Impact) * The Problem: Rigid 7 AM notifications or static reminders often feel intrusive or misaligned with user behavior3, contributing to friction during the trial week. * The Fix: Dynamically adapt push notifications based on user behavior. If a user misses their preferred time, send a gentle, encouraging nudge offering a shorter alternative workout rather than a rigid command.

#### 4. Deprioritize Onboarding and Paywall Redesigns (Strategic Pivot) * The Problem: The design team’s current plan to spend the quarter redesigning a 71% completion onboarding flow and a 52% trial-conversion paywall addresses healthy metrics. * The Fix: Reallocate design and engineering resources away from onboarding/paywalls and toward activation features (such as the micro-workout player, catch-up flows, and early streak tracking).

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticStart again

    Source: Funnelstarted trial 52% → first workout 48% → three workouts in week one 19%

    What we’d changeRead each step against the one before: 92% of trial starters do a first workout, and 40% of those reach three. That second drop is the leak. Add a test plan and success metrics.

  2. 2Invented evidenceVerify or remove the claimSubstantial rework

    What we’d changeRemove 'healthy': nothing in the brief benchmarks it. And note that a quarter of finishers still drop at the paywall.

  3. 3Invented evidenceVerify or remove the claimQuick edit

    Source: First week, step by stepDays 2–7: one push a day at the preferred time.

    What we’d changeOnly the day-1 push is at 7am. The rest already go at the user's preferred time. The fix is the day-1 push, not all reminders.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changePresent it as a likely reason to test. The brief shows the drop-off, not how users felt.

Claims the judge couldn’t find in the brief

  • Rigid 7 AM notifications or static reminders often feel intrusive or misaligned with user behavior, contributing to friction during the trial week

Grades and run details

Decision model 67 · LLM judge 5 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly47%
  • passAddresses the actual decision93%
  • passRespects explicit constraints45%
  • failIdentifies material uncertainty99%
  • partialAvoids unsupported claims31%
  • passProduces the required deliverable81%
  • partialPrioritises by impact38%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right96%
Run
Run
#1
Time to output
6 s
Submitted
25 Sept 2026

Usable with a quick edit9 of 9 checks passed

It makes the right activation-focused call, ranks concrete changes, and flags the key measurement caveat. A PM could act on it with only light wording edits.

Output

Memo

To: Growth PM Subject: Priorities for Stride’s first week

I would not make a full onboarding and paywall redesign the quarter’s main bet. The clearest retention signal is that people who complete three workouts in week one are much more likely to stay subscribed. Prioritize helping more new users reach that behavior, then address the earlier funnel friction.

  1. Make the first workout easy to start and finish. Users currently wait until the next morning and face a 30–45-minute session. Offer an optional workout immediately, plus a 10–15-minute beginner session. Let people choose when to start rather than defaulting to 7am; the Day 1 notification should respect their chosen time.
  1. Help users recover when they miss a session. A missed workout currently pushes the plan forward, with no shorter or catch-up option. Let users reschedule, swap in a quick workout, or resume without feeling behind. Build the week around three achievable sessions, not a rigid sequence.
  1. Show progress early. Don’t wait seven days to introduce the streak counter. Celebrate the first workout and show progress toward the second and third. Test milestone prompts that make the three-workout goal feel attainable without shaming users who miss a day.
  1. Reduce onboarding friction selectively. Nine questions—including weight and injury details—come before users experience the product. Test asking only what is needed to recommend a safe first workout, and collect the rest later or as optional. Preserve essential safety questions.
  1. Improve trial clarity and measurement. Keep the card-up-front model for now, but ensure the price, trial end date, and cancellation terms are unmistakable. Before committing to a paywall redesign, test specific changes. Also confirm whether funnel percentages are install-based or step-to-step, and instrument workout starts, completions, and reschedules.

Run focused experiments on these changes and judge them by first-workout completion and three workouts in week one—not just trial starts. The reported day-60 relationship is promising, but it is correlational; verify that lifting week-one completion improves retention.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly40%
  • passAddresses the actual decision98%
  • passRespects explicit constraints66%
  • passIdentifies material uncertainty69%
  • passAvoids unsupported claims67%
  • passProduces the required deliverable93%
  • passPrioritises by impact96%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right73%
Run
Run
#1
API response time
13 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyWrongWrongRight
Gemini 3.8 Flash · API

It states unsupported current-situation claims such as high initial friction, missed days, guilt, churn, and a psychological hook as facts rather than labelled hypotheses.

Gemini 3.5 Flash-Lite · Gemini

States that notifications often feel intrusive or misaligned, which is not in the supplied context.

GPT-6 Luna · API

All current-situation statements in the output are drawn correctly from the supplied context.

Addresses the actual decisionRightMixedRight
Gemini 3.8 Flash · API

It clearly commits to prioritising week-one activation over onboarding/paywall redesign and gives a ranked set of changes for the growth PM.

Gemini 3.5 Flash-Lite · Gemini

Does not indicate what result or condition would change the recommended priority or actions.

GPT-6 Luna · API

The memo commits early to not making onboarding/paywall redesign the main bet and states that lifting week-one completion would need to be verified as the condition that could change the call.

Identifies material uncertaintyWrongWrongRight
Gemini 3.8 Flash · API

It does not name material unknowns or conditions that would change the call, such as whether shorter workouts preserve habit formation or whether the 4x relationship is causal.

Gemini 3.5 Flash-Lite · Gemini

Does not name any unknowns that could change the decision or say how they would be resolved.

GPT-6 Luna · API

It names specific unknowns—funnel percentage basis and whether lifting week-one completion improves retention—and says how to resolve them.

Avoids unsupported claimsWrongWrongRight
Gemini 3.8 Flash · API

It presents several causes and psychological effects as established facts when the supplied evidence only supports correlations or funnel drop-offs.

Gemini 3.5 Flash-Lite · Gemini

Presents the claim that notifications feel intrusive as established fact without labelling it as a hypothesis.

GPT-6 Luna · API

The one causal caveat is explicitly labelled as correlational, and confident claims are supported by the supplied evidence.

All got right 5

Respects explicit constraintsRightRightRight
Gemini 3.8 Flash · API

It is a memo to the growth PM, under 500 words, and recommends changes in priority order.

Gemini 3.5 Flash-Lite · Gemini

Delivered as a memo to the growth PM under 500 words.

GPT-6 Luna · API

It is a memo to the Growth PM, under 500 words, and respects the requested prioritised recommendation format.

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

The requested memo is present, complete, addressed to the growth PM, and actionable within the length limit.

Gemini 3.5 Flash-Lite · Gemini

Memo is present, complete, within length, and gives actionable prioritized recommendations.

GPT-6 Luna · API

The required memo is complete, actionable, reader-appropriate, and within length.

Prioritises by impactRightRightRight
Gemini 3.8 Flash · API

The fixes are ranked in a plausible impact order centred on reaching three workouts in week one.

Gemini 3.5 Flash-Lite · Gemini

Fixes are ranked by priority with impact levels (high/medium), targeting activation.

GPT-6 Luna · API

Changes are ranked by likely impact on getting users to three workouts in week one, with onboarding/paywall changes deprioritised.

Activation defined by what predicts retentionRightRightRight
Gemini 3.8 Flash · API

It defines activation as three workouts in week one, links that to day-60 retention, and ranks fixes by their effect on that activation event.

Gemini 3.5 Flash-Lite · Gemini

Activation is defined as three workouts, linked to retention (4x), and ranking is based on effect on this event.

GPT-6 Luna · API

It defines three workouts in week one as the activation behaviour linked to day-60 retention and ranks fixes by their effect on that event.

Gets the base of every number rightRightRightRight
Gemini 3.8 Flash · API

The percentages and 4x relationship are used against the correct funnel steps and denominators.

Gemini 3.5 Flash-Lite · Gemini

All numbers are directly from the supplied funnel; no derived figures require base checks.

GPT-6 Luna · API

No percentages are derived incorrectly, and the memo flags ambiguous funnel bases rather than misusing them.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 82% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT94.4100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude80.665.02None
5Gemini 3.5 Flash-LitewithGemini69.455.02None
6Sonnet 5.5withAPI77.845.02None
7Gemini 3.8 FlashwithAPI69.450.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review