Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 58% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable92% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims60% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly63% pass
    The output claims that users charged on day 7 “churn or ask for refunds”, which is not in the supplied context and is not supported by the data.
    Opus 5.5 · Claude · Fitness app first week
  3. Identifies material uncertainty69% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review our onboarding flow and funnel below. Write a memo to the onboarding squad saying where we are losing people and what to fix, in the order you would fix it, with the reason for each. Keep it under 500 words.

ScenarioChartwise is a self-serve product-analytics tool. Most people who sign up are product managers or marketers, not data engineers. Pricing starts at $49 a month after a 14-day free trial.
Funnel (last 90 days, 12,400 signups)Signup 100% → Connect data source 41% → First chart 33% → Invite teammate 9%. Week-2 retention: 62% for users who reach a first chart, 11% otherwise.
Onboarding flow, screen by screen1. Sign up: email, password, company name. 2. Verify email: the flow pauses until the link is clicked. 3. “What's your role?” (seven options). 4. “How big is your team?” (five options). 5. Connect your data: choose Postgres, BigQuery, Snowflake or CSV upload. There is no sample data and no skip. 6. Connection form: host, port, database, username, password and SSL mode. The help link opens a 14-page setup guide. 7. Choose tables: a list of every table in the database. 8. Syncing: a progress bar that can take up to 20 minutes; leaving the page cancels the sync. 9. Pick a dashboard template (twelve options). 10. First chart: a chart builder with an empty canvas. 11. Invite teammates: email fields, with a “Skip for now” link.
SupportThe most common onboarding ticket (41% of them) is some version of “I don't have the database credentials.” The next is “my sync failed” (18%).
What a strong answer does

The data-source connection step (screens 5–8) is the activation bottleneck; prioritise a sample-data or CSV-first path to a first chart, then non-admin credential handoff and a sync that survives leaving the page, ahead of cosmetic fixes.

Case

v1.6 · anonymised real · B2B SaaS, data

Usable with a quick edit9 of 9 checks passed

The memo is a clear, prioritized, evidence-based list with a straightforward measurement plan; a PM could use it with minimal edits.

Output

To: Onboarding squad Subject: Fix access and sync before optimising chart creation

Our biggest measured loss is before data connection: 59% of signups, roughly 7,300 people, never connect a source. We ask PMs and marketers to complete a database setup that many cannot do themselves, then make a lengthy sync depend on keeping a page open.

Another 8 percentage points disappear between connection and first chart: about 20% of those who connect. Teammate invitations fall to 9% of signups, but we should establish whether collaboration is necessary for value before treating that as our next bottleneck.

First-chart users have much higher week-2 retention, 62% versus 11%. That makes chart creation a promising activation milestone, but does not prove that pushing everyone through a chart will cause better retention.

Fix in this order

1. Give users a route forward without database credentials. Offer a clearly labelled sample-data experience and make CSV upload a prominent alternative.1 Add a setup hand-off that lets users ask a colleague to connect the source, then resume when it is ready. Replace the generic setup guide with short, source-specific instructions.

This addresses the largest measured loss and the strongest support signal: missing credentials account for 41% of onboarding tickets. That is a share of tickets, not signups, so we cannot yet quantify how much drop-off it explains. Sample data should demonstrate value while preserving a clear path to the user’s own data.

2. Make syncing durable and recoverable. Run syncs in the background so closing the page does not cancel them. Show meaningful status, explain failures and allow retries without repeating setup. Notify users when their data is ready.

A 20-minute foreground wait is a fragile requirement. Sync failures generate 18% of onboarding tickets and can waste the effort users have already invested.

3. Turn connected data into a useful first chart. Help users identify relevant tables instead of presenting every table equally. Use their chosen template to produce an editable chart from compatible data, with guided choices where mapping is needed. Do not finish template selection with an empty canvas.

This targets the roughly one in five connected users who never reach a chart. Measure whether they produce something useful, not merely whether a chart renders.

4. Remove avoidable preliminary work. Defer role and team-size questions unless their answers immediately improve setup. Review whether email verification must block exploration, retaining verification where required for account security.

These steps add friction before value, but we lack evidence that they drive the main loss.

Measure alongside delivery: instrument each screen, connection method, error and return visit. Test changes against first charts using customers’ own data, week-2 retention and trial-to-paid conversion. Track sample-data charts separately. Revisit invitations once we understand whether users need collaboration to succeed. To: Onboarding squad Subject: Fix access and sync before optimising chart creation

Our biggest measured loss is before data connection: 59% of signups, roughly 7,300 people, never connect a source. We ask PMs and marketers to complete a database setup that many cannot do themselves, then make a lengthy sync depend on keeping a page open.

Another 8 percentage points disappear between connection and first chart: about 20% of those who connect. Teammate invitations fall to 9% of signups, but we should establish whether collaboration is necessary for value before treating that as our next bottleneck.

First-chart users have much higher week-2 retention, 62% versus 11%. That makes chart creation a promising activation milestone, but does not prove that pushing everyone through a chart will cause better retention.

Fix in this order

1. Give users a route forward without database credentials. Offer a clearly labelled sample-data experience and make CSV upload a prominent alternative. Add a setup hand-off that lets users ask a colleague to connect the source, then resume when it is ready. Replace the generic setup guide with short, source-specific instructions.

This addresses the largest measured loss and the strongest support signal: missing credentials account for 41% of onboarding tickets. That is a share of tickets, not signups, so we cannot yet quantify how much drop-off it explains. Sample data should demonstrate value while preserving a clear path to the user’s own data.

2. Make syncing durable and recoverable. Run syncs in the background so closing the page does not cancel them. Show meaningful status, explain failures and allow retries without repeating setup. Notify users when their data is ready.

A 20-minute foreground wait is a fragile requirement. Sync failures generate 18% of onboarding tickets and can waste the effort users have already invested.

3. Turn connected data into a useful first chart. Help users identify relevant tables instead of presenting every table equally. Use their chosen template to produce an editable chart from compatible data, with guided choices where mapping is needed. Do not finish template selection with an empty canvas.

This targets the roughly one in five connected users who never reach a chart. Measure whether they produce something useful, not merely whether a chart renders.

4. Remove avoidable preliminary work. Defer role and team-size questions unless their answers immediately improve setup. Review whether email verification must block exploration, retaining verification where required for account security.

These steps add friction before value, but we lack evidence that they drive the main loss.

Measure alongside delivery: instrument each screen, connection method, error and return visit. Test changes against first charts using customers’ own data, week-2 retention and trial-to-paid conversion. Track sample-data charts separately. Revisit invitations once we understand whether users need collaboration to succeed.

What a PM had to fix

  1. 1Decision deferredMake the callQuick edit

    What we’d changeRank the four parts of fix 1 so the squad knows what ships first, and make the calls fix 4 leaves open.

Check by check

Got right · 9
  • Uses the supplied evidence correctlyAll statements about the current situation are taken correctly from the brief or derived arithmetically; nothing is invented.
  • Addresses the actual decisionCommits early to a ranked order of fixes, addressed to the onboarding squad, and says team invitations will be revisited after understanding if collaboration is needed.
  • Respects explicit constraintsMemo format, under 500 words, directed to the onboarding squad, respecting all explicit constraints.
  • Identifies material uncertaintyIdentifies that first-chart correlation does not prove causation, ticket share is not user drop-off, and role/team questions lack evidence; proposes measurement to resolve.
  • Avoids unsupported claimsInterpretations are grounded in the data and qualified where evidence is suggestive, so no factual claims are presented as established without support.
  • Produces the required deliverableDelivers a complete memo with prioritized fixes, reasons, and measurement plan that the squad can act on with light edits.
  • Prioritises by impactFixes are ranked by likely impact on activation: credential-free path first (largest drop), then sync durability, then first-chart creation, then friction removal.
  • Activation defined by what predicts retentionActivation is defined as first chart, linked to week-2 retention (62% vs 11%), and fixes are judged by their effect on reaching that milestone.
  • Gets the base of every number rightAll percentages are computed from the correct funnel steps, and ticket shares are explicitly noted as share of tickets, not users.

Grades and run details

Decision model 94 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly31%
  • passAddresses the actual decision85%
  • partialRespects explicit constraints17%
  • passIdentifies material uncertainty85%
  • passAvoids unsupported claims30%
  • passProduces the required deliverable63%
  • passPrioritises by impact98%
  • passActivation defined by what predicts retention94%
  • passGets the base of every number right67%
Run
Run
#1
Time to output
40 s
Submitted
25 Sept 2026

Needs more than a quick edit5 of 9 checks passed

The output presents causal interpretations as certain and omits conditions that could change the priority, so a product manager would need to add caveats and uncertainty before sending. The underlying evidence and reasoning are sound, but it requires more than just a quick edit to align with a demanding senior PM's standard for acknowledging uncertainty.

Output

MEMORANDUM

To: Onboarding Squad From: Product Team Subject: Onboarding Funnel Analysis & Prioritized Fixes

Our last 90 days show 12,400 signups, but we face a steep 59% drop-off immediately after signup (reaching only 41% at data source connection), followed by a massive drop at team invites (down to 9%). Because week-2 retention jumps from 11% to 62% for users who build a first chart, removing friction before that milestone is critical.

Here is where we are losing people and the fixes, ordered by priority:

1. The Data Connection Blockade (Fix First)

  • Where we lose people: We drop 59% of users right at the data connection step.2 Support tickets confirm why: 41% lack database credentials, and 18% experience sync failures. Forcing product managers and marketers—who are not data engineers—to input host, port, SSL modes, and select from raw database tables while blocking the UI for up to 20 minutes is a conversion killer.
  • What to fix:
  • Add a pre-loaded Sample Data / Demo Workspace option so users can experience the product immediately without connecting a live database.
  • Make data connection skippable or move it after the first chart experience.
  • Simplify the CSV upload option or provide pre-built connectors with clear, inline tooltips instead of a 14-page setup guide.

2. The Premature Team Invite Wall (Fix Second)

  • Where we lose people: Reaching a first chart drops to 33%, but inviting teammates plummets to 9%—a catastrophic 24-percentage-point drop right before value is fully realized.1 Users are being asked to advocate for a tool before they’ve even proven its worth to themselves.
  • What to fix:
  • Remove the team invite step from the core linear onboarding path entirely.
  • Relocate the invite prompt to the dashboard or settings menu once the user has successfully built and saved their first chart.

3. Redundant Survey Friction (Fix Third)

  • Where we lose people: Screens 3 ("What's your role?") and 4 ("How big is your team?") add unnecessary micro-friction right after email verification.
  • What to fix:
  • Combine these into a single, clean optional screen or infer them from workspace setup later to streamline time-to-value.

Fixing data connection roadblocks and removing premature team invites will directly protect our week-2 retention (which hinges on reaching that first chart).

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: Funnel (last 90 days, 12,400 signups)Connect data source 41% → First chart 33% → Invite teammate 9%.

    What we’d changeRank the 41% to 33% drop between connecting and a first chart second. The invite comes after the first chart, so it isn't what stands between users and value.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeThe 59% is lost somewhere between screens 2 and 8. Say the tickets point to the connection form, and instrument the screens to confirm it.

Check by check

Got wrong · 3
  • Addresses the actual decisionThe output commits to a ranked order but does not state what result or condition would change that answer, as required by the criterion.
  • Identifies material uncertaintyThe output does not identify unknowns that could change the decision (e.g., whether sample data actually drives first-chart creation) or say how they could be resolved.
  • Avoids unsupported claimsCausal claims like 'conversion killer', 'premature team invite wall', and 'unnecessary micro-friction' are stated as facts rather than labelled as hypotheses or inferences.
Mixed · 1
  • Uses the supplied evidence correctlyAll factual claims about the current situation are directly from the brief or context, or derived correctly by arithmetic.The two graders disagreed on this one.
Got right · 5
  • Respects explicit constraintsThe output is a memo to the onboarding squad, well under 500 words, and addresses the requested points.
  • Produces the required deliverableThe memo is complete, in the right form, under the word limit, and the squad could act on it with minimal edits.
  • Prioritises by impactFixes are ranked by likely impact on reaching the first-chart activation milestone that predicts retention.
  • Activation defined by what predicts retentionThe memo explicitly names reaching a first chart as the activation event, links it to the retention jump, and uses that to prioritize fixes.
  • Gets the base of every number rightAll percentages and differences are correctly calculated from the supplied data, and the base for support-ticket figures is clearly ticket counts.

Grades and run details

Decision model 61 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly17%
  • failAddresses the actual decision47%
  • passRespects explicit constraints69%
  • failIdentifies material uncertainty99%
  • partialAvoids unsupported claims40%
  • passProduces the required deliverable86%
  • passPrioritises by impact73%
  • passActivation defined by what predicts retention98%
  • passGets the base of every number right44%
Run
Run
#1
Time to output
8 s
Submitted
25 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT97.2100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude86.165.02None
5Sonnet 5.5withAPI77.845.02None
6Gemini 3.5 Flash-LitewithGemini61.155.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review