Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 58% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable92% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims60% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly63% pass
    The output claims that users charged on day 7 “churn or ask for refunds”, which is not in the supplied context and is not supported by the data.
    Opus 5.5 · Claude · Fitness app first week
  3. Identifies material uncertainty69% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review our onboarding flow and funnel below. Write a memo to the onboarding squad saying where we are losing people and what to fix, in the order you would fix it, with the reason for each. Keep it under 500 words.

ScenarioChartwise is a self-serve product-analytics tool. Most people who sign up are product managers or marketers, not data engineers. Pricing starts at $49 a month after a 14-day free trial.
Funnel (last 90 days, 12,400 signups)Signup 100% → Connect data source 41% → First chart 33% → Invite teammate 9%. Week-2 retention: 62% for users who reach a first chart, 11% otherwise.
Onboarding flow, screen by screen1. Sign up: email, password, company name. 2. Verify email: the flow pauses until the link is clicked. 3. “What's your role?” (seven options). 4. “How big is your team?” (five options). 5. Connect your data: choose Postgres, BigQuery, Snowflake or CSV upload. There is no sample data and no skip. 6. Connection form: host, port, database, username, password and SSL mode. The help link opens a 14-page setup guide. 7. Choose tables: a list of every table in the database. 8. Syncing: a progress bar that can take up to 20 minutes; leaving the page cancels the sync. 9. Pick a dashboard template (twelve options). 10. First chart: a chart builder with an empty canvas. 11. Invite teammates: email fields, with a “Skip for now” link.
SupportThe most common onboarding ticket (41% of them) is some version of “I don't have the database credentials.” The next is “my sync failed” (18%).
What a strong answer does

The data-source connection step (screens 5–8) is the activation bottleneck; prioritise a sample-data or CSV-first path to a first chart, then non-admin credential handoff and a sync that survives leaving the page, ahead of cosmetic fixes.

Case

v1.6 · anonymised real · B2B SaaS, data

Usable with a quick edit6 of 9 checks passed

The evidence and priorities are usable, but it needs a sentence stating what would change the order and softer labelling of causal claims; no re-checking or redoing of the reasoning is needed.

Output

To: Onboarding Squad Subject: Where onboarding is leaking—and what to fix first

The biggest loss is before a data source is connected: only 41% of signups get that far. Another 8 percentage points drop before a first chart. Just 9% of signups invite a teammate. The strongest activation signal is the chart: week-2 retention is 62% for users who reach one, versus 11% for those who don’t. Prioritize getting people to a useful chart; treat invites as a later step.

Fixes, in priority order

  1. Offer a sample-data path that reaches a real first chart. Let users skip connecting their own data and explore a preloaded dataset, with a clear route to connect later. Many signups are PMs and marketers, and the most common onboarding ticket is missing database credentials (41%). A required technical connection is an especially poor first step for them. This also gives users value while they wait for access.
  1. Make connecting data simpler and recoverable. Replace or supplement the 14-page guide with short, connector-specific instructions and troubleshooting. Validate credentials early, explain the required access, and offer a clear retry path. Don’t cancel a sync when someone leaves the page; send progress and failure notices and let users resume. The credential and sync tickets point to concrete blockers, and successful connection is the gateway to the high-retention chart experience.
  1. Get users to a chart with fewer decisions. Replace the database-wide table list with suggested tables and plain-language descriptions; cut the 12 dashboard templates to a few recommended starting points. Provide a prebuilt chart or guided first-chart flow rather than an empty canvas. The 8-point drop between connection and first chart, plus the sync burden, suggests too much work remains after connection.
  1. Remove early friction that doesn’t create value. Test deferring email verification until after users enter the product, and move role/team questions until they can personalize something useful. These steps currently delay access before users see value.
  1. Make inviting a teammate contextual, not a hurdle. Keep it optional; prompt after a user has made or shared a chart, and explain the benefit. Only 9% of signups invite someone, but chart reach—not inviting—is the clearest retention signal.

Track step conversion and time-to-first-chart by path, especially sample-data versus live connection, before expanding the changes.

Check by check

Got wrong · 2
  • Identifies material uncertaintyIt does not name specific unknowns that could change the priorities or say how resolving them would alter the call.
  • Avoids unsupported claimsIt presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
Mixed · 1
  • Addresses the actual decisionThe memo commits to a clear ranked order but never states what result or condition would change that order; the closing sentence is only about expanding changes.The two graders disagreed on this one.
Got right · 6
  • Uses the supplied evidence correctlyAll factual statements about the current situation are taken correctly from the brief and supplied context, with no invented facts.
  • Respects explicit constraintsIt is a memo to the onboarding squad, well under 500 words, and respects the requested form and reader.
  • Produces the required deliverableThe required memo is complete, actionable, and usable by the onboarding squad as written or with light edits.
  • Prioritises by impactFixes are ranked by likely impact on reaching a first chart, tied to the retention signal.
  • Activation defined by what predicts retentionIt names first chart as the activation behaviour based on 62% vs 11% retention and ranks fixes by their effect on it.
  • Gets the base of every number rightAll percentages use the correct denominators, and the 41% ticket figure is understood as a share of onboarding tickets rather than users.

Grades and run details

Decision model 89 · LLM judge 7 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly33%
  • passAddresses the actual decision55%
  • passRespects explicit constraints74%
  • partialIdentifies material uncertainty63%
  • partialAvoids unsupported claims52%
  • passProduces the required deliverable97%
  • passPrioritises by impact98%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right80%
Run
Run
#1
API response time
12 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT97.2100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude86.165.02None
5Sonnet 5.5withAPI77.845.02None
6Gemini 3.5 Flash-LitewithGemini61.155.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review