Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable93% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims54% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly54% pass
    It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
    Sonnet 5.5 · API · Fitness app first week
  3. Identifies material uncertainty59% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review our onboarding flow and funnel below. Write a memo to the onboarding squad saying where we are losing people and what to fix, in the order you would fix it, with the reason for each. Keep it under 500 words.

What the model was given4 items: Scenario, Funnel (last 90 days, 12,400 signups), Onboarding flow, screen by screen, Support
ScenarioChartwise is a self-serve product-analytics tool. Most people who sign up are product managers or marketers, not data engineers. Pricing starts at $49 a month after a 14-day free trial.
Funnel (last 90 days, 12,400 signups)Signup 100% → Connect data source 41% → First chart 33% → Invite teammate 9%. Week-2 retention: 62% for users who reach a first chart, 11% otherwise.
Onboarding flow, screen by screen1. Sign up: email, password, company name. 2. Verify email: the flow pauses until the link is clicked. 3. “What's your role?” (seven options). 4. “How big is your team?” (five options). 5. Connect your data: choose Postgres, BigQuery, Snowflake or CSV upload. There is no sample data and no skip. 6. Connection form: host, port, database, username, password and SSL mode. The help link opens a 14-page setup guide. 7. Choose tables: a list of every table in the database. 8. Syncing: a progress bar that can take up to 20 minutes; leaving the page cancels the sync. 9. Pick a dashboard template (twelve options). 10. First chart: a chart builder with an empty canvas. 11. Invite teammates: email fields, with a “Skip for now” link.
SupportThe most common onboarding ticket (41% of them) is some version of “I don't have the database credentials.” The next is “my sync failed” (18%).
What a strong answer doesThe answer key the graders mark against

The data-source connection step (screens 5–8) is the activation bottleneck; prioritise a sample-data or CSV-first path to a first chart, then non-admin credential handoff and a sync that survives leaving the page, ahead of cosmetic fixes.

Case

Mid-level · v1.8 · anonymised real · B2B SaaS, data

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 9 checks passed

The memo is prioritised, evidence-based, and actionable as-is. At most one could add an explicit reprioritization trigger, but that would be a light wording edit.

Output

To: Onboarding squad Subject: Prioritize getting users to a useful first chart

Our biggest leak is before data connection: 59% of signups—7,316 people—never connect a source. Another 8 percentage points disappear before the first chart. Only 9% invite a teammate, but collaboration should come after users see value.

First-chart users have 62% week-2 retention versus 11% for everyone else. That makes time-to-first-chart our priority, although this correlation does not prove that creating a chart causes retention.

Fixes, in priority order

1. Remove the database-credentials gate. Offer a prominent sample-data path and a clearly explained CSV path. Let users request access from a data owner without abandoning onboarding. For database connections, replace the 14-page guide with source-specific instructions, inline examples, and actionable error messages.

Why first: Our audience is mostly PMs and marketers, yet onboarding assumes database expertise and access. “I don’t have the credentials” accounts for 41% of onboarding tickets. Sample data can demonstrate value immediately, but we should separately measure whether those users eventually connect their own data.

2. Make syncing resilient and asynchronous. Run syncs server-side so leaving the page does not cancel them. Show status, provide completion notifications, and support retries with clear failure explanations. Where technically feasible, sync a useful subset first rather than requiring a full load.

Why second: A wait of up to 20 minutes is already costly; requiring an open tab makes it fragile. Sync failures account for another 18% of tickets and block users who have overcome the access hurdle.

3. Deliver a chart, not another blank canvas. Recommend relevant tables rather than listing everything. Reduce twelve templates to a few guided choices, then populate a starter chart that users can edit.

Why third: Choosing a template and then seeing an empty builder breaks the promise of guidance. About 20% of connected users fail to reach a first chart. Focus on a useful result, not merely recording a chart-creation event.

4. Shorten the route to value. Move role and team-size questions until after the first chart, unless they directly personalize onboarding. Allow users to explore before email verification where security requirements permit.

Why fourth: These steps add friction before any payoff, but we lack screen-level evidence that they explain the largest losses.

5. Make invitations contextual. Prompt users to share a useful chart after success; keep invitations optional. Distinguish “ask a data owner for access” from “invite a collaborator.”

Why fifth: Low invitation rates may reflect weak activation, not a broken invitation screen.

Instrument every screen, connection method, and sync failure. Evaluate changes on useful first-chart completion, time to value, own-data activation, and week-2 retention—not just step clicks.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly28%
  • passAddresses the actual decision96%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty34%
  • passAvoids unsupported claims48%
  • passProduces the required deliverable95%
  • passPrioritises by impact98%
  • passActivation defined by what predicts retention99%
  • passGets the base of every number right55%
Run
Run
#1
API response time
17 s
Submitted
29 Sept 2026

Usable with a quick edit6 of 9 checks passed

The memo is actionable and fully reasoned; a PM could trim a few words to meet the word limit without redoing any analysis.

Output

Onboarding: where we lose people and what to fix first

Sep 25, 2026 · @Martin Slaney

To: Onboarding squad

Bottom line

Onboarding works for people who reach a first chart: 62% of them are still here in week 2, against 11% of everyone else. Two-thirds of signups never get there, and most of those stop at "Connect your data", where we ask product managers and marketers for database credentials they don't have.

Where we lose people (last 90 days, 12,400 signups)

  • Signup → connect data: 59% lost (\~7,300 people). The biggest leak by far. The top support ticket (41%) is "I don't have the database credentials."
  • Connect → first chart: \~990 more lost, a fifth of those who connected. Syncs take up to 20 minutes, die if you leave the page, and "my sync failed" is the second-biggest ticket (18%).
  • First chart → invite: 33% → 9%. Real, but downstream of the first two.

Caveat: we have no step-level data for screens 2–8, so we can't split the 59% between verification, the two questions and the connection form. The ticket data points hard at the form.

What to fix, in order

  1. Let people reach a first chart without their own data. Add a sample dataset and a skip on screen 5, and land them on a ready-made chart. Why: it goes around the step that loses most people, and the first chart is the moment that predicts retention.
  2. Let them hand off the credentials step. "Ask a teammate to connect" sends an engineer a secure link to fill in the connection form. Why: it answers the 41% ticket directly and turns the blocker into an invite.
  3. Make syncing forgiving. Run it in the background, keep it going when they leave, email them when it's ready, and load a small sample of rows first so a chart appears in minutes. Show plain-language errors with a retry. Why: a 20-minute wait that punishes leaving the page is where connected users die, and it's the 18% ticket.
  4. Don't start from an empty canvas. Use the role answer from screen 3 to suggest two or three templates instead of twelve, and draw the first chart automatically. Why: we already ask the role and do nothing with it; a blank builder is hard for non-analysts.
  5. Trim the steps before connecting. Let people in before they verify their email (verify within a day), and drop or defer the team-size question. Why: cheap to do, but unmeasured, so instrument first rather than assume it's a big win.
  6. Move the invite prompt to after value is shown, such as when someone saves or shares a chart. Why: people invite colleagues to something useful, not to a setup screen.

How we'll know it worked

  • Instrument every screen, 1 to 11, before shipping anything.1
  • Judge fix 1 by how many sample-data users go on to connect real data and are retained in week 2, not by first-chart rate. A chart of sample data can lift that metric without anyone sticking around, so don't assume the 62% carries over.

What a PM had to fix

  1. 1OtherFix itQuick edit

    What we’d changeInstrument alongside fix 1 rather than before it, and merge fixes 1 and 4, which both land the user on an auto-drawn chart.

Grades and run details

Decision model 94 · LLM judge 7 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly22%
  • passAddresses the actual decision99%
  • passRespects explicit constraints23%
  • passIdentifies material uncertainty93%
  • partialAvoids unsupported claims31%
  • passProduces the required deliverable93%
  • passPrioritises by impact96%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right50%
Run
Run
#1
Time to output
40 s
Submitted
25 Sept 2026

Needs more than a quick edit3 of 9 checks passed

It exceeds the word limit and contains an unsupported user-capability claim. A PM would need to recheck evidence and trim before use.

Output

To: Onboarding squad Re: Where we lose people and what to fix, in order

Summary: Reaching a first chart is what predicts retention (62% at week 2 versus 11% otherwise). Only 33% of signups get there, and the biggest loss, 59% of signups, happens before they connect data. Most of our users are PMs and marketers who can't complete a database connection alone.

Fixes, in priority order

1. Add sample data and a skip option on the "Connect your data" screen. The most common ticket (41%) is "I don't have the database credentials." We ask non-engineers for host, port and SSL mode before they've seen any value. Sample data lets them reach a chart in minutes, and they can connect real data later. Caveat: sample-data users may retain worse than the 62% figure, which likely reflects real data. Track them as a separate cohort and prompt them to connect real data after their first chart.

2. Make it easy to get credentials from someone else. Add a "Send this to your engineer" option: a shareable link or prefilled email containing the connection form and a short setup checklist. Replace the 14-page guide with inline, per-field help. Promote CSV upload as the no-engineer path, since it's currently one of four equal options.

3. Fix the sync experience and its failures. Syncs run up to 20 minutes and cancel if the user leaves. Run them in the background, send an email when they finish, and let users start exploring once the first tables land. Then look at the 18% of tickets that say "my sync failed" and surface specific error messages. In the table-selection step, preselect a sensible default instead of listing every table.

4. Get people to a first chart without an empty canvas. We show twelve templates, then a blank chart builder. Cut the choice to two or three based on role, and auto-generate a first chart from the synced data (or sample data) that the user can edit. This targets the 41% → 33% drop and shortens time to value.

5. Remove friction at the front. Let users into the product before email verification, and enforce it later (e.g., before inviting teammates). Move the role question ahead of template selection, where it drives the default, and drop or defer team size unless something uses it. These are cheap fixes, but I'm ranking them below the others because I can't yet show they're a big loss.

6. Move the invite prompt to after the first chart. The 9% invite rate is 27% of users who reach a chart, so it's mostly a downstream symptom. Fix activation first. Then trigger the invite prompt right after a first chart is created, with a message like "Share this chart," rather than as a generic final step.

Measure first The funnel jumps from signup straight to data connection, so we can't see how many people drop at email verification, role, team size, or the connection form. Add per-screen events this sprint, before changing anything, so we can confirm the ordering above and measure each fix.

Success metrics: signup → first chart (33% today), share of users who connect real data within 7 days, sync failure rate, and week-2 retention split by sample-data versus real-data users.

Claims the judge couldn’t find in the brief

  • Most users are PMs and marketers who can't complete a database connection alone.

Grades and run details

Decision model 83 · LLM judge 3 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly7%
  • passAddresses the actual decision99%
  • passRespects explicit constraints24%
  • passIdentifies material uncertainty90%
  • partialAvoids unsupported claims43%
  • passProduces the required deliverable93%
  • passPrioritises by impact95%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right42%
Run
Run
#1
API response time
15 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyRightRightWrong
GPT-6.1 Sol · API

All current-state facts and figures in the output are drawn or derived correctly from the supplied funnel, scenario, support data, and onboarding flow.

Opus 5.5 · Claude

All factual statements about the current situation are directly from the supplied context or derived by correct arithmetic.

Sonnet 5.5 · API

It presents as fact that PMs and marketers cannot complete a database connection alone, which the supplied context does not establish.

Addresses the actual decisionRightRightMixed
GPT-6.1 Sol · API

It commits early to prioritising time-to-first-chart, gives a ranked fix order for the onboarding squad, and flags retention and own-data activation as conditions that would change the approach.

Opus 5.5 · Claude

Commits to a clear, ranked order of fixes early in the memo, framed for the onboarding squad.

Sonnet 5.5 · API

It commits to a ranked order but does not state a result or condition that would change the answer, only that per-screen events would confirm it.

Respects explicit constraintsRightMixedMixed
GPT-6.1 Sol · API

The output is a memo addressed to the onboarding squad, stays under 500 words, and provides the requested ordered fixes.

Opus 5.5 · Claude

The output is exactly 500 words; the brief requires under 500 words.

Sonnet 5.5 · API

The memo is approximately 541 words, exceeding the explicit 500-word limit.

Identifies material uncertaintyRightMixedMixed
GPT-6.1 Sol · API

It names the causal uncertainty around first-chart retention, the risk that sample-data users may not connect their own data, and the lack of screen-level evidence, with ways to resolve them.

Opus 5.5 · Claude

Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.

Sonnet 5.5 · API

It names sample-data retention and missing funnel events but does not specify what data result would change the prioritisation.

Avoids unsupported claimsRightRightWrong
GPT-6.1 Sol · API

Interpretive statements are either supported by the supplied flow or explicitly hedged as possibilities or caveats rather than established fact.

Opus 5.5 · Claude

Interpretations and causes are presented as reasoning (often within 'Why' sections) and not as established facts.

Sonnet 5.5 · API

The user-capability claim and 'mostly downstream symptom' are stated as established fact rather than labelled interpretations.

Produces the required deliverableRightMixedMixed
GPT-6.1 Sol · API

The requested memo is present, complete, actionable, and usable by the onboarding squad with little or no editing.

Opus 5.5 · Claude

The memo is complete and usable but exceeds the word limit: it is 500 words, not under 500.

Sonnet 5.5 · API

It is a usable memo in form and audience, but it does not meet the required length constraint.

All got right 3

Prioritises by impactRightRightRight
GPT-6.1 Sol · API

Fixes are ranked by likely impact on activation, starting with the data-connection bottleneck and moving to later-stage friction.

Opus 5.5 · Claude

Fixes are ranked by impact on activation, with the largest drop addressed first.

Sonnet 5.5 · API

Fixes are explicitly ranked by likely impact on reaching first chart, with reasons for each.

Activation defined by what predicts retentionRightRightRight
GPT-6.1 Sol · API

It defines the activation event as reaching a useful first chart, ties it to retention, and ranks fixes by effect on that outcome rather than step completion.

Opus 5.5 · Claude

Identifies the first chart as the activation event that predicts retention (62% vs 11%) and ranks fixes by their effect on reaching it.

Sonnet 5.5 · API

It defines first chart as the activation behaviour using the retention split and ranks fixes by effect on that event.

Gets the base of every number rightRightRightRight
GPT-6.1 Sol · API

Percentages and differences use the correct denominators, step-to-step drops are computed from the right stages, and ambiguous bases are clarified where needed.

Opus 5.5 · Claude

All derived figures (59%, ~7,300, ~990, a fifth) are computed from the correct steps and denominators, with bases stated where needed.

Sonnet 5.5 · API

Derived percentages use the correct denominators, including 59% before data, 27% invite of first-chart users, and the 41% to 33% step.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 82% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT94.4100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude80.665.02None
5Gemini 3.5 Flash-LitewithGemini69.455.02None
6Sonnet 5.5withAPI77.845.02None
7Gemini 3.8 FlashwithAPI69.450.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review