Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable93% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims54% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly54% pass
    It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
    Sonnet 5.5 · API · Fitness app first week
  3. Identifies material uncertainty59% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review our onboarding flow and funnel below. Write a memo to the onboarding squad saying where we are losing people and what to fix, in the order you would fix it, with the reason for each. Keep it under 500 words.

What the model was given4 items: Scenario, Funnel (last 90 days, 12,400 signups), Onboarding flow, screen by screen, Support
ScenarioChartwise is a self-serve product-analytics tool. Most people who sign up are product managers or marketers, not data engineers. Pricing starts at $49 a month after a 14-day free trial.
Funnel (last 90 days, 12,400 signups)Signup 100% → Connect data source 41% → First chart 33% → Invite teammate 9%. Week-2 retention: 62% for users who reach a first chart, 11% otherwise.
Onboarding flow, screen by screen1. Sign up: email, password, company name. 2. Verify email: the flow pauses until the link is clicked. 3. “What's your role?” (seven options). 4. “How big is your team?” (five options). 5. Connect your data: choose Postgres, BigQuery, Snowflake or CSV upload. There is no sample data and no skip. 6. Connection form: host, port, database, username, password and SSL mode. The help link opens a 14-page setup guide. 7. Choose tables: a list of every table in the database. 8. Syncing: a progress bar that can take up to 20 minutes; leaving the page cancels the sync. 9. Pick a dashboard template (twelve options). 10. First chart: a chart builder with an empty canvas. 11. Invite teammates: email fields, with a “Skip for now” link.
SupportThe most common onboarding ticket (41% of them) is some version of “I don't have the database credentials.” The next is “my sync failed” (18%).
What a strong answer doesThe answer key the graders mark against

The data-source connection step (screens 5–8) is the activation bottleneck; prioritise a sample-data or CSV-first path to a first chart, then non-admin credential handoff and a sync that survives leaving the page, ahead of cosmetic fixes.

Case

Mid-level · v1.8 · anonymised real · B2B SaaS, data

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit4 of 9 checks passed

The output misinterprets the support ticket data as a user proportion and presents unsupported causal claims, so a PM would need to re-check the evidence before using it.

Output

To: Onboarding Squad Subject: Fixing Our Activation Funnel: Priorities & Action Plan

Where We Are Losing People

Our biggest drop-off occurs at Step 5–8 (Connect Data): 59% of signups leave without connecting a data source.

We are asking PMs and marketers for database ports, SSL modes, and passwords they do not possess. 41% of our support tickets are “I don't have credentials,” followed by 18% for failed syncs. Furthermore, requiring a 20-minute, open-tab sync kills momentum.

Because Week-2 retention jumps from 11% to 62% once a user builds their first chart, our sole objective is getting non-technical users to that "aha!" moment in under two minutes.

---

What to Fix (In Priority Order)

#### 1. Add “Explore with Sample Data” & a Clear “Skip” Button (Step 5) * What: Let users bypass database connection entirely using a fully populated e-commerce or SaaS demo dataset. * Why: This is the highest-leverage fix. It immediately unblocks the 59% who drop off at Step 5 and the 41% who lack DB credentials, giving them an instant path to create a first chart and experience the product’s value on Day 1.

#### 2. Introduce “Invite a Data Teammate” at Connection (Step 6) * What: Add a 1-click option: “Don't have credentials? Send an invite link to your data lead to connect your warehouse.” * Why: Replaces a dead end with collaboration. The PM stays in the app using sample data while the engineer handles credentials asynchronously.

#### 3. Make Syncing Asynchronous (Step 8) * What: Run syncing in the background. Notify users via email/in-app alert when ready instead of forcing them to keep a 20-minute tab open. * Why: Directly addresses our second-biggest support issue (18% sync failures/cancellations) and eliminates the drop-off caused by tab abandonment.

#### 4. Pre-Populate the Dashboard; Kill the Blank Canvas (Steps 9 & 10) * What: When a user selects a template, generate it automatically with 3–4 standard charts instead of dropping them onto an empty canvas. Prompt them to tweak one existing chart. * Why: Bridges the 8% drop-off between connecting data (41%) and creating a first chart (33%). Editing an existing chart is far lower friction than building from scratch.

#### 5. Defer Email Verification and Profiling Questions (Steps 2, 3, & 4) * What: Move verification to a non-blocking banner at the top of the app, and cut or defer the role/team-size questions until after the user creates their first chart. * Why: Every barrier before core value reduces trial-to-paid conversion. Let users reach the product before demanding administrative tasks.

---

Next Step

Let’s implement Priority 1 (Sample Data) this sprint. It requires zero infrastructure changes to our warehouse sync logic and directly targets our biggest activation leak.

Claims the judge couldn’t find in the brief

  • Requiring a 20-minute, open-tab sync kills momentum.
  • The 41% who lack DB credentials (as a proportion of users).
  • The drop-off caused by tab abandonment.

Grades and run details

Decision model 72 · LLM judge 4 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly22%
  • passAddresses the actual decision93%
  • passRespects explicit constraints40%
  • failIdentifies material uncertainty99%
  • partialAvoids unsupported claims39%
  • passProduces the required deliverable96%
  • passPrioritises by impact98%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right44%
Run
Run
#1
API response time
21 s
Submitted
1 Oct 2026

Needs more than a quick edit5 of 9 checks passed

The output presents causal interpretations as certain and omits conditions that could change the priority, so a product manager would need to add caveats and uncertainty before sending. The underlying evidence and reasoning are sound, but it requires more than just a quick edit to align with a demanding senior PM's standard for acknowledging uncertainty.

Output

MEMORANDUM

To: Onboarding Squad From: Product Team Subject: Onboarding Funnel Analysis & Prioritized Fixes

Our last 90 days show 12,400 signups, but we face a steep 59% drop-off immediately after signup (reaching only 41% at data source connection), followed by a massive drop at team invites (down to 9%). Because week-2 retention jumps from 11% to 62% for users who build a first chart, removing friction before that milestone is critical.

Here is where we are losing people and the fixes, ordered by priority:

1. The Data Connection Blockade (Fix First)

  • Where we lose people: We drop 59% of users right at the data connection step.2 Support tickets confirm why: 41% lack database credentials, and 18% experience sync failures. Forcing product managers and marketers—who are not data engineers—to input host, port, SSL modes, and select from raw database tables while blocking the UI for up to 20 minutes is a conversion killer.
  • What to fix:
  • Add a pre-loaded Sample Data / Demo Workspace option so users can experience the product immediately without connecting a live database.
  • Make data connection skippable or move it after the first chart experience.
  • Simplify the CSV upload option or provide pre-built connectors with clear, inline tooltips instead of a 14-page setup guide.

2. The Premature Team Invite Wall (Fix Second)

  • Where we lose people: Reaching a first chart drops to 33%, but inviting teammates plummets to 9%—a catastrophic 24-percentage-point drop right before value is fully realized.1 Users are being asked to advocate for a tool before they’ve even proven its worth to themselves.
  • What to fix:
  • Remove the team invite step from the core linear onboarding path entirely.
  • Relocate the invite prompt to the dashboard or settings menu once the user has successfully built and saved their first chart.

3. Redundant Survey Friction (Fix Third)

  • Where we lose people: Screens 3 ("What's your role?") and 4 ("How big is your team?") add unnecessary micro-friction right after email verification.
  • What to fix:
  • Combine these into a single, clean optional screen or infer them from workspace setup later to streamline time-to-value.

Fixing data connection roadblocks and removing premature team invites will directly protect our week-2 retention (which hinges on reaching that first chart).

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: Funnel (last 90 days, 12,400 signups)Connect data source 41% → First chart 33% → Invite teammate 9%.

    What we’d changeRank the 41% to 33% drop between connecting and a first chart second. The invite comes after the first chart, so it isn't what stands between users and value.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeThe 59% is lost somewhere between screens 2 and 8. Say the tickets point to the connection form, and instrument the screens to confirm it.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly20%
  • passAddresses the actual decision87%
  • passRespects explicit constraints68%
  • failIdentifies material uncertainty99%
  • partialAvoids unsupported claims38%
  • passProduces the required deliverable82%
  • passPrioritises by impact64%
  • passActivation defined by what predicts retention98%
  • passGets the base of every number right37%
Run
Run
#1
Time to output
8 s
Submitted
25 Sept 2026

Usable with a quick edit6 of 9 checks passed

The memo is actionable and fully reasoned; a PM could trim a few words to meet the word limit without redoing any analysis.

Output

Onboarding: where we lose people and what to fix first

Sep 25, 2026 · @Martin Slaney

To: Onboarding squad

Bottom line

Onboarding works for people who reach a first chart: 62% of them are still here in week 2, against 11% of everyone else. Two-thirds of signups never get there, and most of those stop at "Connect your data", where we ask product managers and marketers for database credentials they don't have.

Where we lose people (last 90 days, 12,400 signups)

  • Signup → connect data: 59% lost (\~7,300 people). The biggest leak by far. The top support ticket (41%) is "I don't have the database credentials."
  • Connect → first chart: \~990 more lost, a fifth of those who connected. Syncs take up to 20 minutes, die if you leave the page, and "my sync failed" is the second-biggest ticket (18%).
  • First chart → invite: 33% → 9%. Real, but downstream of the first two.

Caveat: we have no step-level data for screens 2–8, so we can't split the 59% between verification, the two questions and the connection form. The ticket data points hard at the form.

What to fix, in order

  1. Let people reach a first chart without their own data. Add a sample dataset and a skip on screen 5, and land them on a ready-made chart. Why: it goes around the step that loses most people, and the first chart is the moment that predicts retention.
  2. Let them hand off the credentials step. "Ask a teammate to connect" sends an engineer a secure link to fill in the connection form. Why: it answers the 41% ticket directly and turns the blocker into an invite.
  3. Make syncing forgiving. Run it in the background, keep it going when they leave, email them when it's ready, and load a small sample of rows first so a chart appears in minutes. Show plain-language errors with a retry. Why: a 20-minute wait that punishes leaving the page is where connected users die, and it's the 18% ticket.
  4. Don't start from an empty canvas. Use the role answer from screen 3 to suggest two or three templates instead of twelve, and draw the first chart automatically. Why: we already ask the role and do nothing with it; a blank builder is hard for non-analysts.
  5. Trim the steps before connecting. Let people in before they verify their email (verify within a day), and drop or defer the team-size question. Why: cheap to do, but unmeasured, so instrument first rather than assume it's a big win.
  6. Move the invite prompt to after value is shown, such as when someone saves or shares a chart. Why: people invite colleagues to something useful, not to a setup screen.

How we'll know it worked

  • Instrument every screen, 1 to 11, before shipping anything.1
  • Judge fix 1 by how many sample-data users go on to connect real data and are retained in week 2, not by first-chart rate. A chart of sample data can lift that metric without anyone sticking around, so don't assume the 62% carries over.

What a PM had to fix

  1. 1OtherFix itQuick edit

    What we’d changeInstrument alongside fix 1 rather than before it, and merge fixes 1 and 4, which both land the user on an auto-drawn chart.

Grades and run details

Decision model 94 · LLM judge 7 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly22%
  • passAddresses the actual decision99%
  • passRespects explicit constraints23%
  • passIdentifies material uncertainty93%
  • partialAvoids unsupported claims31%
  • passProduces the required deliverable93%
  • passPrioritises by impact96%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right50%
Run
Run
#1
Time to output
40 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyWrongMixedRight
Gemini 3.8 Flash · API

The output presents 'the 41% who lack DB credentials' as a user proportion, which is not supported by the ticket share data, and states 'kills momentum' and 'drop-off caused by tab abandonment' as facts without evidence.

Gemini 3.5 Flash-Lite · Gemini

All factual claims about the current situation are directly from the brief or context, or derived correctly by arithmetic.

Opus 5.5 · Claude

All factual statements about the current situation are directly from the supplied context or derived by correct arithmetic.

Addresses the actual decisionMixedMixedRight
Gemini 3.8 Flash · API

The output does not state what result or condition would change the prioritised order, failing the requirement to say what would change the answer.

Gemini 3.5 Flash-Lite · Gemini

The output commits to a ranked order but does not state what result or condition would change that answer, as required by the criterion.

Opus 5.5 · Claude

Commits to a clear, ranked order of fixes early in the memo, framed for the onboarding squad.

Respects explicit constraintsRightRightMixed
Gemini 3.8 Flash · API

The memo is under 500 words, addressed to the onboarding squad, and provides a ranked list with reasons.

Gemini 3.5 Flash-Lite · Gemini

The output is a memo to the onboarding squad, well under 500 words, and addresses the requested points.

Opus 5.5 · Claude

The output is exactly 500 words; the brief requires under 500 words.

Identifies material uncertaintyWrongWrongMixed
Gemini 3.8 Flash · API

No unknowns are named, and the output does not discuss what could change the decision or how to resolve uncertainty.

Gemini 3.5 Flash-Lite · Gemini

The output does not identify unknowns that could change the decision (e.g., whether sample data actually drives first-chart creation) or say how they could be resolved.

Opus 5.5 · Claude

Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.

Avoids unsupported claimsWrongWrongRight
Gemini 3.8 Flash · API

Interpretations like 'kills momentum', 'the 41% who lack DB credentials' as a user share, and 'drop-off caused by tab abandonment' are presented as established facts without labelling them as hypotheses.

Gemini 3.5 Flash-Lite · Gemini

Causal claims like 'conversion killer', 'premature team invite wall', and 'unnecessary micro-friction' are stated as facts rather than labelled as hypotheses or inferences.

Opus 5.5 · Claude

Interpretations and causes are presented as reasoning (often within 'Why' sections) and not as established facts.

Produces the required deliverableRightRightMixed
Gemini 3.8 Flash · API

The output is a complete memo with the requested structure, reader, and length, and a PM could act on it with light edits.

Gemini 3.5 Flash-Lite · Gemini

The memo is complete, in the right form, under the word limit, and the squad could act on it with minimal edits.

Opus 5.5 · Claude

The memo is complete and usable but exceeds the word limit: it is 500 words, not under 500.

Gets the base of every number rightMixedRightRight
Gemini 3.8 Flash · API

The output misstates the 41% support-ticket share as a proportion of users lacking credentials, getting the base wrong.

Gemini 3.5 Flash-Lite · Gemini

All percentages and differences are correctly calculated from the supplied data, and the base for support-ticket figures is clearly ticket counts.

Opus 5.5 · Claude

All derived figures (59%, ~7,300, ~990, a fifth) are computed from the correct steps and denominators, with bases stated where needed.

All got right 2

Prioritises by impactRightRightRight
Gemini 3.8 Flash · API

Fixes are ranked by likely impact on activation, starting with the biggest drop-off and linking to the retention-driving first chart.

Gemini 3.5 Flash-Lite · Gemini

Fixes are ranked by likely impact on reaching the first-chart activation milestone that predicts retention.

Opus 5.5 · Claude

Fixes are ranked by impact on activation, with the largest drop addressed first.

Activation defined by what predicts retentionRightRightRight
Gemini 3.8 Flash · API

The output defines activation as building a first chart, cites the retention difference (11% vs 62%), and ranks fixes by their effect on reaching that event.

Gemini 3.5 Flash-Lite · Gemini

The memo explicitly names reaching a first chart as the activation event, links it to the retention jump, and uses that to prioritize fixes.

Opus 5.5 · Claude

Identifies the first chart as the activation event that predicts retention (62% vs 11%) and ranks fixes by their effect on reaching it.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 82% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT94.4100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude80.665.02None
5Gemini 3.5 Flash-LitewithGemini69.455.02None
6Sonnet 5.5withAPI77.845.02None
7Gemini 3.8 FlashwithAPI69.450.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review