Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable93% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims54% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly54% pass
    It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
    Sonnet 5.5 · API · Fitness app first week
  3. Identifies material uncertainty59% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review our onboarding flow and funnel below. Write a memo to the onboarding squad saying where we are losing people and what to fix, in the order you would fix it, with the reason for each. Keep it under 500 words.

What the model was given4 items: Scenario, Funnel (last 90 days, 12,400 signups), Onboarding flow, screen by screen, Support
ScenarioChartwise is a self-serve product-analytics tool. Most people who sign up are product managers or marketers, not data engineers. Pricing starts at $49 a month after a 14-day free trial.
Funnel (last 90 days, 12,400 signups)Signup 100% → Connect data source 41% → First chart 33% → Invite teammate 9%. Week-2 retention: 62% for users who reach a first chart, 11% otherwise.
Onboarding flow, screen by screen1. Sign up: email, password, company name. 2. Verify email: the flow pauses until the link is clicked. 3. “What's your role?” (seven options). 4. “How big is your team?” (five options). 5. Connect your data: choose Postgres, BigQuery, Snowflake or CSV upload. There is no sample data and no skip. 6. Connection form: host, port, database, username, password and SSL mode. The help link opens a 14-page setup guide. 7. Choose tables: a list of every table in the database. 8. Syncing: a progress bar that can take up to 20 minutes; leaving the page cancels the sync. 9. Pick a dashboard template (twelve options). 10. First chart: a chart builder with an empty canvas. 11. Invite teammates: email fields, with a “Skip for now” link.
SupportThe most common onboarding ticket (41% of them) is some version of “I don't have the database credentials.” The next is “my sync failed” (18%).
What a strong answer doesThe answer key the graders mark against

The data-source connection step (screens 5–8) is the activation bottleneck; prioritise a sample-data or CSV-first path to a first chart, then non-admin credential handoff and a sync that survives leaving the page, ahead of cosmetic fixes.

Case

Mid-level · v1.8 · anonymised real · B2B SaaS, data

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit6 of 9 checks passed

The evidence and priorities are usable, but it needs a sentence stating what would change the order and softer labelling of causal claims; no re-checking or redoing of the reasoning is needed.

Output

To: Onboarding Squad Subject: Where onboarding is leaking—and what to fix first

The biggest loss is before a data source is connected: only 41% of signups get that far. Another 8 percentage points drop before a first chart. Just 9% of signups invite a teammate. The strongest activation signal is the chart: week-2 retention is 62% for users who reach one, versus 11% for those who don’t. Prioritize getting people to a useful chart; treat invites as a later step.

Fixes, in priority order

  1. Offer a sample-data path that reaches a real first chart. Let users skip connecting their own data and explore a preloaded dataset, with a clear route to connect later. Many signups are PMs and marketers, and the most common onboarding ticket is missing database credentials (41%). A required technical connection is an especially poor first step for them. This also gives users value while they wait for access.
  1. Make connecting data simpler and recoverable. Replace or supplement the 14-page guide with short, connector-specific instructions and troubleshooting. Validate credentials early, explain the required access, and offer a clear retry path. Don’t cancel a sync when someone leaves the page; send progress and failure notices and let users resume. The credential and sync tickets point to concrete blockers, and successful connection is the gateway to the high-retention chart experience.
  1. Get users to a chart with fewer decisions. Replace the database-wide table list with suggested tables and plain-language descriptions; cut the 12 dashboard templates to a few recommended starting points. Provide a prebuilt chart or guided first-chart flow rather than an empty canvas. The 8-point drop between connection and first chart, plus the sync burden, suggests too much work remains after connection.
  1. Remove early friction that doesn’t create value. Test deferring email verification until after users enter the product, and move role/team questions until they can personalize something useful. These steps currently delay access before users see value.
  1. Make inviting a teammate contextual, not a hurdle. Keep it optional; prompt after a user has made or shared a chart, and explain the benefit. Only 9% of signups invite someone, but chart reach—not inviting—is the clearest retention signal.

Track step conversion and time-to-first-chart by path, especially sample-data versus live connection, before expanding the changes.

Grades and run details

Decision model 89 · LLM judge 7 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly33%
  • passAddresses the actual decision98%
  • passRespects explicit constraints71%
  • partialIdentifies material uncertainty65%
  • partialAvoids unsupported claims48%
  • passProduces the required deliverable97%
  • passPrioritises by impact99%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right79%
Run
Run
#1
API response time
12 s
Submitted
29 Sept 2026

Usable with a quick edit9 of 9 checks passed

The memo is prioritised, evidence-based, and actionable as-is. At most one could add an explicit reprioritization trigger, but that would be a light wording edit.

Output

To: Onboarding squad Subject: Prioritize getting users to a useful first chart

Our biggest leak is before data connection: 59% of signups—7,316 people—never connect a source. Another 8 percentage points disappear before the first chart. Only 9% invite a teammate, but collaboration should come after users see value.

First-chart users have 62% week-2 retention versus 11% for everyone else. That makes time-to-first-chart our priority, although this correlation does not prove that creating a chart causes retention.

Fixes, in priority order

1. Remove the database-credentials gate. Offer a prominent sample-data path and a clearly explained CSV path. Let users request access from a data owner without abandoning onboarding. For database connections, replace the 14-page guide with source-specific instructions, inline examples, and actionable error messages.

Why first: Our audience is mostly PMs and marketers, yet onboarding assumes database expertise and access. “I don’t have the credentials” accounts for 41% of onboarding tickets. Sample data can demonstrate value immediately, but we should separately measure whether those users eventually connect their own data.

2. Make syncing resilient and asynchronous. Run syncs server-side so leaving the page does not cancel them. Show status, provide completion notifications, and support retries with clear failure explanations. Where technically feasible, sync a useful subset first rather than requiring a full load.

Why second: A wait of up to 20 minutes is already costly; requiring an open tab makes it fragile. Sync failures account for another 18% of tickets and block users who have overcome the access hurdle.

3. Deliver a chart, not another blank canvas. Recommend relevant tables rather than listing everything. Reduce twelve templates to a few guided choices, then populate a starter chart that users can edit.

Why third: Choosing a template and then seeing an empty builder breaks the promise of guidance. About 20% of connected users fail to reach a first chart. Focus on a useful result, not merely recording a chart-creation event.

4. Shorten the route to value. Move role and team-size questions until after the first chart, unless they directly personalize onboarding. Allow users to explore before email verification where security requirements permit.

Why fourth: These steps add friction before any payoff, but we lack screen-level evidence that they explain the largest losses.

5. Make invitations contextual. Prompt users to share a useful chart after success; keep invitations optional. Distinguish “ask a data owner for access” from “invite a collaborator.”

Why fifth: Low invitation rates may reflect weak activation, not a broken invitation screen.

Instrument every screen, connection method, and sync failure. Evaluate changes on useful first-chart completion, time to value, own-data activation, and week-2 retention—not just step clicks.

Grades and run details

Decision model 100 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly28%
  • passAddresses the actual decision96%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty34%
  • passAvoids unsupported claims48%
  • passProduces the required deliverable95%
  • passPrioritises by impact98%
  • passActivation defined by what predicts retention99%
  • passGets the base of every number right55%
Run
Run
#1
API response time
17 s
Submitted
29 Sept 2026

Needs more than a quick edit5 of 9 checks passed

The output presents causal interpretations as certain and omits conditions that could change the priority, so a product manager would need to add caveats and uncertainty before sending. The underlying evidence and reasoning are sound, but it requires more than just a quick edit to align with a demanding senior PM's standard for acknowledging uncertainty.

Output

MEMORANDUM

To: Onboarding Squad From: Product Team Subject: Onboarding Funnel Analysis & Prioritized Fixes

Our last 90 days show 12,400 signups, but we face a steep 59% drop-off immediately after signup (reaching only 41% at data source connection), followed by a massive drop at team invites (down to 9%). Because week-2 retention jumps from 11% to 62% for users who build a first chart, removing friction before that milestone is critical.

Here is where we are losing people and the fixes, ordered by priority:

1. The Data Connection Blockade (Fix First)

  • Where we lose people: We drop 59% of users right at the data connection step.2 Support tickets confirm why: 41% lack database credentials, and 18% experience sync failures. Forcing product managers and marketers—who are not data engineers—to input host, port, SSL modes, and select from raw database tables while blocking the UI for up to 20 minutes is a conversion killer.
  • What to fix:
  • Add a pre-loaded Sample Data / Demo Workspace option so users can experience the product immediately without connecting a live database.
  • Make data connection skippable or move it after the first chart experience.
  • Simplify the CSV upload option or provide pre-built connectors with clear, inline tooltips instead of a 14-page setup guide.

2. The Premature Team Invite Wall (Fix Second)

  • Where we lose people: Reaching a first chart drops to 33%, but inviting teammates plummets to 9%—a catastrophic 24-percentage-point drop right before value is fully realized.1 Users are being asked to advocate for a tool before they’ve even proven its worth to themselves.
  • What to fix:
  • Remove the team invite step from the core linear onboarding path entirely.
  • Relocate the invite prompt to the dashboard or settings menu once the user has successfully built and saved their first chart.

3. Redundant Survey Friction (Fix Third)

  • Where we lose people: Screens 3 ("What's your role?") and 4 ("How big is your team?") add unnecessary micro-friction right after email verification.
  • What to fix:
  • Combine these into a single, clean optional screen or infer them from workspace setup later to streamline time-to-value.

Fixing data connection roadblocks and removing premature team invites will directly protect our week-2 retention (which hinges on reaching that first chart).

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticTargeted repair

    Source: Funnel (last 90 days, 12,400 signups)Connect data source 41% → First chart 33% → Invite teammate 9%.

    What we’d changeRank the 41% to 33% drop between connecting and a first chart second. The invite comes after the first chart, so it isn't what stands between users and value.

  2. 2Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeThe 59% is lost somewhere between screens 2 and 8. Say the tickets point to the connection form, and instrument the screens to confirm it.

Grades and run details

Decision model 72 · LLM judge 6 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly20%
  • passAddresses the actual decision87%
  • passRespects explicit constraints68%
  • failIdentifies material uncertainty99%
  • partialAvoids unsupported claims38%
  • passProduces the required deliverable82%
  • passPrioritises by impact64%
  • passActivation defined by what predicts retention98%
  • passGets the base of every number right37%
Run
Run
#1
Time to output
8 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyRightRightMixed
GPT-6 Luna · API

All factual statements about the current situation are taken correctly from the brief and supplied context, with no invented facts.

GPT-6.1 Sol · API

All current-state facts and figures in the output are drawn or derived correctly from the supplied funnel, scenario, support data, and onboarding flow.

Gemini 3.5 Flash-Lite · Gemini

All factual claims about the current situation are directly from the brief or context, or derived correctly by arithmetic.

Addresses the actual decisionMixedRightMixed
GPT-6 Luna · API

The memo commits to a clear ranked order but never states what result or condition would change that order; the closing sentence is only about expanding changes.

GPT-6.1 Sol · API

It commits early to prioritising time-to-first-chart, gives a ranked fix order for the onboarding squad, and flags retention and own-data activation as conditions that would change the approach.

Gemini 3.5 Flash-Lite · Gemini

The output commits to a ranked order but does not state what result or condition would change that answer, as required by the criterion.

Identifies material uncertaintyWrongRightWrong
GPT-6 Luna · API

It does not name specific unknowns that could change the priorities or say how resolving them would alter the call.

GPT-6.1 Sol · API

It names the causal uncertainty around first-chart retention, the risk that sample-data users may not connect their own data, and the lack of screen-level evidence, with ways to resolve them.

Gemini 3.5 Flash-Lite · Gemini

The output does not identify unknowns that could change the decision (e.g., whether sample data actually drives first-chart creation) or say how they could be resolved.

Avoids unsupported claimsWrongRightWrong
GPT-6 Luna · API

It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.

GPT-6.1 Sol · API

Interpretive statements are either supported by the supplied flow or explicitly hedged as possibilities or caveats rather than established fact.

Gemini 3.5 Flash-Lite · Gemini

Causal claims like 'conversion killer', 'premature team invite wall', and 'unnecessary micro-friction' are stated as facts rather than labelled as hypotheses or inferences.

All got right 5

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

It is a memo to the onboarding squad, well under 500 words, and respects the requested form and reader.

GPT-6.1 Sol · API

The output is a memo addressed to the onboarding squad, stays under 500 words, and provides the requested ordered fixes.

Gemini 3.5 Flash-Lite · Gemini

The output is a memo to the onboarding squad, well under 500 words, and addresses the requested points.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The required memo is complete, actionable, and usable by the onboarding squad as written or with light edits.

GPT-6.1 Sol · API

The requested memo is present, complete, actionable, and usable by the onboarding squad with little or no editing.

Gemini 3.5 Flash-Lite · Gemini

The memo is complete, in the right form, under the word limit, and the squad could act on it with minimal edits.

Prioritises by impactRightRightRight
GPT-6 Luna · API

Fixes are ranked by likely impact on reaching a first chart, tied to the retention signal.

GPT-6.1 Sol · API

Fixes are ranked by likely impact on activation, starting with the data-connection bottleneck and moving to later-stage friction.

Gemini 3.5 Flash-Lite · Gemini

Fixes are ranked by likely impact on reaching the first-chart activation milestone that predicts retention.

Activation defined by what predicts retentionRightRightRight
GPT-6 Luna · API

It names first chart as the activation behaviour based on 62% vs 11% retention and ranks fixes by their effect on it.

GPT-6.1 Sol · API

It defines the activation event as reaching a useful first chart, ties it to retention, and ranks fixes by effect on that outcome rather than step completion.

Gemini 3.5 Flash-Lite · Gemini

The memo explicitly names reaching a first chart as the activation event, links it to the retention jump, and uses that to prioritize fixes.

Gets the base of every number rightRightRightRight
GPT-6 Luna · API

All percentages use the correct denominators, and the 41% ticket figure is understood as a share of onboarding tickets rather than users.

GPT-6.1 Sol · API

Percentages and differences use the correct denominators, step-to-step drops are computed from the right stages, and ambiguous bases are clarified where needed.

Gemini 3.5 Flash-Lite · Gemini

All percentages and differences are correctly calculated from the supplied data, and the base for support-ticket figures is clearly ticket counts.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 82% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT94.4100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude80.665.02None
5Gemini 3.5 Flash-LitewithGemini69.455.02None
6Sonnet 5.5withAPI77.845.02None
7Gemini 3.8 FlashwithAPI69.450.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review