Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 50% were usable with at most a quick edit.

Reliably right

  1. Activation defined by what predicts retention100% pass
    Defines activation as three workouts linked to day-60 retention, and evaluates fixes by effect on that event.
    GPT-6 Astra · ChatGPT · Fitness app first week
  2. Prioritises by impact98% pass
    Ranks changes by likely impact on achieving three workouts, aligned with the identified activation threshold.
    GPT-6 Astra · ChatGPT · Fitness app first week
  3. Produces the required deliverable93% pass
    Provides a complete, actionable memo that a growth PM could implement without major gaps.
    GPT-6 Astra · ChatGPT · Fitness app first week

Where it slips

  1. Avoids unsupported claims54% pass
    It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.
    GPT-6 Luna · API · Analytics tool losing users at setup
  2. Uses the supplied evidence correctly54% pass
    It includes an unsupported motivation claim and a miscomputed '4% of trial starters' figure, so not every current-situation statement is supported by the brief or arithmetic.
    Sonnet 5.5 · API · Fitness app first week
  3. Identifies material uncertainty59% pass
    Identifies the lack of step-level data but does not state what result would change the call or how it would affect the ranking.
    Opus 5.5 · Claude · Analytics tool losing users at setup

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Review our onboarding flow and funnel below. Write a memo to the onboarding squad saying where we are losing people and what to fix, in the order you would fix it, with the reason for each. Keep it under 500 words.

What the model was given4 items: Scenario, Funnel (last 90 days, 12,400 signups), Onboarding flow, screen by screen, Support
ScenarioChartwise is a self-serve product-analytics tool. Most people who sign up are product managers or marketers, not data engineers. Pricing starts at $49 a month after a 14-day free trial.
Funnel (last 90 days, 12,400 signups)Signup 100% → Connect data source 41% → First chart 33% → Invite teammate 9%. Week-2 retention: 62% for users who reach a first chart, 11% otherwise.
Onboarding flow, screen by screen1. Sign up: email, password, company name. 2. Verify email: the flow pauses until the link is clicked. 3. “What's your role?” (seven options). 4. “How big is your team?” (five options). 5. Connect your data: choose Postgres, BigQuery, Snowflake or CSV upload. There is no sample data and no skip. 6. Connection form: host, port, database, username, password and SSL mode. The help link opens a 14-page setup guide. 7. Choose tables: a list of every table in the database. 8. Syncing: a progress bar that can take up to 20 minutes; leaving the page cancels the sync. 9. Pick a dashboard template (twelve options). 10. First chart: a chart builder with an empty canvas. 11. Invite teammates: email fields, with a “Skip for now” link.
SupportThe most common onboarding ticket (41% of them) is some version of “I don't have the database credentials.” The next is “my sync failed” (18%).
What a strong answer doesThe answer key the graders mark against

The data-source connection step (screens 5–8) is the activation bottleneck; prioritise a sample-data or CSV-first path to a first chart, then non-admin credential handoff and a sync that survives leaving the page, ahead of cosmetic fixes.

Case

Mid-level · v1.8 · anonymised real · B2B SaaS, data

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit6 of 9 checks passed

The evidence and priorities are usable, but it needs a sentence stating what would change the order and softer labelling of causal claims; no re-checking or redoing of the reasoning is needed.

Output

To: Onboarding Squad Subject: Where onboarding is leaking—and what to fix first

The biggest loss is before a data source is connected: only 41% of signups get that far. Another 8 percentage points drop before a first chart. Just 9% of signups invite a teammate. The strongest activation signal is the chart: week-2 retention is 62% for users who reach one, versus 11% for those who don’t. Prioritize getting people to a useful chart; treat invites as a later step.

Fixes, in priority order

  1. Offer a sample-data path that reaches a real first chart. Let users skip connecting their own data and explore a preloaded dataset, with a clear route to connect later. Many signups are PMs and marketers, and the most common onboarding ticket is missing database credentials (41%). A required technical connection is an especially poor first step for them. This also gives users value while they wait for access.
  1. Make connecting data simpler and recoverable. Replace or supplement the 14-page guide with short, connector-specific instructions and troubleshooting. Validate credentials early, explain the required access, and offer a clear retry path. Don’t cancel a sync when someone leaves the page; send progress and failure notices and let users resume. The credential and sync tickets point to concrete blockers, and successful connection is the gateway to the high-retention chart experience.
  1. Get users to a chart with fewer decisions. Replace the database-wide table list with suggested tables and plain-language descriptions; cut the 12 dashboard templates to a few recommended starting points. Provide a prebuilt chart or guided first-chart flow rather than an empty canvas. The 8-point drop between connection and first chart, plus the sync burden, suggests too much work remains after connection.
  1. Remove early friction that doesn’t create value. Test deferring email verification until after users enter the product, and move role/team questions until they can personalize something useful. These steps currently delay access before users see value.
  1. Make inviting a teammate contextual, not a hurdle. Keep it optional; prompt after a user has made or shared a chart, and explain the benefit. Only 9% of signups invite someone, but chart reach—not inviting—is the clearest retention signal.

Track step conversion and time-to-first-chart by path, especially sample-data versus live connection, before expanding the changes.

Grades and run details

Decision model 89 · LLM judge 7 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly33%
  • passAddresses the actual decision98%
  • passRespects explicit constraints71%
  • partialIdentifies material uncertainty65%
  • partialAvoids unsupported claims48%
  • passProduces the required deliverable97%
  • passPrioritises by impact99%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right79%
Run
Run
#1
API response time
12 s
Submitted
29 Sept 2026

Needs more than a quick edit3 of 9 checks passed

It exceeds the word limit and contains an unsupported user-capability claim. A PM would need to recheck evidence and trim before use.

Output

To: Onboarding squad Re: Where we lose people and what to fix, in order

Summary: Reaching a first chart is what predicts retention (62% at week 2 versus 11% otherwise). Only 33% of signups get there, and the biggest loss, 59% of signups, happens before they connect data. Most of our users are PMs and marketers who can't complete a database connection alone.

Fixes, in priority order

1. Add sample data and a skip option on the "Connect your data" screen. The most common ticket (41%) is "I don't have the database credentials." We ask non-engineers for host, port and SSL mode before they've seen any value. Sample data lets them reach a chart in minutes, and they can connect real data later. Caveat: sample-data users may retain worse than the 62% figure, which likely reflects real data. Track them as a separate cohort and prompt them to connect real data after their first chart.

2. Make it easy to get credentials from someone else. Add a "Send this to your engineer" option: a shareable link or prefilled email containing the connection form and a short setup checklist. Replace the 14-page guide with inline, per-field help. Promote CSV upload as the no-engineer path, since it's currently one of four equal options.

3. Fix the sync experience and its failures. Syncs run up to 20 minutes and cancel if the user leaves. Run them in the background, send an email when they finish, and let users start exploring once the first tables land. Then look at the 18% of tickets that say "my sync failed" and surface specific error messages. In the table-selection step, preselect a sensible default instead of listing every table.

4. Get people to a first chart without an empty canvas. We show twelve templates, then a blank chart builder. Cut the choice to two or three based on role, and auto-generate a first chart from the synced data (or sample data) that the user can edit. This targets the 41% → 33% drop and shortens time to value.

5. Remove friction at the front. Let users into the product before email verification, and enforce it later (e.g., before inviting teammates). Move the role question ahead of template selection, where it drives the default, and drop or defer team size unless something uses it. These are cheap fixes, but I'm ranking them below the others because I can't yet show they're a big loss.

6. Move the invite prompt to after the first chart. The 9% invite rate is 27% of users who reach a chart, so it's mostly a downstream symptom. Fix activation first. Then trigger the invite prompt right after a first chart is created, with a message like "Share this chart," rather than as a generic final step.

Measure first The funnel jumps from signup straight to data connection, so we can't see how many people drop at email verification, role, team size, or the connection form. Add per-screen events this sprint, before changing anything, so we can confirm the ordering above and measure each fix.

Success metrics: signup → first chart (33% today), share of users who connect real data within 7 days, sync failure rate, and week-2 retention split by sample-data versus real-data users.

Claims the judge couldn’t find in the brief

  • Most users are PMs and marketers who can't complete a database connection alone.

Grades and run details

Decision model 83 · LLM judge 3 of 10 checks
Decision model checks
  • failUses the supplied evidence correctly7%
  • passAddresses the actual decision99%
  • passRespects explicit constraints24%
  • passIdentifies material uncertainty90%
  • partialAvoids unsupported claims43%
  • passProduces the required deliverable93%
  • passPrioritises by impact95%
  • passActivation defined by what predicts retention100%
  • passGets the base of every number right42%
Run
Run
#1
API response time
15 s
Submitted
29 Sept 2026

Usable with a quick edit8 of 9 checks passed

The memo is a clear, prioritized, evidence-based list with a straightforward measurement plan; a PM could use it with minimal edits.

Output

To: Onboarding squad Subject: Fix access and sync before optimising chart creation

Our biggest measured loss is before data connection: 59% of signups, roughly 7,300 people, never connect a source. We ask PMs and marketers to complete a database setup that many cannot do themselves, then make a lengthy sync depend on keeping a page open.

Another 8 percentage points disappear between connection and first chart: about 20% of those who connect. Teammate invitations fall to 9% of signups, but we should establish whether collaboration is necessary for value before treating that as our next bottleneck.

First-chart users have much higher week-2 retention, 62% versus 11%. That makes chart creation a promising activation milestone, but does not prove that pushing everyone through a chart will cause better retention.

Fix in this order

1. Give users a route forward without database credentials. Offer a clearly labelled sample-data experience and make CSV upload a prominent alternative.1 Add a setup hand-off that lets users ask a colleague to connect the source, then resume when it is ready. Replace the generic setup guide with short, source-specific instructions.

This addresses the largest measured loss and the strongest support signal: missing credentials account for 41% of onboarding tickets. That is a share of tickets, not signups, so we cannot yet quantify how much drop-off it explains. Sample data should demonstrate value while preserving a clear path to the user’s own data.

2. Make syncing durable and recoverable. Run syncs in the background so closing the page does not cancel them. Show meaningful status, explain failures and allow retries without repeating setup. Notify users when their data is ready.

A 20-minute foreground wait is a fragile requirement. Sync failures generate 18% of onboarding tickets and can waste the effort users have already invested.

3. Turn connected data into a useful first chart. Help users identify relevant tables instead of presenting every table equally. Use their chosen template to produce an editable chart from compatible data, with guided choices where mapping is needed. Do not finish template selection with an empty canvas.

This targets the roughly one in five connected users who never reach a chart. Measure whether they produce something useful, not merely whether a chart renders.

4. Remove avoidable preliminary work. Defer role and team-size questions unless their answers immediately improve setup. Review whether email verification must block exploration, retaining verification where required for account security.

These steps add friction before value, but we lack evidence that they drive the main loss.

Measure alongside delivery: instrument each screen, connection method, error and return visit. Test changes against first charts using customers’ own data, week-2 retention and trial-to-paid conversion. Track sample-data charts separately. Revisit invitations once we understand whether users need collaboration to succeed. To: Onboarding squad Subject: Fix access and sync before optimising chart creation

Our biggest measured loss is before data connection: 59% of signups, roughly 7,300 people, never connect a source. We ask PMs and marketers to complete a database setup that many cannot do themselves, then make a lengthy sync depend on keeping a page open.

Another 8 percentage points disappear between connection and first chart: about 20% of those who connect. Teammate invitations fall to 9% of signups, but we should establish whether collaboration is necessary for value before treating that as our next bottleneck.

First-chart users have much higher week-2 retention, 62% versus 11%. That makes chart creation a promising activation milestone, but does not prove that pushing everyone through a chart will cause better retention.

Fix in this order

1. Give users a route forward without database credentials. Offer a clearly labelled sample-data experience and make CSV upload a prominent alternative. Add a setup hand-off that lets users ask a colleague to connect the source, then resume when it is ready. Replace the generic setup guide with short, source-specific instructions.

This addresses the largest measured loss and the strongest support signal: missing credentials account for 41% of onboarding tickets. That is a share of tickets, not signups, so we cannot yet quantify how much drop-off it explains. Sample data should demonstrate value while preserving a clear path to the user’s own data.

2. Make syncing durable and recoverable. Run syncs in the background so closing the page does not cancel them. Show meaningful status, explain failures and allow retries without repeating setup. Notify users when their data is ready.

A 20-minute foreground wait is a fragile requirement. Sync failures generate 18% of onboarding tickets and can waste the effort users have already invested.

3. Turn connected data into a useful first chart. Help users identify relevant tables instead of presenting every table equally. Use their chosen template to produce an editable chart from compatible data, with guided choices where mapping is needed. Do not finish template selection with an empty canvas.

This targets the roughly one in five connected users who never reach a chart. Measure whether they produce something useful, not merely whether a chart renders.

4. Remove avoidable preliminary work. Defer role and team-size questions unless their answers immediately improve setup. Review whether email verification must block exploration, retaining verification where required for account security.

These steps add friction before value, but we lack evidence that they drive the main loss.

Measure alongside delivery: instrument each screen, connection method, error and return visit. Test changes against first charts using customers’ own data, week-2 retention and trial-to-paid conversion. Track sample-data charts separately. Revisit invitations once we understand whether users need collaboration to succeed.

What a PM had to fix

  1. 1Decision deferredMake the callQuick edit

    What we’d changeRank the four parts of fix 1 so the squad knows what ships first, and make the calls fix 4 leaves open.

Grades and run details

Decision model 89 · LLM judge 10 of 10 checks
Decision model checks
  • passUses the supplied evidence correctly22%
  • passAddresses the actual decision98%
  • failRespects explicit constraints40%
  • passIdentifies material uncertainty87%
  • passAvoids unsupported claims26%
  • passProduces the required deliverable59%
  • passPrioritises by impact97%
  • passActivation defined by what predicts retention93%
  • passGets the base of every number right66%
Run
Run
#1
Time to output
40 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyRightWrongRight
GPT-6 Luna · API

All factual statements about the current situation are taken correctly from the brief and supplied context, with no invented facts.

Sonnet 5.5 · API

It presents as fact that PMs and marketers cannot complete a database connection alone, which the supplied context does not establish.

GPT-6 Astra · ChatGPT

All statements about the current situation are taken correctly from the brief or derived arithmetically; nothing is invented.

Addresses the actual decisionMixedMixedRight
GPT-6 Luna · API

The memo commits to a clear ranked order but never states what result or condition would change that order; the closing sentence is only about expanding changes.

Sonnet 5.5 · API

It commits to a ranked order but does not state a result or condition that would change the answer, only that per-screen events would confirm it.

GPT-6 Astra · ChatGPT

Commits early to a ranked order of fixes, addressed to the onboarding squad, and says team invitations will be revisited after understanding if collaboration is needed.

Respects explicit constraintsRightMixedMixed
GPT-6 Luna · API

It is a memo to the onboarding squad, well under 500 words, and respects the requested form and reader.

Sonnet 5.5 · API

The memo is approximately 541 words, exceeding the explicit 500-word limit.

GPT-6 Astra · ChatGPT

Memo format, under 500 words, directed to the onboarding squad, respecting all explicit constraints.

Identifies material uncertaintyWrongMixedRight
GPT-6 Luna · API

It does not name specific unknowns that could change the priorities or say how resolving them would alter the call.

Sonnet 5.5 · API

It names sample-data retention and missing funnel events but does not specify what data result would change the prioritisation.

GPT-6 Astra · ChatGPT

Identifies that first-chart correlation does not prove causation, ticket share is not user drop-off, and role/team questions lack evidence; proposes measurement to resolve.

Avoids unsupported claimsWrongWrongRight
GPT-6 Luna · API

It presents causal or evaluative claims such as 'successful connection is the gateway' and 'a required technical connection is an especially poor first step' as established fact rather than hypothesis.

Sonnet 5.5 · API

The user-capability claim and 'mostly downstream symptom' are stated as established fact rather than labelled interpretations.

GPT-6 Astra · ChatGPT

Interpretations are grounded in the data and qualified where evidence is suggestive, so no factual claims are presented as established without support.

Produces the required deliverableRightMixedRight
GPT-6 Luna · API

The required memo is complete, actionable, and usable by the onboarding squad as written or with light edits.

Sonnet 5.5 · API

It is a usable memo in form and audience, but it does not meet the required length constraint.

GPT-6 Astra · ChatGPT

Delivers a complete memo with prioritized fixes, reasons, and measurement plan that the squad can act on with light edits.

All got right 3

Prioritises by impactRightRightRight
GPT-6 Luna · API

Fixes are ranked by likely impact on reaching a first chart, tied to the retention signal.

Sonnet 5.5 · API

Fixes are explicitly ranked by likely impact on reaching first chart, with reasons for each.

GPT-6 Astra · ChatGPT

Fixes are ranked by likely impact on activation: credential-free path first (largest drop), then sync durability, then first-chart creation, then friction removal.

Activation defined by what predicts retentionRightRightRight
GPT-6 Luna · API

It names first chart as the activation behaviour based on 62% vs 11% retention and ranks fixes by their effect on it.

Sonnet 5.5 · API

It defines first chart as the activation behaviour using the retention split and ranks fixes by effect on that event.

GPT-6 Astra · ChatGPT

Activation is defined as first chart, linked to week-2 retention (62% vs 11%), and fixes are judged by their effect on reaching that milestone.

Gets the base of every number rightRightRightRight
GPT-6 Luna · API

All percentages use the correct denominators, and the 41% ticket figure is understood as a share of onboarding tickets rather than users.

Sonnet 5.5 · API

Derived percentages use the correct denominators, including 59% before data, 27% invite of first-chart users, and the 41% to 33% step.

GPT-6 Astra · ChatGPT

All percentages are computed from the correct funnel steps, and ticket shares are explicitly noted as share of tickets, not users.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 82% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.0100.02None
2GPT-6 AstrawithChatGPT94.4100.02None
3GPT-6 LunawithAPI94.485.02None
4Opus 5.5withClaude80.665.02None
5Gemini 3.5 Flash-LitewithGemini69.455.02None
6Sonnet 5.5withAPI77.845.02None
7Gemini 3.8 FlashwithAPI69.450.02None

About the task

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Defines activation as the behaviour that predicts retention, not finishing onboarding

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Decision model and LLM judge, calibrated against a blind PM review