Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Pollen. Our Head of Growth, Sam Okoro, has two engineers for next quarter and four ideas for how to use them. Write Sam a memo of no more than 800 words that says what our primary growth loop is, whether it's compounding, and where the two engineers should go, with how we'll know it worked. Everything we know is below.

What the model was given5 items: About Pollen, Creators, Where new creators come from (last month), Revenue, The four ideas on the table
About PollenA free form and survey builder. Every published form shows a small 'Made with Pollen: make your own' badge at the bottom. Creators can upgrade to Pro for $20 a month to remove the badge and unlock logic and integrations.
Creators18,000 creators published at least one form last month, publishing 40,000 forms between them. Each form gets 120 respondents on average. 78% of last month's active creators were active again this month.
Where new creators come from (last month)Badge: 0.9% of respondents clicked the badge and 11% of those signed up; 38% of badge sign-ups published a form within 30 days, and 85% of those were active again the next month. Template pages (written by our team, ranking in search): 6,000 sign-ups, 14% published, 71% active the next month. Paid search: 2,100 sign-ups at $38 per sign-up, 21% published, 64% active the next month.
Revenue9% of creators who publish upgrade to Pro, and Pro customers stay for 14 months on average.
The four ideas on the table1. Double the paid search budget (Finance has approved it). 2. 'Build our SEO loop': 200 more template pages. 3. A referral programme: $10 of Pro credit for each friend who signs up. 4. Replace the badge with 'Make a form like this', which opens the editor with a copy of the form the respondent just filled in. A two-week pilot on 500 forms raised badge clicks from 0.9% to 1.6% of respondents; 11% of them signed up, as before, and 52% of those published within 30 days.
What a strong answer doesThe answer key the graders mark against

Names the badge as the primary loop: creators publish forms, respondents see the badge, some become creators who publish more forms. It's chosen because badge creators publish and stay best (38% publish, 85% retained), not because it brings the most sign-ups. Sizes it: 120 × 0.9% × 11% × 38% ≈ 0.045 new publishing creators per form, about 1,800 a month from 40,000 forms. Puts that against churn: 22% of 18,000 is about 3,960 creators lost a month, while all sources together add about 3,090 (1,806 + 840 + 441), so the base is shrinking and the loop contributes rather than compounds. Shows paid isn't a loop: $38 ÷ 21% ≈ $181 per publishing creator against about $25 of revenue each (9% × $20 × 14), so doubling it burns money. Calls the template pages a channel, not a loop, since our team writes them and creators don't. Gives the referral credit little weight, because it only matters to the 9% who pay. Puts both engineers on 'Make a form like this': the pilot gives 120 × 1.6% × 11% × 52% ≈ 0.11 creators per form, about 2.4× today, which would add roughly 2,600 creators a month if it held. Flags that the pilot was small and short, and proposes a full test with a threshold (for example, publishing creators per 1,000 respondents) and a stop condition. Mentions retention as the other lever.

Critical failures (cap the score at 40)
  • Recommends doubling paid acquisition, or calls it a growth loop
Task

Mid-level · v1.0 · synthetic · B2C, PLG, viral loop, unit economics

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit6 of 11 checks passed

The test lacks a numeric threshold and measurement window, so a PM would need to add those before acting. The paid math is per signup rather than per publishing creator, but still shows the loss.

Output

To: Sam Okoro, Head of Growth Subject: Next-quarter engineering allocation

Recommendation: Put both engineers on the respondent-to-creator loop: ship and rigorously measure “Make a form like this.” Don’t double paid search or build the referral programme this quarter. Keep producing SEO templates through the content team; the available evidence doesn’t show that engineering is the constraint there.

Our primary loop is: a respondent sees a Pollen badge → visits or copies the form → signs up → publishes a form → brings in more respondents. It’s our clearest product-led acquisition loop because creators’ forms distribute the invitation to make a form.

It is not currently compounding. An average active creator publishes 2.22 forms (40,000 / 18,000), producing about 267 respondent exposures per month. Applying the current badge funnel gives roughly 0.085 next-month-active creators per active creator (267 × 0.9% × 11% × 38% × 85%). That’s well below the 1.0 needed for the loop to reproduce itself.

The pilot is promising: using its 1.6% click rate and 52% publish rate, the same calculation would imply about 0.21 next-month-active creators per creator—roughly 2.4× the current rate, if those new publishers retain at the existing 85% rate. That retention is an assumption, not a pilot result. Even so, the loop would still be subcritical: it could help offset churn, but it would not independently drive exponential growth.

Engineering plan - Engineer 1: Build the copy-to-editor flow, preserving the original form’s structure and making it easy to publish a customized copy. - Engineer 2: Instrument the full funnel and run a randomized rollout against the current badge. Track respondent exposure, click, signup, publish within 30 days, and next-month activity by variant.

How we’ll know it worked: Make next-month-active creators per 1,000 respondent exposures the primary outcome, with publishing and next-month activity as required funnel checks. Compare treatment with control, and don’t assume the pilot’s 52% publish rate will hold at scale. Expand only if the lift is sustained and new creators’ next-month activity doesn’t materially underperform the current badge cohort.

Why not the other ideas: Paid search costs $38 per signup. At its observed 21% publish rate, 9% upgrade rate, and $280 average Pro revenue lifetime, that implies about $5.29 in lifetime revenue per signup before costs—poor grounds for doubling spend. SEO already brings 6,000 signups, but only 14% publish; test additional pages with the content team before assigning engineering. The referral offer has no conversion or retention evidence, and paying credit per signup risks rewarding low-intent users.

Grades and run details

Decision model 68 · LLM judge 8 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly53%
  • passAddresses the actual decision99%
  • passRespects explicit constraints51%
  • passIdentifies material uncertainty81%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable57%
  • partialCalls out the paid maths86%
  • passPicks the lever with the most yield88%
  • passA closed loop, not a channel94%
  • failThe loop maths holds18%
  • partialProposes tests that could fail79%
Run
Run
#1
API response time
38 s
Submitted
2 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 2

Calls out the paid mathsWrong

Computes revenue per signup ($5.29) rather than cost per publishing creator ($181) and revenue per creator ($25), missing the explicit unit economics the criterion requires.

Proposes tests that could failWrong

Proposes a randomized rollout but gives no numeric threshold for 'sustained lift' or 'materially underperform', and no measurement window.

Mixed 3

Uses the supplied evidence correctlyMixed

All factual claims about the current situation are directly from the supplied context or correct arithmetic.

Picks the lever with the most yieldMixed

Chooses the copy-as-template badge and sizes its yield (0.21 per creator, 2.4×), but does not flag the pilot's small size (500 forms, two weeks) explicitly.

The loop maths holdsMixed

Correctly computes current coefficient (0.085) and pilot coefficient (0.21), applies retention, and gives plain verdicts (not compounding, subcritical).

Got right 6

Addresses the actual decisionRight

Commits early to putting both engineers on 'Make a form like this', framed for Sam, and says to expand only if lift is sustained and next-month activity doesn't underperform.

Respects explicit constraintsRight

Memo format, addressed to Sam, under 800 words, and respects the brief's request.

Identifies material uncertaintyRight

Identifies that retention from the pilot is an assumption, that the 52% publish rate may not hold at scale, and proposes a randomized rollout to resolve.

Avoids unsupported claimsRight

Assumptions are labelled (retention assumption, pilot result not guaranteed), and judgments are grounded in arithmetic.

Produces the required deliverableRight

Complete memo with recommendation, rationale, and measurement plan, usable as is.

A closed loop, not a channelRight

Names the badge as the primary closed loop, uses the highest-retention cohort's numbers (38% publish, 85% retention), and computes its coefficient with retention.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review