Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Pollen. Our Head of Growth, Sam Okoro, has two engineers for next quarter and four ideas for how to use them. Write Sam a memo of no more than 800 words that says what our primary growth loop is, whether it's compounding, and where the two engineers should go, with how we'll know it worked. Everything we know is below.

What the model was given5 items: About Pollen, Creators, Where new creators come from (last month), Revenue, The four ideas on the table
About PollenA free form and survey builder. Every published form shows a small 'Made with Pollen: make your own' badge at the bottom. Creators can upgrade to Pro for $20 a month to remove the badge and unlock logic and integrations.
Creators18,000 creators published at least one form last month, publishing 40,000 forms between them. Each form gets 120 respondents on average. 78% of last month's active creators were active again this month.
Where new creators come from (last month)Badge: 0.9% of respondents clicked the badge and 11% of those signed up; 38% of badge sign-ups published a form within 30 days, and 85% of those were active again the next month. Template pages (written by our team, ranking in search): 6,000 sign-ups, 14% published, 71% active the next month. Paid search: 2,100 sign-ups at $38 per sign-up, 21% published, 64% active the next month.
Revenue9% of creators who publish upgrade to Pro, and Pro customers stay for 14 months on average.
The four ideas on the table1. Double the paid search budget (Finance has approved it). 2. 'Build our SEO loop': 200 more template pages. 3. A referral programme: $10 of Pro credit for each friend who signs up. 4. Replace the badge with 'Make a form like this', which opens the editor with a copy of the form the respondent just filled in. A two-week pilot on 500 forms raised badge clicks from 0.9% to 1.6% of respondents; 11% of them signed up, as before, and 52% of those published within 30 days.
What a strong answer doesThe answer key the graders mark against

Names the badge as the primary loop: creators publish forms, respondents see the badge, some become creators who publish more forms. It's chosen because badge creators publish and stay best (38% publish, 85% retained), not because it brings the most sign-ups. Sizes it: 120 × 0.9% × 11% × 38% ≈ 0.045 new publishing creators per form, about 1,800 a month from 40,000 forms. Puts that against churn: 22% of 18,000 is about 3,960 creators lost a month, while all sources together add about 3,090 (1,806 + 840 + 441), so the base is shrinking and the loop contributes rather than compounds. Shows paid isn't a loop: $38 ÷ 21% ≈ $181 per publishing creator against about $25 of revenue each (9% × $20 × 14), so doubling it burns money. Calls the template pages a channel, not a loop, since our team writes them and creators don't. Gives the referral credit little weight, because it only matters to the 9% who pay. Puts both engineers on 'Make a form like this': the pilot gives 120 × 1.6% × 11% × 52% ≈ 0.11 creators per form, about 2.4× today, which would add roughly 2,600 creators a month if it held. Flags that the pilot was small and short, and proposes a full test with a threshold (for example, publishing creators per 1,000 respondents) and a stop condition. Mentions retention as the other lever.

Critical failures (cap the score at 40)
  • Recommends doubling paid acquisition, or calls it a growth loop
Task

Mid-level · v1.0 · synthetic · B2C, PLG, viral loop, unit economics

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit9 of 11 checks passed

The memo contains an unsupported equilibrium number (10,700) and an unsupported claim about rising cost per sign-up, so a PM would need to re-check those parts before using it.

Output

To: Sam Okoro, Head of Growth From: [PM], Growth Re: Our growth loop, and where the two engineers should go next quarter

Summary: Our primary loop is the badge. Today it does not compound, and the active creator base is shrinking. The "Make a form like this" pilot is the only one of the four ideas that could make it compound. Both engineers should ship it properly, with a holdout. We should not double paid search.

1. The loop

Creators publish forms → respondents see the badge → some respondents become creators → they publish forms.

Last month the loop worked like this: - 40,000 forms × 120 respondents = 4.8M respondents - 0.9% clicked the badge = 43,200 clicks - 11% of clicks signed up = 4,752 sign-ups - 38% of sign-ups published = 1,806 new publishing creators

The badge is our largest source of creators. Templates produced 840 publishers and paid search 441. Badge creators also retain best: 85% are active the next month, against 71% for templates and 64% for paid.

2. Is it compounding? Not yet.

Each active creator produces about 267 respondents a month, which yields 0.10 new publishers per creator per month. With 78% monthly retention, a creator stays about 4.5 months. Over that lifetime, each creator brings in about 0.46 new creators.

A loop compounds only when this number exceeds 1. At 0.46, the badge amplifies our other channels by about 1.8×, but it cannot sustain itself.

The overall picture is worse: - We keep 14,040 of last month's 18,000 active creators. - We add about 3,100 new publishers. - That gives about 17,100 active creators, a 5% monthly decline. - At current rates the base settles at roughly 10,700 creators.

This assumes the 78% retention figure covers the whole base, which is how I read the data.

3. The four ideas

Double paid search: no. - Each sign-up costs $38. Only 21% publish, so each publishing creator costs about $181. - Each publisher is worth about $25 in revenue (9% upgrade × $20 × 14 months). - Adding the loop's downstream creators raises that to only about $46. - Doubling spend would add about 440 publishers a month for an extra $80k a month, and cost per sign-up usually rises with spend. - This idea doesn't need engineers. I'd ask Finance to hold the budget rather than spend it because it's approved.

200 more template pages: not now. - This is not a loop, because new creators don't create more pages. Our team writes them, so output grows only as fast as we write. - We don't know how many pages produce today's 6,000 sign-ups, so we can't estimate the yield. - Templates publish at only 14%. - This is mainly content work, not engineering work.

Referral credit: no. - It rewards sign-ups, not publishing, which invites low-quality and fake sign-ups. - $10 of Pro credit means nothing to the 91% of creators who never upgrade.

"Make a form like this": yes. The pilot changed every stage of the loop:

Current badgePilot
Click rate0.9%1.6%
Sign-up rate11%11%
Publish rate38%52%
New publishers per 1,000 respondents0.380.92

At network scale, that would mean about 4,400 badge publishers a month instead of 1,800. Each creator's lifetime yield would rise from 0.46 to about 1.1, so the loop would start to compound. Next month's active base would be about 19,700 instead of 17,100.

The margin is thin. The loop reaches the break-even of 1.0 at 0.83 publishers per 1,000 respondents. The pilot beat that by only about 10%. We also need to treat the pilot result with caution: - It covered 500 forms, which may not be representative. - It ran for two weeks, so check whether the 30-day publish window had fully elapsed. - We have no next-month retention data for these creators. People who copy a form may be less committed than people who build one.

4. What the engineers build

1. Staged rollout with a 20% holdout of forms on the old badge, randomised by form and run for at least 8 weeks so we see 30-day publishing and next-month retention. 2. Creator controls and safety: - Copies take the form structure only, never responses. - Creators can turn off copying for sensitive forms. 3. Instrumentation for each stage: click, sign-up, publish, next-month activity, and Pro upgrade. Results should be split by form type so we learn which forms spread best.

5. How we'll know it worked

Primary metric: new publishing creators per 1,000 respondents, treatment vs holdout. - Ship to 100% if it stays at or above 0.83, the level needed for the loop to compound. - Below about 0.6, the change is still positive but the loop won't compound. We would keep it and look for other levers.

Quality checks: - Next-month retention of copy sign-ups should be at least 70%. If it falls much lower, the extra volume won't last. - The upgrade rate of these creators to Pro should be close to 9%.

Guardrails: - Form completion rate is unchanged. - Creator complaints and copy opt-outs are low. - Pro upgrade rate across the base is unchanged.

Business outcome: the active creator base grows month on month by the end of the quarter, instead of falling about 5%.

Bottom line: paid search and templates rent growth one sign-up at a time. The copy-this-form change is our one realistic chance to make growth compound, and a proper rollout will tell us within a quarter whether it does.

Claims the judge couldn’t find in the brief

  • At current rates the base settles at roughly 10,700 creators.
  • Cost per sign-up usually rises with spend.

Grades and run details

Decision model 91 · LLM judge 9 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly27%
  • passAddresses the actual decision100%
  • passRespects explicit constraints33%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims44%
  • passProduces the required deliverable88%
  • passCalls out the paid maths99%
  • passPicks the lever with the most yield98%
  • passA closed loop, not a channel100%
  • passThe loop maths holds11%
  • partialProposes tests that could fail52%
Run
Run
#1
Time to output
55 s
Submitted
2 Oct 2026

Needs more than a quick edit10 of 11 checks passed

The memo omits a numeric success threshold for the experiment, so a PM would need to define one before acting; otherwise the reasoning is sound and the evidence is correct.

Output

To: Sam Okoro, Head of Growth From: Product Subject: Put both engineers on the respondent-to-creator loop

Recommendation

Use both engineers to ship and validate “Make a form like this.” It strengthens our existing growth loop, has encouraging experimental evidence, and attracts creators who historically activate and retain better than creators from other channels. Don’t double paid search or build a referral programme next quarter. Keep existing SEO acquisition running, but don’t call 200 team-written pages a growth loop.

Our primary loop—and whether it compounds

Our loop is:

Creator publishes → respondents encounter Pollen → some become creators → those creators publish → more respondents encounter Pollen.

The badge connects product usage to new acquisition. SEO and paid search feed creators into that loop; they are not themselves loops. Team-written template pages don’t automatically generate more template pages.

At last month’s volume, 40,000 forms × 120 respondents implies 4.8 million respondent encounters. Applying the badge funnel gives:

Monthly outputCurrent badgeProposed badge*
------:---:
Clicks43,20076,800
Sign-ups4,7528,448
New publishing creators1,8064,393
Active again the following month1,5353,734

Assumes pilot conversion rates hold at scale and new creators retain at the historical badge cohort’s 85% rate. These are illustrative ceilings if some of the 4.8 million encounters are on Pro forms without badges; actual forecasts must use badge-visible encounters.

The loop recurs, but we have not demonstrated self-sustaining compounding. Each existing active creator generates approximately 0.10 new publishing creators per cycle under the current funnel, or 0.085 who remain active the following month. Meanwhile, 22% of existing active creators lapse—about 3,960 of an 18,000-creator base.

The proposed experience raises those figures to approximately 0.244 new publishers and 0.207 subsequently retained creators per existing creator. That approaches churn replacement, but does not establish durable compounding. Activation and retention occur over different windows, so this is a replacement diagnostic, not an exact monthly growth forecast. Pro upgrades also remove badges, reducing future distribution.

Why this beats the alternatives

  • Paid search is economically unattractive on the supplied numbers. Its $38 sign-up cost becomes approximately $181 per publishing creator. Expected lifetime revenue per publishing creator is only $25.20: 9% upgrade × $20 × 14 months, before costs. Finance approval doesn’t fix that gap.
  • SEO produces useful acquisition, but the marginal return is unknown. Existing pages generated 840 publishers, versus roughly 1,806 from the badge. We don’t know what 200 additional pages would deliver, and their production remains dependent on our team.
  • Referral credit rewards the wrong milestone. Paying for sign-ups invites low-intent referrals and abuse. We have no evidence of referral demand or incremental retained creators.
  • The badge pilot improves two stages: clicks rise from 0.9% to 1.6%, and publishing from 38% to 52%. Together, that is 2.43× more publishers per respondent, not merely better click-through.

Engineering allocation and success criteria

Assign one engineer to the copy-to-editor experience and one to experimentation, instrumentation, and rollout safeguards, working as a single team. Copy only reusable form structure—not responses or sensitive data—and provide appropriate creator controls.

Run a persistent randomized experiment, assigned by creator to avoid mixing experiences across their forms. Measure actual badge-visible encounters, not total respondents.

Primary metric: incremental next-month-active creators per 100,000 badge-visible encounters. The baseline is approximately 32; the pilot implies 78, if retention holds. Track clicks, sign-ups, and 30-day publishing as diagnostic metrics, not the final win.

Wait for the activation and subsequent retention windows to mature. Check respondent completion, creator complaints, copied-form publishing quality, Pro conversion, and badge exposure lost through upgrades. Attribute downstream acquisition consistently and test for cannibalization of other channels.

Success means a statistically credible increase in retained creators without material guardrail deterioration. At full illustrative exposure, the opportunity is roughly 2,200 additional retained creators per monthly cohort. Separately, claim compounding only when observed cohort reproduction—including repeat publishing, churn, and badge removal—demonstrably exceeds replacement.

Grades and run details

Decision model 91 · LLM judge 10 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly35%
  • passAddresses the actual decision99%
  • passRespects explicit constraints38%
  • passIdentifies material uncertainty90%
  • passAvoids unsupported claims39%
  • passProduces the required deliverable76%
  • passCalls out the paid maths100%
  • passPicks the lever with the most yield95%
  • passA closed loop, not a channel100%
  • partialThe loop maths holds15%
  • partialProposes tests that could fail78%
Run
Run
#1
API response time
54 s
Submitted
2 Oct 2026

1 of 11 checks passed

Output

MEMORANDUM

TO: Sam Okoro, Head of Growth FROM: Product Manager, Pollen DATE: March 30, 2026 SUBJECT: Growth Analysis & Q3 Engineering Allocation

---

1. Our Primary Growth Loop and Compounding Status

Pollen’s primary growth loop is the Viral Product-Led Growth (PLG) Badge Loop.

Here is how it functions: 1. Creation: A creator builds and publishes a form on Pollen. 2. Distribution: Creators send their forms out, generating an average of 120 respondents per form (40,000 forms $\times$ 120 = 4.8 million monthly respondent impressions). 3. Acquisition: 0.9% of those respondents click the "Made with Pollen" badge, and 11% of those clickers sign up. This currently yields approximately 4,752 sign-ups per month at zero media cost. 4. Activation & Retention: 38% of those sign-ups publish a form within 30 days, and 85% of those active creators retain into the following month, immediately re-feeding the top of the loop.

Is it compounding? Yes, it is structurally a compounding loop, but its current compounding velocity is fragile. Because our monthly creator retention is strong (78% overall, and 85% for badge cohorts), every cohort of users added to the system continually outputs thousands of new forms, generating millions of new respondent impressions month after month.

However, the loop’s current conversion efficiency means it is barely breaking even with natural churn. To accelerate compounding, we need to increase the velocity of respondents turning into active creators. That brings us to our Q2 engineering allocation.

---

2. Where the Two Engineers Should Go

We have four ideas on the table. We should allocate both engineers to Idea 4: Replace the badge with 'Make a form like this' (Cloning Loop).

Here is why we must reject the other three: * Idea 1 (Double Paid Search): Generates 2,100 sign-ups at $38/ea ($79.8k/month additional spend). While profitable over a 14-month Pro lifespan (9% conversion $\times$ $20/mo $\times$ 14 mos = $25.20 LTV vs. $38 CAC—wait, our LTV is actually around $252, making CAC look fine on paper), paid search yields lower activation (21%) and worse retention (64%) than organic loops. It is linear growth, not compounding. * Idea 2 (200 SEO Template Pages): Good, but it relies on static editorial content. It lacks a viral coefficient and has a lower activation rate (14%) and month-2 retention (71%) than the badge loop. * Idea 3 (Referral Programme): A $10 Pro credit is structurally misaligned for a product where 91% of our base is on a free tier. Free users cannot easily utilize Pro credits, rendering this incentive largely inert for the vast majority of our creators.

Why Idea 4 is the clear winner: The pilot data for "Make a form like this" is exceptional. By reducing friction and showing respondents a working replica of the exact form they just successfully interacted with, we saw: * Badge clicks nearly doubled: From 0.9% to 1.6%. * Sign-up rate held steady: 11% conversion. * Activation surged: 52% of sign-ups published within 30 days (up from 38% on the old badge).

This isn't just an incremental tweak; it transforms our passive watermark into an active, high-intent replication engine. It directly attacks the weakest link in our primary loop (activation) while supercharging top-of-funnel discovery.

---

3. How We Will Know It Worked (Success Metrics)

We will roll this out to 100% of published forms on day one of the quarter. We will measure success across three core KPIs over a 60-day post-launch window:

  1. Primary Metric (Loop Velocity): Monthly Free-to-Published Creator conversion from badge traffic. Target: A sustained lift in 30-day activation from the baseline 38% toward the pilot’s 52% benchmark.
  2. Secondary Metric (Top-of-Funnel Conversion): Respondent-to-sign-up volume. Target: Maintain or exceed the pilot's 1.6% click-through rate across our full 4.8M monthly respondent pool, yielding ~8,400 sign-ups/month from the badge alone (nearly double current output).
  3. Downstream Revenue Impact: Monthly Recurring Revenue (MRR) added via Pro upgrades from this specific acquisition channel. Target: Since 9% of publishing creators upgrade to Pro ($20/mo), the uplift in activated creators from this loop must translate directly to an accelerated net-new Pro ARR growth rate compared to Q1.

Put both engineers on Idea 4. Let's compound the loop.

Grades and run details

Decision model 32
Decision model checks
  • failUses the supplied evidence correctly80%
  • partialAddresses the actual decision36%
  • failRespects explicit constraints81%
  • failIdentifies material uncertainty44%
  • failAvoids unsupported claims95%
  • partialProduces the required deliverable70%
  • partialCalls out the paid maths54%
  • partialPicks the lever with the most yield99%
  • passA closed loop, not a channel25%
  • failThe loop maths holds95%
  • partialProposes tests that could fail53%
Run
Run
#1
Time to output
6 s
Submitted
2 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 10

Uses the supplied evidence correctlyMixedRightWrong
Opus 5.5 · Claude

The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.

GPT-6.1 Sol · API

All claims about the current situation are taken directly from the supplied context or derived by correct arithmetic.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Addresses the actual decisionRightRightMixed
Opus 5.5 · Claude

Commits early to putting both engineers on 'Make a form like this' with a holdout, says not to double paid search, and specifies what metric would change the call.

GPT-6.1 Sol · API

The memo commits early to putting both engineers on 'Make a form like this', addresses Sam directly, and states that success means a statistically credible increase in retained creators without guardrail deterioration.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Respects explicit constraintsRightRightWrong
Opus 5.5 · Claude

Memo is addressed to Sam Okoro, stays within ~800 words, and proposes enforcement mechanisms (holdout, controls).

GPT-6.1 Sol · API

The output is a memo to Sam, under 800 words, and respects the brief's request for a recommendation and success criteria.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Identifies material uncertaintyRightRightWrong
Opus 5.5 · Claude

Flags pilot's small size, two-week window, missing retention data, and representativeness, and resolves them with an 8-week holdout test and thresholds.

GPT-6.1 Sol · API

It flags the small pilot, the assumption that retention holds, the need to measure badge-visible encounters, and says to wait for retention windows to mature.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Avoids unsupported claimsWrongRightWrong
Opus 5.5 · Claude

Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.

GPT-6.1 Sol · API

Forecasts and assumptions are clearly labelled, and confident claims are backed by the supplied evidence.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Produces the required deliverableRightRightMixed
Opus 5.5 · Claude

Delivers a complete memo to Sam with a clear recommendation, how to measure success, and is within the word limit.

GPT-6.1 Sol · API

The memo is complete, in the right form, within the word limit, and a PM could act on it with light edits.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Calls out the paid mathsRightRightMixed
Opus 5.5 · Claude

Calculates $181 cost per publishing creator and $25 revenue, states doubling would lose money, and recommends against it.

GPT-6.1 Sol · API

It calculates $181 per publishing creator and $25.20 revenue, and says plainly that doubling paid search would be economically unattractive.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Picks the lever with the most yieldRightRightMixed
Opus 5.5 · Claude

Chooses the copy-as-template badge, sizes its yield from the pilot (0.92 per 1,000 respondents, ~4,400 publishers, ~2.4× current), and notes the pilot's small size.

GPT-6.1 Sol · API

It chooses the copy-as-template badge, sizes its yield at about 0.11 creators per form (2.43× today), and notes the pilot was small and short.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

The loop maths holdsRightRightWrong
Opus 5.5 · Claude

Correctly computes yield per creator (0.10/month, 0.46 lifetime), retention (4.5 months), amplification (1.8×), and break-even (0.83), and says the loop does not compound.

GPT-6.1 Sol · API

The yield (0.10 publishers, 0.085 retained) is correct, retention is applied, and it gives a plain verdict that the loop does not compound.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Proposes tests that could failRightWrongMixed
Opus 5.5 · Claude

Proposes a holdout test with a numeric threshold (0.83 to ship, below 0.6 to keep but seek other levers), an 8-week window, and quality checks with thresholds.

GPT-6.1 Sol · API

The proposed experiment lacks a numeric threshold for success; it only says 'statistically credible increase' without a specific number, so it does not meet the requirement for a threshold.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

All got right 1

A closed loop, not a channelRightRightRight
Opus 5.5 · Claude

Names the badge as the primary closed loop, grounds it in highest-retention creators, and sizes its yield with retention applied.

GPT-6.1 Sol · API

It names the badge-driven respondent-to-creator loop as the primary loop, grounds it in the highest-retention creators, and computes its yield with retention.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review