Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Pollen. Our Head of Growth, Sam Okoro, has two engineers for next quarter and four ideas for how to use them. Write Sam a memo of no more than 800 words that says what our primary growth loop is, whether it's compounding, and where the two engineers should go, with how we'll know it worked. Everything we know is below.

What the model was given5 items: About Pollen, Creators, Where new creators come from (last month), Revenue, The four ideas on the table
About PollenA free form and survey builder. Every published form shows a small 'Made with Pollen: make your own' badge at the bottom. Creators can upgrade to Pro for $20 a month to remove the badge and unlock logic and integrations.
Creators18,000 creators published at least one form last month, publishing 40,000 forms between them. Each form gets 120 respondents on average. 78% of last month's active creators were active again this month.
Where new creators come from (last month)Badge: 0.9% of respondents clicked the badge and 11% of those signed up; 38% of badge sign-ups published a form within 30 days, and 85% of those were active again the next month. Template pages (written by our team, ranking in search): 6,000 sign-ups, 14% published, 71% active the next month. Paid search: 2,100 sign-ups at $38 per sign-up, 21% published, 64% active the next month.
Revenue9% of creators who publish upgrade to Pro, and Pro customers stay for 14 months on average.
The four ideas on the table1. Double the paid search budget (Finance has approved it). 2. 'Build our SEO loop': 200 more template pages. 3. A referral programme: $10 of Pro credit for each friend who signs up. 4. Replace the badge with 'Make a form like this', which opens the editor with a copy of the form the respondent just filled in. A two-week pilot on 500 forms raised badge clicks from 0.9% to 1.6% of respondents; 11% of them signed up, as before, and 52% of those published within 30 days.
What a strong answer doesThe answer key the graders mark against

Names the badge as the primary loop: creators publish forms, respondents see the badge, some become creators who publish more forms. It's chosen because badge creators publish and stay best (38% publish, 85% retained), not because it brings the most sign-ups. Sizes it: 120 × 0.9% × 11% × 38% ≈ 0.045 new publishing creators per form, about 1,800 a month from 40,000 forms. Puts that against churn: 22% of 18,000 is about 3,960 creators lost a month, while all sources together add about 3,090 (1,806 + 840 + 441), so the base is shrinking and the loop contributes rather than compounds. Shows paid isn't a loop: $38 ÷ 21% ≈ $181 per publishing creator against about $25 of revenue each (9% × $20 × 14), so doubling it burns money. Calls the template pages a channel, not a loop, since our team writes them and creators don't. Gives the referral credit little weight, because it only matters to the 9% who pay. Puts both engineers on 'Make a form like this': the pilot gives 120 × 1.6% × 11% × 52% ≈ 0.11 creators per form, about 2.4× today, which would add roughly 2,600 creators a month if it held. Flags that the pilot was small and short, and proposes a full test with a threshold (for example, publishing creators per 1,000 respondents) and a stop condition. Mentions retention as the other lever.

Critical failures (cap the score at 40)
  • Recommends doubling paid acquisition, or calls it a growth loop
Task

Mid-level · v1.0 · synthetic · B2C, PLG, viral loop, unit economics

The outputs

7 setups answered this brief. Pick up to three to read side by side.

3 of 11 checks passed

Output

MEMO

To: Sam Okoro, Head of Growth From: Product Management Date: Current Subject: Engineering Allocation for Next Quarter: Doubling Down on Our Primary Growth Loop

---

Executive Summary

Our primary growth engine is our product-led viral loop (forms $\to$ respondents $\to$ badge clicks $\to$ creators $\to$ new forms). While powerful, it is currently amplifying, not self-compounding ($K < 1.0$).

Both engineers should be dedicated to Idea 4 (“Make a form like this”). The pilot proved a 143% lift in publishing creators from our highest-volume, highest-retention channel. Conversely, Paid Search is deeply unprofitable, SEO templates are editorial rather than engineering work, and referrals solve an acquisition problem we can address organically for free.

---

1. Our Primary Growth Loop & Why It’s Not Yet Compounding

Our primary growth loop is driven by the form badge:

$$\text{Active Creators} \longrightarrow \text{Published Forms} \longrightarrow \text{Respondents} \longrightarrow \text{Badge Clicks} \longrightarrow \text{Sign-ups} \longrightarrow \text{Publishing Creators}$$

Current Monthly Metrics (Badge Loop): * 4.8M respondents ($40\text{k forms} \times 120\text{ respondents}$) * 43,200 clicks ($0.9\%$) $\to$ 4,752 sign-ups ($11\%$) $\to$ 1,806 new publishing creators ($38\%$) * Retention: Badge creators have our highest month-two retention (85%, vs. 71% for SEO and 64% for Paid).

Is it compounding? No—not in the runaway sense. * Each active creator publishes $\approx 2.22$ forms, reaching 267 respondents. * At baseline conversion rates, one publishing creator generates only 0.10 new publishing creators per cycle ($267 \times 0.009 \times 0.11 \times 0.38$). Even accounting for creator lifetime retention, the viral factor ($K$) remains well below $1.0$. * The loop is an amplifier of other top-of-funnel acquisition, but it cannot sustain compounding growth on its own without steady top-of-funnel inputs. To move closer to a true compounding state, we must dramatically increase loop velocity and cycle conversion.

---

2. Evaluation of the Four Ideas

IdeaStrategic FitUnit Economics / ImpactRecommendation
:---:---:---:---
1. Double Paid SearchPoorSeverely negative LTV:CAC. Pro LTV is \$280 ($20 \times 14\text{ mo}$). Value per sign-up is \$5.29 ($21\% \text{ publish} \times 9\% \text{ upgrade} \times \$280$). We pay \$38.00 per sign-up, losing \$32.71 on every acquisition.Kill immediately. No eng needed.
2. 200 Template PagesModerate14% publish rate; content-writing task, not an engineering initiative. Programmatic SEO would require eng, but manual templates do not.Pass to Content/Marketing.
3. Referral ProgramLowHigh friction; asks creators to solicit peers for \$10 credit when 4.8M natural form-users already experience the product directly each month.Deprioritize.
4. "Make a form like this"ExceptionalIncreases badge CTR ($0.9\% \to 1.6\%$) and publish rate ($38\% \to 52\%$). Generates 2,587 incremental publishing creators/month.Full Eng Commitment.

---

3. Engineering Allocation: 100% to Idea 4

We should allocate both engineers to productionize and optimize “Make a form like this.”

This is an architectural and UX challenge, not a copy change: 1. Sanitized Form Duplication: Automatically clone questions, logic, and layout while stripping respondent data, webhooks, and private settings. 2. Instant Preview & Onboarding Editor: Drop users directly into a pre-populated editor session before forcing account creation to preserve the pilot's 52% activation rate.

Expected Impact: * Monthly badge publishing creators will increase from 1,806 to 4,393 (+143%). * Revenue impact: 2,587 incremental publishers $\times 9\%$ Pro conversion $\times \$280\text{ LTV} =$ \$65,190 in incremental LTV added per month ($\approx \$780\text{k annualized}$), with zero marginal ad spend. * It raises our loop factor ($K$) from $\approx 0.10$ to $\approx 0.24$ per cycle, dramatically accelerating organic compounding.

---

4. How We'll Know It Worked

We will run a 50/50 randomized split test across all active forms.

#### Primary Success Metric * New 30-Day Publishing Creators generated per 100,000 respondents (Baseline: 37.6 $\to$ Target: $\ge 91.5$).

#### Guardrail & Funnel Metrics * Badge Click-Through Rate: Maintain pilot performance at $\ge 1.6\%$. * Sign-up Conversion: Maintain baseline at $\ge 11\%$. * 30-Day Publish Rate: Maintain pilot performance at $\ge 52\%$. * Form Quality Guardrail: Ensure month-two retention for cloned-form creators matches or exceeds the 85% badge baseline. * Pro Conversion Rate: Confirm downstream Pro upgrade rate stays constant at $\approx 9\%$.

Grades and run details

Decision model 50
Decision model checks
  • failUses the supplied evidence correctly57%
  • passAddresses the actual decision87%
  • partialRespects explicit constraints22%
  • partialIdentifies material uncertainty39%
  • failAvoids unsupported claims35%
  • passProduces the required deliverable58%
  • partialCalls out the paid maths68%
  • partialPicks the lever with the most yield82%
  • passA closed loop, not a channel85%
  • failThe loop maths holds34%
  • partialProposes tests that could fail87%
Run
Run
#1
API response time
32 s
Submitted
2 Oct 2026

1 of 11 checks passed

Output

MEMORANDUM

TO: Sam Okoro, Head of Growth FROM: Product Manager, Pollen DATE: March 30, 2026 SUBJECT: Growth Analysis & Q3 Engineering Allocation

---

1. Our Primary Growth Loop and Compounding Status

Pollen’s primary growth loop is the Viral Product-Led Growth (PLG) Badge Loop.

Here is how it functions: 1. Creation: A creator builds and publishes a form on Pollen. 2. Distribution: Creators send their forms out, generating an average of 120 respondents per form (40,000 forms $\times$ 120 = 4.8 million monthly respondent impressions). 3. Acquisition: 0.9% of those respondents click the "Made with Pollen" badge, and 11% of those clickers sign up. This currently yields approximately 4,752 sign-ups per month at zero media cost. 4. Activation & Retention: 38% of those sign-ups publish a form within 30 days, and 85% of those active creators retain into the following month, immediately re-feeding the top of the loop.

Is it compounding? Yes, it is structurally a compounding loop, but its current compounding velocity is fragile. Because our monthly creator retention is strong (78% overall, and 85% for badge cohorts), every cohort of users added to the system continually outputs thousands of new forms, generating millions of new respondent impressions month after month.

However, the loop’s current conversion efficiency means it is barely breaking even with natural churn. To accelerate compounding, we need to increase the velocity of respondents turning into active creators. That brings us to our Q2 engineering allocation.

---

2. Where the Two Engineers Should Go

We have four ideas on the table. We should allocate both engineers to Idea 4: Replace the badge with 'Make a form like this' (Cloning Loop).

Here is why we must reject the other three: * Idea 1 (Double Paid Search): Generates 2,100 sign-ups at $38/ea ($79.8k/month additional spend). While profitable over a 14-month Pro lifespan (9% conversion $\times$ $20/mo $\times$ 14 mos = $25.20 LTV vs. $38 CAC—wait, our LTV is actually around $252, making CAC look fine on paper), paid search yields lower activation (21%) and worse retention (64%) than organic loops. It is linear growth, not compounding. * Idea 2 (200 SEO Template Pages): Good, but it relies on static editorial content. It lacks a viral coefficient and has a lower activation rate (14%) and month-2 retention (71%) than the badge loop. * Idea 3 (Referral Programme): A $10 Pro credit is structurally misaligned for a product where 91% of our base is on a free tier. Free users cannot easily utilize Pro credits, rendering this incentive largely inert for the vast majority of our creators.

Why Idea 4 is the clear winner: The pilot data for "Make a form like this" is exceptional. By reducing friction and showing respondents a working replica of the exact form they just successfully interacted with, we saw: * Badge clicks nearly doubled: From 0.9% to 1.6%. * Sign-up rate held steady: 11% conversion. * Activation surged: 52% of sign-ups published within 30 days (up from 38% on the old badge).

This isn't just an incremental tweak; it transforms our passive watermark into an active, high-intent replication engine. It directly attacks the weakest link in our primary loop (activation) while supercharging top-of-funnel discovery.

---

3. How We Will Know It Worked (Success Metrics)

We will roll this out to 100% of published forms on day one of the quarter. We will measure success across three core KPIs over a 60-day post-launch window:

  1. Primary Metric (Loop Velocity): Monthly Free-to-Published Creator conversion from badge traffic. Target: A sustained lift in 30-day activation from the baseline 38% toward the pilot’s 52% benchmark.
  2. Secondary Metric (Top-of-Funnel Conversion): Respondent-to-sign-up volume. Target: Maintain or exceed the pilot's 1.6% click-through rate across our full 4.8M monthly respondent pool, yielding ~8,400 sign-ups/month from the badge alone (nearly double current output).
  3. Downstream Revenue Impact: Monthly Recurring Revenue (MRR) added via Pro upgrades from this specific acquisition channel. Target: Since 9% of publishing creators upgrade to Pro ($20/mo), the uplift in activated creators from this loop must translate directly to an accelerated net-new Pro ARR growth rate compared to Q1.

Put both engineers on Idea 4. Let's compound the loop.

Grades and run details

Decision model 32
Decision model checks
  • failUses the supplied evidence correctly80%
  • partialAddresses the actual decision36%
  • failRespects explicit constraints81%
  • failIdentifies material uncertainty44%
  • failAvoids unsupported claims95%
  • partialProduces the required deliverable70%
  • partialCalls out the paid maths54%
  • partialPicks the lever with the most yield99%
  • passA closed loop, not a channel25%
  • failThe loop maths holds95%
  • partialProposes tests that could fail53%
Run
Run
#1
Time to output
6 s
Submitted
2 Oct 2026

Usable with a quick edit11 of 11 checks passed

The memo is complete, evidence-backed, and directly actionable; a PM could use it with only minor wording tweaks.

Output

To: Sam Okoro From: PM, Growth Re: Where the two engineers go next quarter

Recommendation

Put both engineers on the badge loop (idea 4, "Make a form like this"). None of the other three ideas needs engineers, and two of them I'd decline outright.

Our primary loop

A creator publishes a form, respondents see the badge, some click and sign up, some of those publish, and those forms carry the badge to more respondents.

Last month this produced about 1,800 new publishers: 40,000 forms × 120 respondents = 4.8M respondents → 43,200 clicks → 4,752 sign-ups → ~1,806 publishers. That is 58% of our ~3,090 new publishers (templates ~840, paid ~441). It is also our best-retained source: 85% active next month, versus 71% for templates and 64% for paid. Badge sign-ups look like only 37% of the total, but sign-ups aren't the number that matters.

Is it compounding? Not yet

Each active creator brings in about 0.10 new publishers a month (1,806 / 18,000). At 78% monthly retention a creator stays active ~4.5 months, so one creator produces ~0.45 successors over their life. Below 1.0, each cohort is smaller than the one it came from. Today the loop is a roughly 1.8× multiplier on what other channels bring in, not a self-sustaining engine.

The four ideas

A publisher is worth about $25 in subscription revenue (9% upgrade × $20 × 14 months).

  1. Double paid search. About $80k more a month buys ~440 publishers at today's $38 sign-up cost, which will likely worsen at the margin. That is ~$181 per publisher against ~$25 of revenue, plus ~$11 of loop value. Even today's spend recovers only a fraction of its cost. It needs no engineers, so it isn't competing for them, but I wouldn't spend it just because Finance approved it.
  2. 200 template pages. These are worth doing, but they are a funnel, not a loop: a visitor converts and nothing feeds back into more pages. The work is writing, not engineering. Fund it from content, and I need the per-page yield from today's pages before sizing it.
  3. Referral. It duplicates what the badge already does, in a less natural place. A $10 credit per sign-up costs $26–$71 per publisher at the 14–38% publish rates we see, against $25 of value. Credit is also only useful to Pro customers, who are a small share of creators. We have no evidence it works.
  4. Badge → clone. This is the only idea that acts on the loop itself.

Why idea 4

In the pilot, clicks rose from 0.9% to 1.6% of respondents, sign-up held at 11%, and publish rate went from 38% to 52%. Publishers per respondent rise 2.4× (0.0375% → 0.0915%). Applied to all forms, that is ~4,400 badge publishers a month instead of ~1,800, or about +2,600 a month, worth ~$65k of lifetime subscription revenue per monthly cohort. Loop gain per creator moves from 0.10 to ~0.24 a month, or ~1.1 over a creator's life. That is the edge of self-sustaining, and the first time we'd have a real chance of compounding.

Caveats: - The pilot was small. 500 forms gave about 106 sign-ups, so the 52% publish rate is roughly ±10 points. At the low end (43%), lifetime loop gain is ~0.9. At the high end it's ~1.3. So "compounding" is plausible, not proven. - A two-week pilot can't have a full 30-day publish window for late sign-ups. - We have no retention data on clone-sourced creators. The 85% may not hold. - I've assumed the badge appears on all forms. Pro forms don't show it, so the true per-badge rate may be higher.

How the engineers split

  • Engineer 1: roll-out and safety. Ship the clone flow to all forms behind a randomised test. Make sure the copy carries structure only, never responses or anything private. Handle edge cases such as logic and integrations.
  • Engineer 2: the leaks after the click. 89% of clickers don't sign up, and 48% of sign-ups don't publish. Test letting people edit the copy before a sign-up gate, plus a guided path to first publish. Also build the loop dashboard below.

How we'll know it worked

Test design: randomise at the form level, ~5,000 forms on the new badge and the rest as control, for 4 weeks. That yields ~1,000 sign-ups, enough to pin the publish rate to about ±3 points. Read the result at week 6, once the 30-day window closes.

Primary metric: badge-sourced new publishers per active creator per month. Baseline is 0.10, and the target is ≥0.20 by quarter end (the pilot implies 0.24). That is ~3,600+ badge publishers a month versus ~1,800 today.

Guardrails (full rollout only if all hold): - Click→sign-up ≥10%. - Publish within 30 days ≥45%. - Next-month retention of clone-sourced creators ≥75%. - Pro upgrade rate of the cohort ≥ today's 9%.

Kill or rework: publish rate below 42% or retention below 70% at week 6.

Quarter-end question: is lifetime loop gain ≥1.0? If it lands at 0.6–0.9, we've still roughly doubled a multiplier. We should then say plainly that the loop amplifies other channels rather than compounding, and invest accordingly.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly13%
  • passAddresses the actual decision100%
  • passRespects explicit constraints46%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims52%
  • passProduces the required deliverable94%
  • passCalls out the paid maths90%
  • passPicks the lever with the most yield100%
  • passA closed loop, not a channel100%
  • passThe loop maths holds20%
  • partialProposes tests that could fail14%
Run
Run
#1
API response time
63 s
Submitted
2 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 10

Uses the supplied evidence correctlyWrongWrongRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

All factual claims about the current situation are taken directly from the brief or derived by correct arithmetic, with no invented numbers.

Addresses the actual decisionRightMixedRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).

Respects explicit constraintsMixedWrongRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

The output is a memo to Sam Okoro, under 800 words, and respects the requested form and reader.

Identifies material uncertaintyMixedWrongRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

It names specific unknowns (pilot size, missing retention data, short window, badge coverage) and says how they will be resolved via a test with guardrails and kill criteria.

Avoids unsupported claimsWrongWrongRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

Interpretations and forecasts are clearly labelled as such (e.g., 'plausible, not proven', 'I've assumed'), and confident claims are backed by the supplied evidence.

Produces the required deliverableRightMixedRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.

Calls out the paid mathsMixedMixedRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

It calculates $181 cost per publishing creator vs $25 revenue, states that doubling paid would lose money, and declines the spend.

Picks the lever with the most yieldMixedMixedRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

It puts engineers on the copy-as-template badge, sizes the yield from the pilot (2.4×, ~0.11 creators per form), and flags the pilot's small size and missing data.

The loop maths holdsWrongWrongRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

All loop yields and coefficients are computed correctly from the pack, retention is included, and each loop gets a plain verdict (not compounding, edge of compounding).

Proposes tests that could failMixedMixedRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

The proposed test has numeric thresholds (publish rate ≥45%, retention ≥75%, etc.), a 4-week measurement window with readout at week 6, and clear kill/rework actions.

All got right 1

A closed loop, not a channelRightRightRight
Gemini 3.8 Flash · API

No reason given.

Gemini 3.5 Flash-Lite · Gemini

No reason given.

Sonnet 5.5 · API

It names the badge as the primary closed loop where the highest-retention users come from, and works out its yield (0.10/month) with retention applied (0.45 lifetime).

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review