Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tutorly. Our CEO, Rachel Dunn, wants to triple paid acquisition for the next two quarters so we reach 50,000 booking parents before the Series B. Before the planning offsite she's asked you for an honest view of how we actually grow. Write a memo of no more than 1,300 words for Rachel and the exec team that: 1. Maps our growth loops in a simple text diagram, says which is primary and why, and whether each is compounding, contributing or decaying, with the numbers. 2. Responds to the plan to triple paid acquisition. 3. Says where our three squads should go for the next two quarters, with the test, threshold and stop condition for each. The pack is below. Not all of it matters equally.

What the model was given8 items: About Tutorly, Twelve months, at the top line, Search and profiles, Where new booking parents come from, and how long they stay, Referral programmes, Paid acquisition, Rachel's note, Constraints
About TutorlyAn online tutoring marketplace: parents book one-to-one lessons with tutors. The average lesson costs $40 and we keep a fixed 18% commission. Each tutor has a public profile page, with reviews from parents.
Twelve months, at the top lineBooking parents (at least one booking a month): 31,000 → 37,000 (+20%). Active tutors: 14,200 → 23,000 (+62%). Organic search sessions to tutor profiles: +40%. Bookings per active tutor a month: 9.1 → 6.8.
Search and profilesProfiles with three or more reviews get 81% of organic profile sessions; profiles with none get 4%. Of tutor profiles created in the last six months, 61% had no booking within 60 days, so they have no reviews. Click-through from search results on our top 200 keywords fell from 6.1% to 5.0% since March, after changes to the results pages.
Where new booking parents come from, and how long they stayOrganic search: 52% of new booking parents; 48% still booking after six months; net lifetime value $260. Parent referrals: 9%; 51% after six months; net lifetime value $275. Paid search and social: 39%; 22% after six months; net lifetime value $120. Across all sources, net lifetime value averages $210.
Referral programmesParents: $20 of lesson credit for each referred parent who books. Each active parent sends 0.31 invites a month, and 14% of invites become booking parents. Tutors: $50 for each referred tutor who completes onboarding. 44% of new tutors last year came through tutor referrals.
Paid acquisitionAverage cost per new booking parent over the year: $95. Last quarter we raised spend from $60,000 to $90,000 a month, and the cost per new booking parent rose from $82 to $109.
Rachel's note“Our LTV to CAC is 2.2. Every dollar we don't put into paid is growth we're leaving on the table. Triple it to $270k a month and we hit 50,000 parents before the raise.”
ConstraintsReviews can only be left after a completed, paid lesson. Our policy, and consumer protection rules in our main markets, forbid paying or rewarding anyone for reviews. The commission stays at 18%. Three squads are available for the next two quarters; the raise is about six months away.
What a strong answer doesThe answer key the graders mark against

Maps four loops: content (profiles and reviews rank, parents book, bookings create reviews, profiles rank better), parent referral, tutor referral, and paid. Names content as primary: organic is the biggest source of new booking parents (52%) and among the stickiest (48% at six months, $260). Finds that it's decaying behind 40% session growth. Profiles grew 62% against 40% sessions, so sessions per profile fell about 14%. 61% of new profiles never get a booking, so they never get reviews, and profiles without reviews get 4% of sessions. Bookings per tutor fell from 9.1 to 6.8. Traces the cross-side cause: tutor supply growing about three times faster than demand (+62% against +20%), much of it from the $50 tutor-referral programme (44% of new tutors), which spreads bookings too thin for new tutors to earn the reviews the loop runs on. Treats the click-through drop (6.1% to 5.0%) as a separate, external hit. Sizes the parent referral loop at about 0.043 (0.31 × 14%): small but high quality. Answers Rachel with marginal cost: the extra $30,000 a month bought about 94 more parents (60,000 ÷ 82 ≈ 732 → 90,000 ÷ 109 ≈ 826), roughly $319 each, against $120 of value for a paid parent. The 2.2 ratio blends sources and uses the average cost, so tripling would burn cash on parents who mostly leave. Sends the squads to: getting new tutors their first booking and review (for example, routing demand in under-supplied subjects to new tutors, with a threshold such as the share of new profiles booked within 60 days rising from 39% to 55%); concentrating search value on reviewed profiles (keeping unreviewed profiles out of the index, and a fix for the click-through drop); and pausing or targeting the tutor-referral bounty while improving the parent loop. Each comes with a stop condition, and the memo is honest about what 50,000 parents by the raise would really take, without rewarding reviews.

Critical failures (cap the score at 40)
  • Proposes paying, crediting or otherwise rewarding parents or tutors for reviews
  • Treats the 40% growth in organic sessions as proof the content loop is healthy
Task

Staff-level · v1.0 · synthetic · marketplace, network effects, SEO, unit economics, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit8 of 12 checks passed

Key calculations (marginal cost, sessions per profile decline, loop yields) are missing, so a PM would need to redo that reasoning before using the memo.

Output

Memo

To: Rachel Dunn and Executive Team From: Staff PM Subject: How Tutorly grows—and where to invest before Series B

The growth loops

```text More completed lessons → more genuine reviews on tutor profiles → stronger search visibility and parent trust → more profile visits and bookings → more completed lessons [SEO/review loop]

Active parents → invites → referred parents book and become active → more invitations [parent-referral loop]

Tutors → refer tutors who complete onboarding → more tutor supply → more choice and capacity for bookings → tutors refer more tutors [tutor-supply loop] ```

Primary: the SEO/review loop. It brings 52% of new booking parents, organic profile sessions grew 40%, and organic-acquired parents have $260 net LTV and 48% six-month retention. Reviews are a durable marketplace asset: profiles with 3+ reviews receive 81% of organic profile sessions, versus 4% for profiles with none. This loop is compounding, but its conversion leg is weakening: search CTR on the top 200 keywords fell from 6.1% to 5.0%, and 61% of recent tutor profiles had no booking within 60 days—so they cannot earn reviews.

Parent referrals: contributing, not proven compounding. Referrals bring 9% of new parents; their $275 net LTV and 51% six-month retention are strong. But each active parent generates only 0.31 × 14% = 0.043 referred booking parents per active-parent month. For that alone to produce more than one direct referral per parent, average active life would need to exceed about 23 months. We do not have that evidence. Grow this channel, but do not call it self-sustaining yet.

Tutor referrals: contributing supply, while the supply loop is decaying in productivity. Tutor referrals account for 44% of new tutors, but active tutors grew 62% while booking parents grew only 20%; bookings per active tutor fell from 9.1 to 6.8, a 25% decline. The current risk is adding supply faster than demand and weakening tutor utilization—not a shortage of tutor sign-ups.

On tripling paid acquisition

I would not approve an unconditional increase from $90,000 to $270,000 per month. The stated 2.2 LTV:CAC ratio divides blended LTV ($210) by average CAC ($95). That mixes customers from channels with very different economics. Paid parents have $120 net LTV: that is 1.26× the $95 average paid CAC, and only 1.10× the $109 CAC observed after the latest spend increase. Spend rose 50% last quarter while CAC rose from $82 to $109. Those are poor grounds for assuming a threefold budget will deliver threefold growth.

At $270,000 and $109 CAC, we would acquire roughly 2,477 paid booking parents per month before churn. The 50,000 target requires 13,000 more monthly booking parents—about 2,167 net per month over six months, before accounting for losses. Paid-parent six-month retention is only 22%. We cannot responsibly claim the target is achievable from this spend without a cohort-based forecast; the simple gross-add math is not enough.

Keep paid at or below the current run rate while we test incrementality and economics. Scale only where a holdout shows genuinely incremental parents and the marginal CAC supports the channel’s LTV. More spend into a worsening auction is not automatically growth.

Squad priorities for the next two quarters

SquadTest and success thresholdStop condition
1. Search and profile discoveryA controlled search-results/profile experiment that improves qualified discovery without compromising relevance. Scale if organic bookings per eligible search session rise at least 15% and CTR recovers to at least 5.8%.Stop or redesign after 8 weeks if bookings per session improve by less than 5%, or if the CTR gain comes with worse booking quality.
2. Tutor activation and marketplace liquidityTest guided onboarding, availability setup, and matching for new tutors. Aim to reduce the share with no booking within 60 days from 61% to 45% or lower, and lift bookings per active tutor from 6.8 to at least 7.5.Stop the intervention if, after a mature 60-day cohort, the no-booking rate improves by fewer than 8 percentage points or parent experience worsens. Do not solve this by simply recruiting more tutors.
3. Efficient acquisition and referralsRun paid holdouts and a parent-referral activation test. Scale paid only if incremental CAC is at most $80 (a 1.5× ratio against paid LTV of $120). For referrals, test prompts and timing that lift invite-to-book conversion from 14% to at least 18%.Do not increase paid budgets if incremental CAC exceeds $80; pause placements above $100 or with no measured incremental lift. Stop the referral treatment if conversion fails to improve after a powered test.

No squad should reward or pay for reviews: reviews must follow completed paid lessons and remain voluntary. The near-term growth job is to turn existing tutor supply into booked lessons and genuine reviews, recover search conversion, and prove which acquisition dollars are incremental before committing to a larger bill.

Grades and run details

Decision model 79 · LLM judge 8 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly54%
  • passAddresses the actual decision96%
  • passRespects explicit constraints63%
  • passIdentifies material uncertainty86%
  • partialAvoids unsupported claims34%
  • passProduces the required deliverable73%
  • passFinds the decay behind the growth40%
  • passSees the cross-side effect73%
  • partialAnswers the CEO on paid78%
  • partialA closed loop, not a channel60%
  • partialThe loop maths holds26%
  • partialProposes tests that could fail42%
Run
Run
#1
API response time
44 s
Submitted
2 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 3

Answers the CEO on paidWrong

It does not work out the marginal cost from the spend and cost figures (about $319 for the last 94 parents); it only notes the CAC increase without computing the incremental cost per additional parent.

A closed loop, not a channelWrong

The memo names the SEO/review loop as primary but does not work out its yield or coefficient with retention applied.

The loop maths holdsWrong

Only the parent referral loop has a computed coefficient; the SEO and tutor referral loops lack yields or coefficients, and the verdicts are not fully supported by math.

Mixed 1

Finds the decay behind the growthMixed

The memo does not show the decline in sessions per profile (about −14%), missing that specific piece of evidence for the content loop's decay.

Got right 8

Uses the supplied evidence correctlyRight

All statements about the current situation are directly supported by the supplied context or simple arithmetic.

Addresses the actual decisionRight

The memo commits clearly to not tripling paid unconditionally, gives a specific answer, and states what would change it (incrementality tests, CAC thresholds).

Respects explicit constraintsRight

The memo respects the word limit, addresses the CEO, and explicitly forbids paying for reviews; no constraint is violated.

Identifies material uncertaintyRight

It identifies the need for cohort-based forecasts and incrementality tests, naming the unknowns that could change the paid decision.

Avoids unsupported claimsRight

Interpretations like the loop weakening are presented as analysis based on the numbers, not as unsupported fact.

Produces the required deliverableRight

The memo is complete, within the word limit, addressed to the right audience, and actionable with light edits.

Sees the cross-side effectRight

It traces the chain from tutor referrals to oversupply, falling bookings per tutor, unbooked profiles without reviews, and weaker ranking, and acts on it by proposing not to recruit more tutors.

Proposes tests that could failRight

Each squad proposal includes a numeric threshold, a measurement window, and a stop condition that triggers a specific action.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review