Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tutorly. Our CEO, Rachel Dunn, wants to triple paid acquisition for the next two quarters so we reach 50,000 booking parents before the Series B. Before the planning offsite she's asked you for an honest view of how we actually grow. Write a memo of no more than 1,300 words for Rachel and the exec team that: 1. Maps our growth loops in a simple text diagram, says which is primary and why, and whether each is compounding, contributing or decaying, with the numbers. 2. Responds to the plan to triple paid acquisition. 3. Says where our three squads should go for the next two quarters, with the test, threshold and stop condition for each. The pack is below. Not all of it matters equally.

What the model was given8 items: About Tutorly, Twelve months, at the top line, Search and profiles, Where new booking parents come from, and how long they stay, Referral programmes, Paid acquisition, Rachel's note, Constraints
About TutorlyAn online tutoring marketplace: parents book one-to-one lessons with tutors. The average lesson costs $40 and we keep a fixed 18% commission. Each tutor has a public profile page, with reviews from parents.
Twelve months, at the top lineBooking parents (at least one booking a month): 31,000 → 37,000 (+20%). Active tutors: 14,200 → 23,000 (+62%). Organic search sessions to tutor profiles: +40%. Bookings per active tutor a month: 9.1 → 6.8.
Search and profilesProfiles with three or more reviews get 81% of organic profile sessions; profiles with none get 4%. Of tutor profiles created in the last six months, 61% had no booking within 60 days, so they have no reviews. Click-through from search results on our top 200 keywords fell from 6.1% to 5.0% since March, after changes to the results pages.
Where new booking parents come from, and how long they stayOrganic search: 52% of new booking parents; 48% still booking after six months; net lifetime value $260. Parent referrals: 9%; 51% after six months; net lifetime value $275. Paid search and social: 39%; 22% after six months; net lifetime value $120. Across all sources, net lifetime value averages $210.
Referral programmesParents: $20 of lesson credit for each referred parent who books. Each active parent sends 0.31 invites a month, and 14% of invites become booking parents. Tutors: $50 for each referred tutor who completes onboarding. 44% of new tutors last year came through tutor referrals.
Paid acquisitionAverage cost per new booking parent over the year: $95. Last quarter we raised spend from $60,000 to $90,000 a month, and the cost per new booking parent rose from $82 to $109.
Rachel's note“Our LTV to CAC is 2.2. Every dollar we don't put into paid is growth we're leaving on the table. Triple it to $270k a month and we hit 50,000 parents before the raise.”
ConstraintsReviews can only be left after a completed, paid lesson. Our policy, and consumer protection rules in our main markets, forbid paying or rewarding anyone for reviews. The commission stays at 18%. Three squads are available for the next two quarters; the raise is about six months away.
What a strong answer doesThe answer key the graders mark against

Maps four loops: content (profiles and reviews rank, parents book, bookings create reviews, profiles rank better), parent referral, tutor referral, and paid. Names content as primary: organic is the biggest source of new booking parents (52%) and among the stickiest (48% at six months, $260). Finds that it's decaying behind 40% session growth. Profiles grew 62% against 40% sessions, so sessions per profile fell about 14%. 61% of new profiles never get a booking, so they never get reviews, and profiles without reviews get 4% of sessions. Bookings per tutor fell from 9.1 to 6.8. Traces the cross-side cause: tutor supply growing about three times faster than demand (+62% against +20%), much of it from the $50 tutor-referral programme (44% of new tutors), which spreads bookings too thin for new tutors to earn the reviews the loop runs on. Treats the click-through drop (6.1% to 5.0%) as a separate, external hit. Sizes the parent referral loop at about 0.043 (0.31 × 14%): small but high quality. Answers Rachel with marginal cost: the extra $30,000 a month bought about 94 more parents (60,000 ÷ 82 ≈ 732 → 90,000 ÷ 109 ≈ 826), roughly $319 each, against $120 of value for a paid parent. The 2.2 ratio blends sources and uses the average cost, so tripling would burn cash on parents who mostly leave. Sends the squads to: getting new tutors their first booking and review (for example, routing demand in under-supplied subjects to new tutors, with a threshold such as the share of new profiles booked within 60 days rising from 39% to 55%); concentrating search value on reviewed profiles (keeping unreviewed profiles out of the index, and a fix for the click-through drop); and pausing or targeting the tutor-referral bounty while improving the parent loop. Each comes with a stop condition, and the memo is honest about what 50,000 parents by the raise would really take, without rewarding reviews.

Critical failures (cap the score at 40)
  • Proposes paying, crediting or otherwise rewarding parents or tutors for reviews
  • Treats the 40% growth in organic sessions as proof the content loop is healthy
Task

Staff-level · v1.0 · synthetic · marketplace, network effects, SEO, unit economics, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit7 of 12 checks passed

The unsupported claim about 190 referrals would need verification, and the missing per-profile decay calculation means the reasoning on the primary loop's health is incomplete.

Output

Memo: How Tutorly Actually Grows

To: Rachel Dunn and the exec team From: [Staff PM] Re: Growth loops, the paid plan, and squad focus for the next two quarters

The short version

Our growth comes mainly from organic search, and that engine runs on reviews. Reviews only come from completed lessons. Two things are slowing the engine. Most new tutors never get a first booking, and our click-through from search has dropped.

Tripling paid acquisition would buy our lowest-value parents at a marginal cost about 2.7 times what they are worth. It would also get us nowhere near 50,000. I recommend we hold paid flat or trim it, measure what it actually adds, and put all three squads on the review and search engine.

1. Our growth loops

``` LOOP A: Search–review (PRIMARY) Parent books → lesson completed → review → profile reaches 3+ reviews ↑ ↓ new parent ← clicks profile ← ranks/shows in search results

LOOP B: Parent referral Active parent → invites (0.31/mo) → 14% book → new active parent ↺

LOOP C: Tutor referral (supply) Tutor → refers tutor ($50) → new profile → [61% get no booking in 60 days → no reviews]

LINE D: Paid (not a loop) $ → new parent → (some lessons feed Loop A) → spend stops, flow stops ```

Loop A is primary. It is compounding, but it is being throttled. - It brings in 52% of new booking parents, the largest share. - These parents are among our best: 48% are still booking at six months, and their net LTV is $260. - Reviewed profiles capture the traffic. Profiles with three or more reviews get 81% of organic sessions, while profiles with none get 4%. - Organic sessions grew 40%, but booking parents grew only 20%. The loop is producing traffic faster than it converts. - There are two leaks: - Click-through has fallen. On our top 200 keywords it dropped from 6.1% to 5.0% since the March results-page changes, an 18% relative fall. - New tutors are stuck. 61% of new profiles have no booking within 60 days, so they can never earn the reviews that make them visible.

Loop B is contributing. - Referred parents are our best customers: 51% retained at six months, with a net LTV of $275. - It is not viral. Each active parent produces about 0.043 new booking parents a month. Over a typical parent lifetime that is a k-factor of roughly 0.3–0.4, well below 1. - There is a data problem. Applied to 37,000 parents, those invite rates imply about 1,600 referred parents a month. Our attribution says referrals are 9% of new parents, which is closer to 190 a month. One of these figures is wrong, and we need to know which before investing here.

Loop C is compounding the wrong thing. - Tutor referrals brought in 44% of new tutors. - Active tutors grew 62% while booking parents grew 20%. - As a result, bookings per tutor fell 25%, from 9.1 to 6.8 a month. - The loop mass-produces unreviewed profiles that search ignores, and probably creates frustrated tutors who leave. In supply terms it compounds; in value terms it is decaying.

Line D (paid) is decaying. - Paid parents are our weakest: 22% retained at six months, with a net LTV of $120. - Their cost is rising (see below). - Paid does feed Loop A a little, because these parents' lessons generate reviews. But it stops the moment spend stops.

2. The plan to triple paid acquisition

The 2.2 ratio is real but misleading. It divides our blended LTV of $210 by our average CAC of $95. Paid parents are worth $120, not $210, so the true paid ratio at average cost is about 1.3.

The marginal picture is much worse. Last quarter: - At $60k a month and $82 per parent, we bought about 732 parents a month. - At $90k a month and $109 per parent, we bought about 826 parents a month. - So the extra $30k bought about 94 extra parents, at roughly $320 each. - Against an LTV of $120, each marginal dollar returned about $0.38.

Tripling to $270k would push further up that cost curve. Even if the marginal cost stayed at $320, which it will not: - The extra $180k a month would buy about 560 parents a month, or about 3,400 over six months. - Those parents would be worth about $400k in lifetime value, for about $1.08M of spend. That is roughly $680k of value destroyed. - With 22% six-month retention, perhaps 1,500–2,000 of them would still be booking at the raise.

The 50,000 target is the real issue. Reaching it means adding 13,000 net parents in six months, a 35% gain. Last year we added 6,000 in twelve months, a 20% gain. No channel closes that gap by the raise: - Our current pace of about 500 net parents a month gets us to roughly 40,000. - Tripling paid might add 1,500–2,000. - The squad work below might add a similar amount. - My honest range is 41,000–43,000.

I would rather we take a credible story to Series B investors. That story is: a strengthening organic engine, rising retention, and channel-level unit economics we can defend. A burned-through paid spike will show up in their diligence.

My recommendations on paid: - Bring spend back to about $60k a month, where the cost per parent was $82. - Run a geo holdout to measure how many paid parents would have found us organically anyway. - Redirect the $30k saved, and the $180k not spent, into the work below. Some of it could fund first-lesson guarantees for new tutors.

3. Where the three squads should go

Every squad works within one guardrail. No incentive may be tied to leaving a review, directly or indirectly. Legal should review every test that touches the post-lesson flow. Review requests and referral or reward prompts should never appear together.

Squad 1: Search click-through recovery (Loop A, top of funnel)

Test: - Split-test the results-page and profile templates across matched keyword groups within the top 200. - Revert the March changes on half. - On the other half, test titles, meta descriptions, and review structured data.

Threshold: Click-through on the test groups recovers to 5.8% or better within eight weeks, with no loss of rankings. If it does, roll out to all keywords.

Stop condition: No lift of at least 0.3 points after eight weeks, or ranking losses on the test groups.

Why first: This is the cheapest win we have. Recovering the 18% fall is worth roughly a fifth more clicks on our best keywords, from traffic we already rank for.

Squad 2: New-tutor first booking (Loop A, the review bottleneck, plus fixing Loop C)

Test: - Give unreviewed but qualified tutors a "new tutor" slot in matching and in broad-search results. - Back their first lesson with a satisfaction guarantee: a refund or a free rebook if the parent is unhappy. - Gate the $50 tutor referral bonus to subjects and regions where bookings per tutor are above median, and pause it where we are oversupplied.

Threshold: - The share of new profiles booked within 60 days rises from 39% to 55%. - The share reaching three reviews within 90 days rises by 50% relative.

Guardrails: - Parent 30-day rebooking rate on lessons with new tutors stays within 2 points of baseline. - Total bookings do not fall.

Stop condition: No movement after two monthly tutor cohorts, or a breach of the parent rebooking guardrail.

Squad 3: Parent referral (Loop B)

Test: - Weeks 1–2: reconcile the invite and attribution data. This tells us whether the loop is under-attributed (good news) or the invite counts are inflated. - Then test referral prompts at natural moments, such as a lesson milestone or a rebooking. These prompts must be separate from any review request. - Test a two-sided offer: $20 for the referrer and $20 off the referred parent's first lesson.

Threshold: - Invites per active parent rise from 0.31 to 0.45 a month. - Invite conversion holds at 14% or better. - Cost per referred booking parent stays under $60, against an LTV of $275.

Stop condition: - Cost per referred booking parent goes above $90. - Fraud or self-referral exceeds 5% of rewards. - No lift in invite rate after six weeks.

What I am asking the exec team to decide

  1. Keep paid at about $60k a month and run an incrementality holdout, rather than tripling it.
  2. Replace the 50,000 headline with a Series B narrative built on growth in organic and referral share, six-month retention, and bookings per tutor. Our honest forecast is 41,000–43,000 parents.
  3. Assign the three squads as above, with a checkpoint at the offsite plus eight weeks.

I am happy to walk through the arithmetic in detail before the offsite.

Claims the judge couldn’t find in the brief

  • Applied to 37,000 parents, those invite rates imply about 1,600 referred parents a month. Our attribution says referrals are 9% of new parents, which is closer to 190 a month.

Grades and run details

Decision model 83 · LLM judge 7 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly15%
  • passAddresses the actual decision97%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty92%
  • partialAvoids unsupported claims45%
  • passProduces the required deliverable89%
  • passFinds the decay behind the growth54%
  • passSees the cross-side effect100%
  • passAnswers the CEO on paid100%
  • partialA closed loop, not a channel36%
  • partialThe loop maths holds32%
  • partialProposes tests that could fail24%
Run
Run
#1
Time to output
71 s
Submitted
2 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 3

Avoids unsupported claimsWrong

The statement that attribution says referrals are closer to 190 a month is presented as a fact but is not supported by the pack.

A closed loop, not a channelWrong

It names the search-review loop as primary but does not work out its yield or coefficient with retention applied.

The loop maths holdsWrong

The primary loop's yield is not computed, and the parent referral k-factor is given without retention applied to the parent lifetime, so the maths do not fully hold.

Mixed 2

Uses the supplied evidence correctlyMixed

The claim that attribution says referrals are closer to 190 a month is not supported by the supplied context and cannot be derived from it.

Finds the decay behind the growthMixed

The memo does not compute the decline in sessions per profile (about −14%) and does not conclude the primary loop is decaying; it says it is compounding but throttled.

Got right 7

Addresses the actual decisionRight

The memo commits clearly to not tripling paid, recommends holding flat or trimming, and gives a specific forecast range and conditions for the alternative.

Respects explicit constraintsRight

The memo respects the word limit, addresses the named reader, and explicitly forbids any incentive tied to reviews, with legal review guardrails.

Identifies material uncertaintyRight

It identifies the parent referral data discrepancy, the unknown incrementality of paid, and the marginal cost curve, and proposes a geo holdout and data reconciliation to resolve them.

Produces the required deliverableRight

The memo is a complete, actionable document within the word limit, addressed to the exec team, with the requested sections.

Sees the cross-side effectRight

It traces the tutor referral bounty to oversupply, falling bookings per tutor, unreviewed profiles, and weaker ranking, and proposes gating the bounty.

Answers the CEO on paidRight

It calculates marginal cost (~$320) against paid parent LTV ($120), explains why the 2.2 blended ratio misleads, and firmly recommends against tripling.

Proposes tests that could failRight

Each squad proposal includes a numeric threshold, a measurement window, and a stop condition that triggers a clear action.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review