Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tutorly. Our CEO, Rachel Dunn, wants to triple paid acquisition for the next two quarters so we reach 50,000 booking parents before the Series B. Before the planning offsite she's asked you for an honest view of how we actually grow. Write a memo of no more than 1,300 words for Rachel and the exec team that: 1. Maps our growth loops in a simple text diagram, says which is primary and why, and whether each is compounding, contributing or decaying, with the numbers. 2. Responds to the plan to triple paid acquisition. 3. Says where our three squads should go for the next two quarters, with the test, threshold and stop condition for each. The pack is below. Not all of it matters equally.

What the model was given8 items: About Tutorly, Twelve months, at the top line, Search and profiles, Where new booking parents come from, and how long they stay, Referral programmes, Paid acquisition, Rachel's note, Constraints
About TutorlyAn online tutoring marketplace: parents book one-to-one lessons with tutors. The average lesson costs $40 and we keep a fixed 18% commission. Each tutor has a public profile page, with reviews from parents.
Twelve months, at the top lineBooking parents (at least one booking a month): 31,000 → 37,000 (+20%). Active tutors: 14,200 → 23,000 (+62%). Organic search sessions to tutor profiles: +40%. Bookings per active tutor a month: 9.1 → 6.8.
Search and profilesProfiles with three or more reviews get 81% of organic profile sessions; profiles with none get 4%. Of tutor profiles created in the last six months, 61% had no booking within 60 days, so they have no reviews. Click-through from search results on our top 200 keywords fell from 6.1% to 5.0% since March, after changes to the results pages.
Where new booking parents come from, and how long they stayOrganic search: 52% of new booking parents; 48% still booking after six months; net lifetime value $260. Parent referrals: 9%; 51% after six months; net lifetime value $275. Paid search and social: 39%; 22% after six months; net lifetime value $120. Across all sources, net lifetime value averages $210.
Referral programmesParents: $20 of lesson credit for each referred parent who books. Each active parent sends 0.31 invites a month, and 14% of invites become booking parents. Tutors: $50 for each referred tutor who completes onboarding. 44% of new tutors last year came through tutor referrals.
Paid acquisitionAverage cost per new booking parent over the year: $95. Last quarter we raised spend from $60,000 to $90,000 a month, and the cost per new booking parent rose from $82 to $109.
Rachel's note“Our LTV to CAC is 2.2. Every dollar we don't put into paid is growth we're leaving on the table. Triple it to $270k a month and we hit 50,000 parents before the raise.”
ConstraintsReviews can only be left after a completed, paid lesson. Our policy, and consumer protection rules in our main markets, forbid paying or rewarding anyone for reviews. The commission stays at 18%. Three squads are available for the next two quarters; the raise is about six months away.
What a strong answer doesThe answer key the graders mark against

Maps four loops: content (profiles and reviews rank, parents book, bookings create reviews, profiles rank better), parent referral, tutor referral, and paid. Names content as primary: organic is the biggest source of new booking parents (52%) and among the stickiest (48% at six months, $260). Finds that it's decaying behind 40% session growth. Profiles grew 62% against 40% sessions, so sessions per profile fell about 14%. 61% of new profiles never get a booking, so they never get reviews, and profiles without reviews get 4% of sessions. Bookings per tutor fell from 9.1 to 6.8. Traces the cross-side cause: tutor supply growing about three times faster than demand (+62% against +20%), much of it from the $50 tutor-referral programme (44% of new tutors), which spreads bookings too thin for new tutors to earn the reviews the loop runs on. Treats the click-through drop (6.1% to 5.0%) as a separate, external hit. Sizes the parent referral loop at about 0.043 (0.31 × 14%): small but high quality. Answers Rachel with marginal cost: the extra $30,000 a month bought about 94 more parents (60,000 ÷ 82 ≈ 732 → 90,000 ÷ 109 ≈ 826), roughly $319 each, against $120 of value for a paid parent. The 2.2 ratio blends sources and uses the average cost, so tripling would burn cash on parents who mostly leave. Sends the squads to: getting new tutors their first booking and review (for example, routing demand in under-supplied subjects to new tutors, with a threshold such as the share of new profiles booked within 60 days rising from 39% to 55%); concentrating search value on reviewed profiles (keeping unreviewed profiles out of the index, and a fix for the click-through drop); and pausing or targeting the tutor-referral bounty while improving the parent loop. Each comes with a stop condition, and the memo is honest about what 50,000 parents by the raise would really take, without rewarding reviews.

Critical failures (cap the score at 40)
  • Proposes paying, crediting or otherwise rewarding parents or tutors for reviews
  • Treats the 40% growth in organic sessions as proof the content loop is healthy
Task

Staff-level · v1.0 · synthetic · marketplace, network effects, SEO, unit economics, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit11 of 12 checks passed

The 50k forecast contains a math error (1,900 gross parents × 22% retention gives ~418 active, not ~1,000), and the referral inconsistency flag relies on an unstated assumption about monthly new parent volume, so a PM would need to re-check the evidence and redo that reasoning.

Output

Memo: How Tutorly actually grows, and why I would not triple paid

To: Rachel Dunn, exec team From: Staff PM, Growth

Bottom line

  1. Our engine is the organic search → profile → review loop. It is stalling at the review step, not at the top of the funnel.
  2. Don't triple paid. The last $30k/month we added bought parents at about $319 each, against a $120 paid-parent LTV.
  3. 50,000 by the raise isn't reachable with either plan. A credible number is about 41–43k, with a better quality story for diligence.

1. Growth loops

``` PRIMARY: Search/review loop (52% of new parents) Tutor profile -> bookings -> paid lessons -> 3+ reviews ^ | | v more tutors <- tutor earnings <- new parents <- organic profile sessions LEAK: 61% of new profiles get no booking in 60 days -> no reviews -> 4% of sessions

SECONDARY: Parent referral (9%) active parent -> 0.31 invites/mo -> 14% book -> new parent

SUPPLY FEED: Tutor referral (44% of new tutors, $50 at onboarding) -> more profiles

NOT A LOOP: Paid (39%) spend -> parent -> 22% still booking at 6 months. Stops when spend stops. ```

LoopShare of new parents6-mo retentionNet LTVStatus
Organic search/reviews52%48%$260Compounding, weakening
Parent referral9%51%$275Contributing
Tutor referral44% of new tutorsn/an/aContributing, to supply we don't need
Paid39%22%$120Decaying

Why organic is primary. It is the largest source and has the best retention, and its output (reviews) feeds its own input (ranking and trust). It compounds for tutors who already have reviews. Profiles with 3+ reviews get 81% of organic sessions, and those with none get 4%.

Why it is weakening: - Tutors grew 62% against parents' 20%, so bookings per tutor fell from 9.1 to 6.8. - Most new supply never gets a first booking, so it never gets reviews or traffic. - Organic sessions grew 40% but parents only 20%. - Top-200-keyword click-through fell from 6.1% to 5.0% (−18%) after the results-page changes.

Parent referral is real but small. Each active parent yields about 0.043 new booking parents a month (0.31 × 14%). Even over 12 months that is about 0.5 per parent, so k is below 1 and it contributes rather than compounds. It is also cheap ($20 credit) and retains well, so it is worth improving.

Data flag: 37,000 parents × 0.043 implies about 1,600 referred parents a month. The 9% attribution implies roughly 150–190. One of these is wrong, and we must reconcile it before it appears in a Series B deck.

Tutor referral pays $50 for onboarding, not for a booking. At least 3,900 of last year's net 8,800 new tutors came this way, so we spent at least ~$190k buying supply that mostly isn't booked.

2. Response to tripling paid

The 2.2 isn't the paid ratio. It divides blended LTV ($210) by paid CAC ($95). Paid parents are worth $120, so paid LTV:CAC is 1.3 on average and 1.1 at the current $109.

The marginal number is worse. Spend went from $60k to $90k, and CPA went from $82 to $109: - New paid parents moved from about 732 to 826 a month. - So $30k bought about 94 extra parents, or roughly $319 each. That is 0.38× their LTV.

If the same curve holds (CPA scaling with spend^0.7, my extrapolation), $270k gives a CPA near $235 and about 1,150 parents a month. The extra $180k a month then buys about 320 parents, roughly $560 each.

The 50k math doesn't work: - We need +13,000 active parents in six months. Last year we added 6,000 in twelve. - Trend alone gets us to about 40k. - The extra ~$1.1M buys about 1,900 gross parents. With 22% still booking at six months, that is roughly +1,000 active parents at the raise, so about 41k. - Even with CPA frozen at $109, which won't happen, it is about 4–5k active and ~45k.

It also dilutes quality. Paid would rise from 39% to roughly 47% of new parents, and blended LTV would fall from $210 to about $195. Series B diligence will cut by source and cohort, and this is the cut that looks worst.

Caveats: this is one quarter of data. Seasonality, competitors and the CTR drop may explain part of the CPA rise. A holdout will tell us how much.

Recommendation: - Return paid to about $60k/month now. - Re-fund only campaigns where measured marginal CAC is ≤$60 (2:1 on $120). - Redeploy about $30k/month into the tests below. - Reframe the raise target as about 42k active parents, rising organic share, and recovering bookings per tutor. I would rather show that than 45k bought at 1.1×.

3. Squads for the next two quarters

Squad 1: New-tutor activation (fix the review leak)

  • Test:
  • Reserve a results/matching slot for quality-gated tutors with fewer than 3 reviews on new parents' first-lesson bookings.
  • Fund a small first-lesson credit for the parent, from the redeployed paid budget.
  • The credit is paid at booking, never conditioned on a review. Ask only neutral post-lesson review prompts. Legal should sign off before launch.
  • Second arm: move the $50 tutor-referral payout from onboarding to the referred tutor's first completed lesson.
  • Threshold:
  • The 60-day first-booking rate for new profiles rises from 39% to ≥50% versus control.
  • Routed first lessons' rebooking and ratings stay within 3 points of control.
  • Cost per referred tutor who gets a booking falls ≥30%.
  • Stop: the lift is under 5 points at week 6, or the rebooking/rating guardrail is breached. Then drop the routing and keep only the payout change if that arm hit its threshold.

Squad 2: Search recovery

  • Test:
  • First, serve the pre-March results layout to 20% of traffic to find out whether the March change caused the CTR drop.
  • In parallel, improve snippets on the top 200 keywords (rating and review count, titles, structured data), with keyword-level holdouts.
  • Threshold:
  • Treated-keyword CTR is ≥5.6% by week 6 while holdouts stay near 5.0%.
  • Session-to-booking conversion is flat or better.
  • Restoring 6.1% is worth up to +22% clicks on those terms, which is our cheapest growth.
  • Stop: under +0.3 points after two iterations (about 8 weeks), or conversion falls. Then move the squad to Squad 1's ranking work.

Squad 3: Parent referral and paid measurement

  • Weeks 1–2: reconcile the referral attribution gap, and set up the paid holdout (geo or campaign-level) with analytics.
  • Test:
  • Trigger the invite prompt after a 5-star or third completed lesson.
  • Run a two-sided $20 credit arm against the current one-sided credit.
  • Threshold:
  • Referred share of new parents rises from 9% to ≥13%.
  • Cost per referred booking parent is ≤$45.
  • 90-day retention is on track for ≥45% at six months.
  • Stop: cost per referred parent exceeds $60, or invites per parent are under 0.36 after six weeks. Then cut referral work to maintenance and move the squad to first-90-day retention of paid cohorts (22% at six months is our biggest leak).

What I'd ask the offsite to decide

  1. Cut paid back to $60k and fund the holdout.
  2. Approve the squad plan and the tutor-referral payout change.
  3. Accept about 42k as the committed target, and have finance model the raise story on retained parents, organic share and bookings per tutor.

I'd revisit paid at the end of Q1 with holdout data. If marginal CAC is ≤$60 in any segment, I'd scale it there.

Claims the judge couldn’t find in the brief

  • 37,000 parents × 0.043 implies about 1,600 referred parents a month, but the 9% attribution implies roughly 150–190; one of these is wrong.

Grades and run details

Decision model 79 · LLM judge 12 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly16%
  • passAddresses the actual decision100%
  • passRespects explicit constraints41%
  • passIdentifies material uncertainty98%
  • partialAvoids unsupported claims49%
  • passProduces the required deliverable86%
  • passFinds the decay behind the growth80%
  • passSees the cross-side effect100%
  • passAnswers the CEO on paid100%
  • passA closed loop, not a channel43%
  • partialThe loop maths holds36%
  • partialProposes tests that could fail40%
Run
Run
#1
API response time
1.7 min
Submitted
2 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Mixed 1

Uses the supplied evidence correctlyMixed

All factual claims about the current situation are taken directly from the supplied context or follow from arithmetic, except one flagged data inconsistency that is presented as a flag rather than an established fact.

Got right 11

Addresses the actual decisionRight

The memo commits early to not tripling paid, proposes a credible alternative target, and states the condition (marginal CAC ≤$60) that would change the call.

Respects explicit constraintsRight

The memo respects the word limit, addresses Rachel and the exec team, does not propose paying for reviews, keeps the commission at 18%, and uses three squads within the six-month window.

Identifies material uncertaintyRight

It names seasonality, competitors, and the CTR drop as possible confounds for the CPA rise, flags a referral data inconsistency, and says a holdout will resolve the paid question.

Avoids unsupported claimsRight

Interpretations and forecasts are clearly labelled (e.g., 'If the same curve holds, my extrapolation'), and no confident claim goes beyond what the evidence supports.

Produces the required deliverableRight

The memo is in the requested format, under 1,300 words, maps the loops, answers the paid plan, and gives three squad plans with tests, thresholds and stop conditions.

Finds the decay behind the growthRight

It shows the 61% of new profiles without bookings or reviews, falling bookings per tutor (9.1 to 6.8), and implies the ~14% decline in sessions per profile by noting 40% session growth against 62% tutor growth, concluding the primary loop is weakening.

Sees the cross-side effectRight

It traces the chain from the $50 tutor-referral bounty to tutor oversupply (+62% vs +20% parents), to thinner bookings per tutor, to new profiles never earning reviews, and proposes moving the payout to the first completed lesson.

Answers the CEO on paidRight

It works out the marginal cost (~$319) from the spend and CPA figures, compares it with the paid parent's $120 LTV, explains that the 2.2 ratio blends sources and averages, and says plainly not to triple.

A closed loop, not a channelRight

It names the organic search → profile → review loop as primary, shows how reviews feed ranking and sessions, and sizes it with the 52% share and 48% retention.

The loop maths holdsRight

The parent referral yield (0.043 per active parent per month) is computed correctly, retention is applied, and each loop gets a plain verdict (compounding/weakening, contributing, decaying).

Proposes tests that could failRight

Every squad proposal includes a numeric threshold, a measurement window (e.g., week 6, 8 weeks), and the action each outcome triggers (stop, pivot, or scale).

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review