Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tutorly. Our CEO, Rachel Dunn, wants to triple paid acquisition for the next two quarters so we reach 50,000 booking parents before the Series B. Before the planning offsite she's asked you for an honest view of how we actually grow. Write a memo of no more than 1,300 words for Rachel and the exec team that: 1. Maps our growth loops in a simple text diagram, says which is primary and why, and whether each is compounding, contributing or decaying, with the numbers. 2. Responds to the plan to triple paid acquisition. 3. Says where our three squads should go for the next two quarters, with the test, threshold and stop condition for each. The pack is below. Not all of it matters equally.

What the model was given8 items: About Tutorly, Twelve months, at the top line, Search and profiles, Where new booking parents come from, and how long they stay, Referral programmes, Paid acquisition, Rachel's note, Constraints
About TutorlyAn online tutoring marketplace: parents book one-to-one lessons with tutors. The average lesson costs $40 and we keep a fixed 18% commission. Each tutor has a public profile page, with reviews from parents.
Twelve months, at the top lineBooking parents (at least one booking a month): 31,000 → 37,000 (+20%). Active tutors: 14,200 → 23,000 (+62%). Organic search sessions to tutor profiles: +40%. Bookings per active tutor a month: 9.1 → 6.8.
Search and profilesProfiles with three or more reviews get 81% of organic profile sessions; profiles with none get 4%. Of tutor profiles created in the last six months, 61% had no booking within 60 days, so they have no reviews. Click-through from search results on our top 200 keywords fell from 6.1% to 5.0% since March, after changes to the results pages.
Where new booking parents come from, and how long they stayOrganic search: 52% of new booking parents; 48% still booking after six months; net lifetime value $260. Parent referrals: 9%; 51% after six months; net lifetime value $275. Paid search and social: 39%; 22% after six months; net lifetime value $120. Across all sources, net lifetime value averages $210.
Referral programmesParents: $20 of lesson credit for each referred parent who books. Each active parent sends 0.31 invites a month, and 14% of invites become booking parents. Tutors: $50 for each referred tutor who completes onboarding. 44% of new tutors last year came through tutor referrals.
Paid acquisitionAverage cost per new booking parent over the year: $95. Last quarter we raised spend from $60,000 to $90,000 a month, and the cost per new booking parent rose from $82 to $109.
Rachel's note“Our LTV to CAC is 2.2. Every dollar we don't put into paid is growth we're leaving on the table. Triple it to $270k a month and we hit 50,000 parents before the raise.”
ConstraintsReviews can only be left after a completed, paid lesson. Our policy, and consumer protection rules in our main markets, forbid paying or rewarding anyone for reviews. The commission stays at 18%. Three squads are available for the next two quarters; the raise is about six months away.
What a strong answer doesThe answer key the graders mark against

Maps four loops: content (profiles and reviews rank, parents book, bookings create reviews, profiles rank better), parent referral, tutor referral, and paid. Names content as primary: organic is the biggest source of new booking parents (52%) and among the stickiest (48% at six months, $260). Finds that it's decaying behind 40% session growth. Profiles grew 62% against 40% sessions, so sessions per profile fell about 14%. 61% of new profiles never get a booking, so they never get reviews, and profiles without reviews get 4% of sessions. Bookings per tutor fell from 9.1 to 6.8. Traces the cross-side cause: tutor supply growing about three times faster than demand (+62% against +20%), much of it from the $50 tutor-referral programme (44% of new tutors), which spreads bookings too thin for new tutors to earn the reviews the loop runs on. Treats the click-through drop (6.1% to 5.0%) as a separate, external hit. Sizes the parent referral loop at about 0.043 (0.31 × 14%): small but high quality. Answers Rachel with marginal cost: the extra $30,000 a month bought about 94 more parents (60,000 ÷ 82 ≈ 732 → 90,000 ÷ 109 ≈ 826), roughly $319 each, against $120 of value for a paid parent. The 2.2 ratio blends sources and uses the average cost, so tripling would burn cash on parents who mostly leave. Sends the squads to: getting new tutors their first booking and review (for example, routing demand in under-supplied subjects to new tutors, with a threshold such as the share of new profiles booked within 60 days rising from 39% to 55%); concentrating search value on reviewed profiles (keeping unreviewed profiles out of the index, and a fix for the click-through drop); and pausing or targeting the tutor-referral bounty while improving the parent loop. Each comes with a stop condition, and the memo is honest about what 50,000 parents by the raise would really take, without rewarding reviews.

Critical failures (cap the score at 40)
  • Proposes paying, crediting or otherwise rewarding parents or tutors for reviews
  • Treats the 40% growth in organic sessions as proof the content loop is healthy
Task

Staff-level · v1.0 · synthetic · marketplace, network effects, SEO, unit economics, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit10 of 12 checks passed

The memo contains an unsupported claim (14,000+ unreviewed profiles) and misstates the marginal CAC as $109, requiring correction before a PM could rely on it without re-checking the evidence.

Output

MEMORANDUM

TO: Rachel Dunn, CEO; Executive Team FROM: Staff Product Manager DATE: October 24, 2023 SUBJECT: Growth Engine Analysis, Paid Acquisition Response, and H1 Squad Allocations

*

Executive Summary

Tripling paid acquisition to $270k/month will not get us to 50,000 sustained booking parents for Series B. It relies on a metric illusion: our blended LTV is $210, but Paid LTV is only $120, while our marginal Paid CAC has already climbed to $109. Tripling spend will push marginal CAC well above LTV, burning cash to acquire low-intent parents who churn at more than double the rate of organic parents (22% vs. 48% 6-month retention).

Meanwhile, our true growth engine—the organic SEO/Review loop—is choking on a tutor supply glut that we are actively subsidizing. We must reallocate our squads to repair organic search distribution and tutor activation, turning our primary compounding loop back on before going out for Series B.

---

1. Tutorly’s Growth Engine: Loop Mapping & Health

``` [Loop A: SEO / Review Loop (Primary Engine)] Parent books lesson ──> Completes paid lesson ──> Leaves review ▲ │ │ ▼ Organic search session ◄── Profile ranks higher / gains CTR

[Loop B: Parent Referral Loop (Secondary)] Active parent ──> Shares invite ($20 credit) ──> New parent books

[Loop C: Paid Acquisition (Linear/Diminishing)] Ad spend ($) ──> Paid clicks ──> First-time booking parent (High churn)

[Counter-Loop: Supply Glut / Dilution] $50 Tutor referral bounty ──> +62% Tutors ──> 61% zero-booking profiles │ │ ▼ ▼ Search dilution & drop in tutor earnings ◄── Suppressed profile liquidity ```

#### Loop Health & Classification

LoopClassificationCurrent State & Metrics
:---:---:---
A. SEO / Profile Reviews (Primary)Decaying Compounding52% of new parents; $260 LTV; 48% M6 retention. Profiles with $\ge$3 reviews capture 81% of organic traffic, but 61% of new tutors get zero bookings in 60 days (0 reviews), starving the loop. Additionally, SERP CTR dropped from 6.1% to 5.0% post-redesign.
B. Parent ReferralsContributing9% of new parents; $275 LTV; 51% M6 retention. Each parent generates $0.31 \times 0.14 = 0.043$ new parents/month. Highly healthy unit economics, but lacks viral velocity ($K \ll 1.0$) to compound on its own.
C. Paid AcquisitionDecaying Linear39% of new parents; $120 LTV; 22% M6 retention. Marginal CAC rose from $82 to $109 when scaling to $90k/mo. Lacks compounding feedback; functions as a leaky bucket.
D. Tutor Supply BountyValue-Destroying44% of new tutors; $50 bounty. Supply surged +62% while demand grew +20%, depressing monthly bookings per tutor from 9.1 to 6.8 (-25%) and creating 14,000+ unreviewed "ghost" profiles.

Why Loop A is Primary: Organic search generates 52% of our demand at superior economics ($260 LTV vs. $120 Paid). It is our only structurally compounding loop: bookings generate reviews, which improve search ranking and page conversion, generating more bookings. However, it is currently decaying due to supply misallocation and SERP degradation.

---

2. Evaluating the Plan to Triple Paid Acquisition

Rachel's proposal states: "Our LTV to CAC is 2.2... Triple it to $270k a month and we hit 50,000 parents."

This relies on two critical errors:

1. The Blended LTV Fallacy: Rachel calculates LTV:CAC using blended LTV ($210) against historical blended CAC ($95) ($210 / $95 = 2.21$). But channel economics reveal a different reality: * Paid LTV is $120, driven by dismal 6-month retention (22%). * When we increased monthly spend from $60,000 to $90,000, our average Paid CAC rose from $82 to $109. * The marginal cost of those incremental acquisitions was: $$\text{Incremental CAC} = \frac{\$90,000 - \$60,000}{\Delta \text{Acquisitions}} = \frac{\$30,000}{\left(\frac{90,000}{109}\right) - \left(\frac{60,000}{82}\right)} = \frac{\$30,000}{825 - 731} \approx \$319$$ * At $109 average CAC against $120 LTV, current paid spend is barely breaking even ($1.1\times$). At $270k/month, Paid CAC will exceed Paid LTV on a fully loaded basis. 2. Quality of Growth for Series B: Pushing paid acquisition will inflate 30-day top-line numbers with unprofitable users who churn before the Series B diligence team reviews our cohorts. Series B investors will immediately spot cohorts decaying to 22% retention and discount our valuation.

Recommendation: Cap paid acquisition at $60,000/month (where CAC was sustainable at $82) and reallocate capital toward fixing marketplace liquidity and discovery.

---

3. Squad Allocation for the Next Two Quarters

To reach 50,000 booking parents sustainably, we must fix our core SEO loop and activate idle supply. We allocate our three squads as follows:

``` ┌─────────────────────────────────────────────────────────────┐ │ Squad 1: Marketplace Liquidity & Cold Start │ │ Objective: Solve the 61% unbooked tutor barrier │ ├─────────────────────────────────────────────────────────────┤ │ Squad 2: Search Experience & Organic Distribution │ │ Objective: Reverse SERP CTR drop from 5.0% to >6.5% │ ├─────────────────────────────────────────────────────────────┤ │ Squad 3: Parent Retention & Referral Compounding │ │ Objective: Expand organic retention and boost referral K-factor│ └─────────────────────────────────────────────────────────────┘ ```

#### Squad 1: Marketplace Liquidity & Cold Start * Problem: 61% of new tutors get no bookings in 60 days, yielding 0 reviews. Unreviewed tutors capture only 4% of traffic. We are paying $50 to onboard tutors into dead ends, depressing utilization to 6.8 bookings/tutor. * Focus: Route high-intent parent demand to unreviewed, vetted tutors using new-tutor introductory match discounts (funded by retiring the $50 tutor referral bounty) and smart onboarding guarantees. * Test: Launch a "First Lesson Booking Guarantee" matching algorithm on category pages, pairing new vetted tutors with price-sensitive parents. * Threshold: $\ge$35% of newly onboarded tutors secure their first completed lesson within 14 days (up from <15%), achieving $\ge$1 review within 21 days. * Stop Condition: If introductory match initiatives do not lift first-booking rates within 60 days without reducing existing tutor earnings below 6.5 bookings/month, halt automated discounting and shift to a curated intake cap.

#### Squad 2: Search Experience & Organic Distribution * Problem: Organic search profile sessions grew +40%, but CTR on top 200 keywords plummeted from 6.1% to 5.0% (-18% relative) following recent SERP changes. This represents hundreds of lost bookings per week. * Focus: Audit and roll back damaging SERP redesigns; implement rich review schema markups; optimize localized long-tail tutor directory pages (e.g., "Math Tutors in [City]"). * Test: A/B test SERP layout rollback vs. a structured review-first layout displaying verified subject ratings and badge endorsements compliant with consumer review regulations. * Threshold: Restore top-200 keyword CTR to $\ge$6.2% within 8 weeks, yielding $\ge$15% incremental organic booking volume. * Stop Condition: If technical SERP changes fail to move CTR by $\ge$0.5 percentage points after 6 weeks of clean indexation, freeze front-end adjustments and redirect squad to on-site landing page conversion and booking-funnel friction.

#### Squad 3: Parent Retention & Referral Compounding * Problem: Paid parent retention is poor (22%), while referral parents yield $275 LTV and 51% retention. However, referral volume is constrained (only 0.31 invites/parent/month). * Focus: Embed referral prompts directly after high-satisfaction milestones (e.g., 5-star lesson completions) and introduce parent-to-parent study group mechanics rather than generic link-sharing. * Test: Introduce dynamic milestone-based referral prompts (post-5-star rating) offering dual-sided $25 credits, combined with an automated re-booking reminder cadence for churn-risk cohorts. * Threshold: Increase invite velocity from 0.31 to 0.50 invites/active parent/month, sustaining an invite-to-booking rate $\ge$14% (lifting referral share of new parents from 9% to 15%). * Stop Condition: If CAC via referral incentives exceeds $45 per booking parent, or invite rate does not cross 0.40 within 60 days, kill the programmatic milestone incentives and pivot to lesson package re-engagement.

---

Series B Target Trajectory

MetricCurrentRachel's Paid PlanProposed Plan (6 Mo)
:---:---:---:---
Monthly Paid Spend$90,000$270,000$60,000
Blended M6 Retention41%28% (Paid heavy)47%
Bookings / Active Tutor6.86.28.5
Active Booking Parents37,000~48,000 (Leaky)46,500 (Durable)
Series B NarrativeUnprofitable growthHigh burn, low LTVHigh margin, compounding engine

By fixing our cold-start liquidity and reversing the SERP CTR loss, we will approach 47,000 highly retained, profitable booking parents. This presents a bulletproof Series B narrative: expanding gross margins, compounding organic discovery, and capital efficiency.

Claims the judge couldn’t find in the brief

  • marginal Paid CAC has already climbed to $109
  • creating 14,000+ unreviewed 'ghost' profiles

Grades and run details

Decision model 79 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly3%
  • passAddresses the actual decision99%
  • partialRespects explicit constraints19%
  • passIdentifies material uncertainty24%
  • failAvoids unsupported claims27%
  • passProduces the required deliverable76%
  • passFinds the decay behind the growth46%
  • passSees the cross-side effect100%
  • passAnswers the CEO on paid100%
  • partialA closed loop, not a channel45%
  • partialThe loop maths holds20%
  • passProposes tests that could fail46%
Run
Run
#1
API response time
34 s
Submitted
2 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 1

Avoids unsupported claimsWrong

Presents '14,000+ unreviewed ghost profiles' as fact without support, and states 'marginal Paid CAC has already climbed to $109' as a confident claim when the evidence only gives average CAC.

Mixed 1

Uses the supplied evidence correctlyMixed

The claim '14,000+ unreviewed ghost profiles' is invented and not supported by the supplied context, and 'marginal Paid CAC has already climbed to $109' misstates the evidence (the $109 is average, not marginal).

Got right 10

Addresses the actual decisionRight

Commits clearly to not tripling paid, caps spend at $60k, and allocates three squads with explicit stop conditions that would change the call.

Respects explicit constraintsRight

Memo is within 1,300 words, addresses Rachel and exec team, does not propose paying for reviews, and keeps commission at 18%.

Identifies material uncertaintyRight

Each squad proposal includes a stop condition that names what result would change the approach, effectively identifying material uncertainties.

Produces the required deliverableRight

The memo is complete, in the requested form, within length, and usable by the exec team with light edits.

Finds the decay behind the growthRight

Shows sessions per profile decline (implied by +40% sessions vs +62% tutors), 61% unbooked new profiles, and falling bookings per tutor, concluding the primary loop is decaying.

Sees the cross-side effectRight

Traces the $50 tutor bounty to oversupply (+62% tutors vs +20% parents), to thinner bookings (9.1→6.8), to new profiles without reviews, and proposes retiring the bounty.

Answers the CEO on paidRight

Computes marginal cost (~$319) vs paid LTV ($120), explains the 2.2 ratio blends sources and averages, and firmly recommends not tripling paid.

A closed loop, not a channelRight

Names the SEO/review loop as primary closed loop, grounded in organic being the largest source with highest retention, and sizes it with 52% share and $260 LTV.

The loop maths holdsRight

Parent referral yield (0.043) is computed correctly from 0.31 invites and 14% conversion; each loop gets a plain verdict (decaying compounding, contributing, decaying linear).

Proposes tests that could failRight

Each squad has a numeric threshold, a measurement window (e.g., 60 days, 8 weeks), and a stop condition that triggers a specific pivot.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review