Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tutorly. Our CEO, Rachel Dunn, wants to triple paid acquisition for the next two quarters so we reach 50,000 booking parents before the Series B. Before the planning offsite she's asked you for an honest view of how we actually grow. Write a memo of no more than 1,300 words for Rachel and the exec team that: 1. Maps our growth loops in a simple text diagram, says which is primary and why, and whether each is compounding, contributing or decaying, with the numbers. 2. Responds to the plan to triple paid acquisition. 3. Says where our three squads should go for the next two quarters, with the test, threshold and stop condition for each. The pack is below. Not all of it matters equally.

What the model was given8 items: About Tutorly, Twelve months, at the top line, Search and profiles, Where new booking parents come from, and how long they stay, Referral programmes, Paid acquisition, Rachel's note, Constraints
About TutorlyAn online tutoring marketplace: parents book one-to-one lessons with tutors. The average lesson costs $40 and we keep a fixed 18% commission. Each tutor has a public profile page, with reviews from parents.
Twelve months, at the top lineBooking parents (at least one booking a month): 31,000 → 37,000 (+20%). Active tutors: 14,200 → 23,000 (+62%). Organic search sessions to tutor profiles: +40%. Bookings per active tutor a month: 9.1 → 6.8.
Search and profilesProfiles with three or more reviews get 81% of organic profile sessions; profiles with none get 4%. Of tutor profiles created in the last six months, 61% had no booking within 60 days, so they have no reviews. Click-through from search results on our top 200 keywords fell from 6.1% to 5.0% since March, after changes to the results pages.
Where new booking parents come from, and how long they stayOrganic search: 52% of new booking parents; 48% still booking after six months; net lifetime value $260. Parent referrals: 9%; 51% after six months; net lifetime value $275. Paid search and social: 39%; 22% after six months; net lifetime value $120. Across all sources, net lifetime value averages $210.
Referral programmesParents: $20 of lesson credit for each referred parent who books. Each active parent sends 0.31 invites a month, and 14% of invites become booking parents. Tutors: $50 for each referred tutor who completes onboarding. 44% of new tutors last year came through tutor referrals.
Paid acquisitionAverage cost per new booking parent over the year: $95. Last quarter we raised spend from $60,000 to $90,000 a month, and the cost per new booking parent rose from $82 to $109.
Rachel's note“Our LTV to CAC is 2.2. Every dollar we don't put into paid is growth we're leaving on the table. Triple it to $270k a month and we hit 50,000 parents before the raise.”
ConstraintsReviews can only be left after a completed, paid lesson. Our policy, and consumer protection rules in our main markets, forbid paying or rewarding anyone for reviews. The commission stays at 18%. Three squads are available for the next two quarters; the raise is about six months away.
What a strong answer doesThe answer key the graders mark against

Maps four loops: content (profiles and reviews rank, parents book, bookings create reviews, profiles rank better), parent referral, tutor referral, and paid. Names content as primary: organic is the biggest source of new booking parents (52%) and among the stickiest (48% at six months, $260). Finds that it's decaying behind 40% session growth. Profiles grew 62% against 40% sessions, so sessions per profile fell about 14%. 61% of new profiles never get a booking, so they never get reviews, and profiles without reviews get 4% of sessions. Bookings per tutor fell from 9.1 to 6.8. Traces the cross-side cause: tutor supply growing about three times faster than demand (+62% against +20%), much of it from the $50 tutor-referral programme (44% of new tutors), which spreads bookings too thin for new tutors to earn the reviews the loop runs on. Treats the click-through drop (6.1% to 5.0%) as a separate, external hit. Sizes the parent referral loop at about 0.043 (0.31 × 14%): small but high quality. Answers Rachel with marginal cost: the extra $30,000 a month bought about 94 more parents (60,000 ÷ 82 ≈ 732 → 90,000 ÷ 109 ≈ 826), roughly $319 each, against $120 of value for a paid parent. The 2.2 ratio blends sources and uses the average cost, so tripling would burn cash on parents who mostly leave. Sends the squads to: getting new tutors their first booking and review (for example, routing demand in under-supplied subjects to new tutors, with a threshold such as the share of new profiles booked within 60 days rising from 39% to 55%); concentrating search value on reviewed profiles (keeping unreviewed profiles out of the index, and a fix for the click-through drop); and pausing or targeting the tutor-referral bounty while improving the parent loop. Each comes with a stop condition, and the memo is honest about what 50,000 parents by the raise would really take, without rewarding reviews.

Critical failures (cap the score at 40)
  • Proposes paying, crediting or otherwise rewarding parents or tutors for reviews
  • Treats the 40% growth in organic sessions as proof the content loop is healthy
Task

Staff-level · v1.0 · synthetic · marketplace, network effects, SEO, unit economics, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit5 of 12 checks passed

Contains a factual error on organic retention (46% vs 48%), mislabels the primary loop as compounding, and miscalculates marginal paid cost. A PM would need to re-check evidence and redo the loop and paid analysis.

Output

MEMORANDUM

TO: Rachel Dunn, CEO, and Executive Team FROM: Staff Product Manager, Tutorly DATE: October 24, 2023 SUBJECT: Growth Analysis, Paid Acquisition Plan Review, and Q1–Q2 Squad Roadmap

*

Executive Summary

To reach 50,000 booking parents before the Series B, we must separate our top-line metrics from the health of our marketplace engine. While growing booking parents from 37,000 to 50,000 (+35%) in two quarters is mathematically possible, tripling paid acquisition to achieve it will accelerate churn, drain cash, and mask structural supply-side liquidity issues that threaten our Series B valuation.

This memo maps our true growth mechanics, evaluates the paid acquisition proposal, and redeploys our three engineering squads to fix marketplace liquidity before we scale demand.

---

Part 1: Mapping Our Growth Loops

Tutorly operates on three core loops. The numbers reveal a marketplace suffering from top-heavy demand generation sitting atop a leaking supply foundation.

``` [1. SEO & Tutor Profile Loop] (PRIMARY) Tutors onboard (44% via referral) --> Create Public Profile Pages --> Indexed by Search Engines --> Organic Search Sessions (52% of parents) --> Completed Lessons & Reviews (81% traffic to profiles w/ 3+ reviews) --> Higher Search Rankings & Marketplace Trust --> More Parent Bookings --> Tutors earn & refer peers (Loop Compounding)

[2. Parent Referral Loop] (CONTRIBUTING) Booking Parents --> Send Invites (0.31/mo) --> Referred Parents Book (14% conversion) --> High Retention (51% at 6 months, $275 LTV) --> More Parents --> More Invites (Loop Compounding, but low volume at 9% share)

[3. Paid Acquisition Loop] (DECAYING AT SCALE) Capital ($90k/mo) --> Paid Search & Social (39% of parents) --> New Booking Parents --> Low Retention (22% at 6 months, $120 LTV) --> Blended LTV:CAC drops from 2.2 to 1.1 at scale --> Capital Inefficient (Loop Decaying) ```

#### Loop Status and Metrics

1. SEO & Tutor Profile Loop (Primary | Compounding): * Status: Compounding at the top, but restricted by a liquidity bottleneck at the bottom. * The Numbers: Organic search drives 52% of new booking parents with a healthy 46% 6-month retention and $260 net LTV. However, active tutors grew +62% (14.2k to 23k), while bookings per tutor dropped from 9.1 to 6.8/month. Because 61% of new tutors get zero bookings in 60 days, they never reach 3+ reviews. Since profiles with 3+ reviews capture 81% of organic profile traffic (vs. 4% for zero reviews), our SEO asset is trapped behind unactivated supply. Furthermore, CTR on our top 200 keywords dropped from 6.1% to 5.0% due to recent SERP changes.

2. Parent Referral Loop (Contributing | Compounding): * Status: Healthy unit economics, but low absolute volume. * The Numbers: Generates 9% of new booking parents. Each active parent sends 0.31 invites/month with a 14% conversion rate. Retains exceptionally well (51% after six months) with a net LTV of $275. It compounds efficiently but lacks the velocity to move the top line alone.

3. Paid Acquisition Loop (Secondary | Decaying): * Status: Decaying rapidly at marginal scale. * The Numbers: Drives 39% of new booking parents, but retention crashes to 22% at six months (net LTV of $120). When we raised spend from $60k to $90k/month last quarter, the cost per new booking parent ($CAC$) jumped from $82 to $109.

---

Part 2: Response to the Plan to Triple Paid Acquisition

Rachel, your note argues: "Our LTV to CAC is 2.2. Every dollar we don't put into paid is growth we're leaving on the table. Triple it to $270k a month and we hit 50,000 parents before the raise."

An honest view of the data shows this plan will undermine our Series B narrative. Here is why:

  1. Blended LTV is a dangerous illusion: Our blended LTV is $210 against a blended CAC of $95 (yielding 2.2). But paid acquisition cohorts have an LTV of only $120. When we pushed spend to $90k/month last quarter, marginal CAC rose to $109. Tripling spend to $270k/month will push marginal CAC past $160–$180, well above the $120 LTV of paid users. We would be acquiring customers at a net loss ($40+ loss per paid parent).
  2. Exacerbating the supply-side liquidity crisis: Active tutors are already seeing their monthly bookings fall from 9.1 to 6.8 (-25%) because supply (23,000) is growing faster than demand (37,000 parents). Pouring cheap, low-intent paid traffic into the top of the funnel without fixing tutor activation will worsen tutor churn, starving our primary SEO loop (which relies on active tutors getting reviews).
  3. Series B Risk: Investors scrutinize cohort retention and payback periods. A 22% six-month retention rate on 50,000 paid parents signals a leaky bucket that will cause churn spikes right after our Series B closes.

Conclusion on Paid: We should cap paid acquisition at a disciplined $90k/month, focusing strictly on high-intent brand and bottom-funnel keywords where marginal CAC remains under $85.

---

Part 3: Q1–Q2 Squad Roadmap

To reach 50,000 booking parents sustainably for the Series B, we must unlock supply liquidity and fix our primary organic engine. We will deploy our three squads to specific operational bottlenecks.

---

#### Squad 1: Search & Discovery Squad * Mission: Reverse the CTR decline on organic search results and optimize tutor profile visibility to capture existing search demand without increasing paid spend. * Test: * Hypothesis: Redesigning the search results page (SERP) layout and introducing structured review snippets will recover our CTR from 5.0% back to baseline (6.1%) and increase profile-to-booking conversion. * Action: A/B test a revised SERP layout prioritizing tutor responsiveness and verified subject badges. * Threshold (Go/No-Go): Recovers organic traffic conversion by $\ge$ 15% within 45 days. * Stop Condition: If CTR remains flat or declines further after two iterative algorithm tests, pivot the squad to programmatic landing page generation for long-tail tutoring subjects.

---

#### Squad 2: Tutor Activation Squad * Mission: Ensure new tutors get their first booking within 30 days, breaking the cold-start problem so they accumulate the 3 reviews required to unlock 81% of organic profile traffic. * Test: * Hypothesis: Lowering the barrier to initial trial lessons (e.g., introducing a standardized 30-minute introductory lesson priced at $15, where Tutorly waives commission) will drive the first completed lesson within 30 days for >50% of new tutors. * Action: Launch "Introductory Lesson Pass" for new tutors to accelerate time-to-first-review. * Threshold (Go/No-Group): Reduces the 60-day zero-booking rate from 61% down to $\le$ 40%. * Stop Condition: If zero-booking rates do not drop below 50% after 60 days, sunset the trial pass and test supply-capping (restricting new tutor onboarding in over-supplied subjects to concentrate demand).

---

#### Squad 3: Parent Retention & Referral Squad * Mission: Improve early-stage retention for paid-acquired parents (who drop to 22% at 6 months) and optimize the parent referral loop. * Test: * Hypothesis: Onboarding paid-acquired parents with an automated, subject-specific tutor matching wizard rather than self-serve browsing will increase 3-month retention by 20%. * Action: Implement an AI/rules-based tutor matching onboarding flow for all non-organic acquisition channels. * Threshold (Go/No-Go): Increases 90-day retention of paid cohorts from current baseline to $\ge$ 35%. * Stop Condition: If paid cohort retention fails to improve, reduce paid spend further to $60k/month and reallocate budget to scaling our Parent Referral program (which yields a $275 LTV and 51% retention).

---

Summary Table: Q1–Q2 Roadmap

SquadFocus AreaCore TestSuccess ThresholdStop Condition
:---:---:---:---:---
1. Search & DiscoverySERP CTR & Profile ConversionSERP layout redesign & review snippets+15% conversion lift in 45 daysPivot to programmatic landing pages
2. Tutor Activation60-Day Zero-Booking Rate"Intro Pass" (zero-commission first lesson)60-day zero bookings drops from 61% to $\le$40%Implement strict supply-capping
3. Retention & ReferralPaid Cohort RetentionGuided tutor-matching onboarding flowPaid cohort 90-day retention $\ge$35%Cut paid spend to $60k; scale referrals

Final Word

Rachel, we can hit 50,000 booking parents before the Series B, but not by brute-forcing paid channels that destroy our unit economics. By fixing tutor activation (Squad 2) and recovering search efficiency (Squad 1), we supercharge our compounding SEO loop. That is the growth story that commands a premium Series B valuation.

Claims the judge couldn’t find in the brief

  • Organic search retention is 46% at six months

Grades and run details

Decision model 50 · LLM judge 7 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly28%
  • passAddresses the actual decision95%
  • failRespects explicit constraints5%
  • partialIdentifies material uncertainty25%
  • failAvoids unsupported claims38%
  • passProduces the required deliverable47%
  • partialFinds the decay behind the growth25%
  • passSees the cross-side effect88%
  • partialAnswers the CEO on paid69%
  • partialA closed loop, not a channel56%
  • failThe loop maths holds25%
  • passProposes tests that could fail26%
Run
Run
#1
Time to output
12 s
Submitted
2 Oct 2026

Check by check

Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Got wrong 5

Uses the supplied evidence correctlyWrong

States organic retention as 46% instead of the correct 48% from the brief, a factual error about the current situation.

Finds the decay behind the growthWrong

Labels the primary SEO loop as 'Compounding' and does not conclude it is decaying, despite having the numbers to show weakening per profile.

Answers the CEO on paidWrong

Treats the $109 average cost as the marginal cost and does not compute the true marginal cost (~$319) of the last spend increment; the marginal analysis is incorrect.

A closed loop, not a channelWrong

Does not compute a yield or coefficient for the primary loop with retention applied; only describes the loop qualitatively.

The loop maths holdsWrong

No loop yields, coefficients or payback periods are calculated; the verdicts (compounding/contributing/decaying) are stated without supporting maths.

Mixed 2

Respects explicit constraintsMixed

Memo is within 1,300 words, addressed to Rachel and exec team, respects no-paying-for-reviews rule, and uses three squads.

Avoids unsupported claimsMixed

Forecasts and hypotheses are clearly labelled; the only error is a factual mistake, not an unsupported interpretation.

Got right 5

Addresses the actual decisionRight

Commits to capping paid at $90k/month and gives a squad roadmap with clear conditions that would change the call.

Identifies material uncertaintyRight

Each squad proposal includes a stop condition that names what result would change the course of action.

Produces the required deliverableRight

Delivers a complete memo with loops, paid response, and three squad plans with tests, thresholds and stop conditions.

Sees the cross-side effectRight

Connects tutor oversupply (+62% tutors vs +20% parents) and the tutor referral bounty to falling bookings per tutor and the review bottleneck.

Proposes tests that could failRight

Each squad test has a numeric threshold, a measurement window (45, 60, 90 days), and a stop condition with a clear next action.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review