Tasks / Experiment

Find the growth loop

Can the model find a product's real growth loop, show whether it compounds, and say which lever to pull?

Measures the modelTask type v1.0 · 2 tasksLast changed 2 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 33% were usable with at most a quick edit.

Reliably right

  1. Sees the cross-side effect100% pass
    It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Addresses the actual decision96% pass
    The memo commits early to putting both engineers on idea 4, names the primary loop and its compounding status, and specifies what results would change the call (kill thresholds, quarter-end loop gain).
    Sonnet 5.5 · API · The badge on every form
  3. Produces the required deliverable96% pass
    The memo answers all parts of the brief (primary loop, compounding, engineer allocation, success measurement) in a usable form for the Head of Growth.
    Sonnet 5.5 · API · The badge on every form

Where it slips

  1. The loop maths holds46% pass
    The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.
    GPT-6.1 Sol · API · Growing on the surface, decaying underneath
  2. Uses the supplied evidence correctly57% pass
    The claim that the base settles at 10,700 creators is unsupported by the pack's arithmetic, and the claim that cost per sign-up usually rises with spend is not in the supplied evidence.
    Opus 5.5 · Claude · The badge on every form
  3. Avoids unsupported claims59% pass
    Presents the 10,700 equilibrium and the rising cost-per-sign-up claim as facts without labelling them as hypotheses or supporting them from the pack.
    Opus 5.5 · Claude · The badge on every form

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Tutorly. Our CEO, Rachel Dunn, wants to triple paid acquisition for the next two quarters so we reach 50,000 booking parents before the Series B. Before the planning offsite she's asked you for an honest view of how we actually grow. Write a memo of no more than 1,300 words for Rachel and the exec team that: 1. Maps our growth loops in a simple text diagram, says which is primary and why, and whether each is compounding, contributing or decaying, with the numbers. 2. Responds to the plan to triple paid acquisition. 3. Says where our three squads should go for the next two quarters, with the test, threshold and stop condition for each. The pack is below. Not all of it matters equally.

What the model was given8 items: About Tutorly, Twelve months, at the top line, Search and profiles, Where new booking parents come from, and how long they stay, Referral programmes, Paid acquisition, Rachel's note, Constraints
About TutorlyAn online tutoring marketplace: parents book one-to-one lessons with tutors. The average lesson costs $40 and we keep a fixed 18% commission. Each tutor has a public profile page, with reviews from parents.
Twelve months, at the top lineBooking parents (at least one booking a month): 31,000 → 37,000 (+20%). Active tutors: 14,200 → 23,000 (+62%). Organic search sessions to tutor profiles: +40%. Bookings per active tutor a month: 9.1 → 6.8.
Search and profilesProfiles with three or more reviews get 81% of organic profile sessions; profiles with none get 4%. Of tutor profiles created in the last six months, 61% had no booking within 60 days, so they have no reviews. Click-through from search results on our top 200 keywords fell from 6.1% to 5.0% since March, after changes to the results pages.
Where new booking parents come from, and how long they stayOrganic search: 52% of new booking parents; 48% still booking after six months; net lifetime value $260. Parent referrals: 9%; 51% after six months; net lifetime value $275. Paid search and social: 39%; 22% after six months; net lifetime value $120. Across all sources, net lifetime value averages $210.
Referral programmesParents: $20 of lesson credit for each referred parent who books. Each active parent sends 0.31 invites a month, and 14% of invites become booking parents. Tutors: $50 for each referred tutor who completes onboarding. 44% of new tutors last year came through tutor referrals.
Paid acquisitionAverage cost per new booking parent over the year: $95. Last quarter we raised spend from $60,000 to $90,000 a month, and the cost per new booking parent rose from $82 to $109.
Rachel's note“Our LTV to CAC is 2.2. Every dollar we don't put into paid is growth we're leaving on the table. Triple it to $270k a month and we hit 50,000 parents before the raise.”
ConstraintsReviews can only be left after a completed, paid lesson. Our policy, and consumer protection rules in our main markets, forbid paying or rewarding anyone for reviews. The commission stays at 18%. Three squads are available for the next two quarters; the raise is about six months away.
What a strong answer doesThe answer key the graders mark against

Maps four loops: content (profiles and reviews rank, parents book, bookings create reviews, profiles rank better), parent referral, tutor referral, and paid. Names content as primary: organic is the biggest source of new booking parents (52%) and among the stickiest (48% at six months, $260). Finds that it's decaying behind 40% session growth. Profiles grew 62% against 40% sessions, so sessions per profile fell about 14%. 61% of new profiles never get a booking, so they never get reviews, and profiles without reviews get 4% of sessions. Bookings per tutor fell from 9.1 to 6.8. Traces the cross-side cause: tutor supply growing about three times faster than demand (+62% against +20%), much of it from the $50 tutor-referral programme (44% of new tutors), which spreads bookings too thin for new tutors to earn the reviews the loop runs on. Treats the click-through drop (6.1% to 5.0%) as a separate, external hit. Sizes the parent referral loop at about 0.043 (0.31 × 14%): small but high quality. Answers Rachel with marginal cost: the extra $30,000 a month bought about 94 more parents (60,000 ÷ 82 ≈ 732 → 90,000 ÷ 109 ≈ 826), roughly $319 each, against $120 of value for a paid parent. The 2.2 ratio blends sources and uses the average cost, so tripling would burn cash on parents who mostly leave. Sends the squads to: getting new tutors their first booking and review (for example, routing demand in under-supplied subjects to new tutors, with a threshold such as the share of new profiles booked within 60 days rising from 39% to 55%); concentrating search value on reviewed profiles (keeping unreviewed profiles out of the index, and a fix for the click-through drop); and pausing or targeting the tutor-referral bounty while improving the parent loop. Each comes with a stop condition, and the memo is honest about what 50,000 parents by the raise would really take, without rewarding reviews.

Critical failures (cap the score at 40)
  • Proposes paying, crediting or otherwise rewarding parents or tutors for reviews
  • Treats the 40% growth in organic sessions as proof the content loop is healthy
Task

Staff-level · v1.0 · synthetic · marketplace, network effects, SEO, unit economics, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 12 checks passed

The analysis is complete and evidence-based; a quick edit to clarify the content loop's decaying verdict would make it fully aligned, but the core reasoning is sound and actionable.

Output

To: Rachel Dunn and Executive Team From: Staff PM, Tutorly Subject: Grow booking demand, not acquisition spend Planning horizon: Next two quarters

Recommendation

Do not commit to $270,000/month in paid acquisition. Our primary growth engine is organic discovery of proven tutors, reinforced by completed lessons and reviews. That engine is growing, but its distribution is concentrated and its conversion into bookings is weakening. Meanwhile, tutor supply is growing much faster than demand, and paid acquisition has deteriorating marginal economics.

Put our three squads on organic discovery, demand-led tutor activation, and parent referrals. Keep paid acquisition as a tightly gated contributor—not the engine assumed to deliver 50,000 booking parents.

1. How we actually grow

```text PRIMARY: Organic discovery / reputation Search → tutor profile → booking → completed paid lesson → genuine review → stronger discovery/conversion → more bookings

Parent referrals Booking parent → invitation → new booking parent → invitations

Tutor referrals Active tutor → referred tutor → onboarding → additional supply → bookings only if matching parent demand exists

Paid acquisition / reinvestment Spend → new booking parent → repeat lessons → commission → acquisition budget → more spend ```

Organic/reputation: our primary compounding mechanism, but uneven and constrained. Organic supplies 52% of new booking parents, versus paid’s 39% and referrals’ 9%. These parents have 48% six-month retention and $260 net LTV. Completed lessons generate reviews, which can improve subsequent discovery and conversion without buying every next customer.

Profiles with at least three reviews capture 81% of organic profile sessions; unreviewed profiles capture only 4%. This supports the mechanism, but does not establish that reviews alone cause the traffic difference.

The warning signs are substantial:

  • Organic profile sessions rose 40%, but booking parents rose only 20%.
  • Search CTR fell from 6.1% to 5.0%, an 18% relative decline.
  • 61% of recently created profiles had no booking within 60 days—and therefore no reviews.

This is a reinforcing loop concentrated in established profiles, not proof that adding more profiles compounds demand.

Parent referrals: contributing, not yet demonstrably self-sustaining. Each active parent generates:

`0.31 invites/month × 14% conversion = 0.0434 new booking parents/month`

That is 4.34 new booking parents per 100 active parents per month. Even six fully active months produce only 0.26 referred booking parents per parent. We lack lifetime active-month data to calculate a full reproduction rate.

Referral customers are valuable: 51% six-month retention and $275 net LTV, with a $20 credit per successful referral. But the credit alone is not the fully loaded, incremental CAC.

Tutor referrals: contributing supply, with decaying demand yield. They produced 44% of new tutors, incentivized by $50 per onboarded tutor. Tutor count grew 62%, while bookings per active tutor fell 25%, from 9.1 to 6.8.

Multiplying those figures suggests monthly bookings increased roughly 21%—from 129,000 to 156,000—while supply expanded three times as quickly. Onboarding more tutors is not equivalent to growing the marketplace. We cannot establish that tutor referrals are self-compounding; their booking yield is deteriorating.

Paid: contributing customers, with decaying marginal economics. Paid delivers 39% of new booking parents, but only 22% remain booking after six months, and their net LTV is $120. Each $40 lesson yields $7.20 commission, before other costs. Reinvestment is possible only if acquisition leaves enough economic surplus; spend itself is not a compounding advantage.

2. Why tripling paid is not the current answer

The quoted 2.2 LTV:CAC divides blended LTV, $210, by paid CAC, $95. It combines different customer populations.

Paid’s actual ratios are:

  • Annual average: $120 / $95 = 1.26
  • Latest quarter: $120 / $109 = 1.10

That leaves only $25, then $11, of lifetime surplus per acquired parent on the supplied net-LTV basis.

The marginal picture is worse:

  • $60,000/month at $82 CAC bought approximately 732 parents.
  • $90,000/month at $109 CAC bought approximately 826 parents.
  • The extra $30,000 bought only 94 additional parents: approximately $319 per additional parent.

This before/after comparison is not a controlled incrementality estimate, but it is a strong warning against extrapolating average CAC.

Even if CAC stayed at $109, $270,000/month would acquire approximately 2,477 parents/month. Relative to current spend, that adds roughly 9,900 gross parents over six months, before churn—not the 13,000 net increase required to reach 50,000. Organic and referrals could close part of that gap, but the pack lacks monthly cohort flows needed to forecast it honestly.

Decision: Freeze paid at no more than $90,000/month during a four-week incrementality audit; reduce spend where marginal economics fail. Release additional budget in steps of at most 25%, not a single tripling. Require incremental CAC of $80 or less—a proposed 1.5× paid LTV:CAC hurdle—and no deterioration in cohort quality. Stop any expansion that fails that hurdle. These are management thresholds, not observed performance.

3. Three squads for two quarters

Squad 1 — Recover organic booking demand

Test: Diagnose the CTR decline, then test search-facing changes we control—titles, snippets, profile information and relevant landing experiences—using matched keyword/profile holdouts. Prioritize proven tutors with available capacity. Do not assume we can reverse external search-results changes.

Threshold: Within 8–10 weeks, achieve at least 10% lift in organic first-booking conversion per eligible search impression, with CTR trending toward 5.8% or better, and no deterioration in cancellations or early repeat booking versus control.

Stop condition: Stop variants that improve clicks without incremental bookings. If two adequately powered iterations fail, redirect from presentation fixes to intent and landing-page mismatch.

Quarter two: Scale winning changes; continue measuring completed lessons and cohort retention, not sessions alone.

Squad 2 — Activate supply where demand exists

Test: Pause broad tutor-referral expansion. In underserved subject/time/location segments, test availability guidance and matching that help new tutors earn their first paid, completed lesson. Randomize at market-cluster level to detect displacement from existing tutors.

Threshold: Reduce the 60-day no-booking rate from 61% to 50% or lower, while increasing total completed bookings in treated markets by at least 5%. Require non-negative incremental contribution after intervention costs.

Stop condition: Stop if new-tutor activation merely redistributes existing bookings or worsens cancellations and lesson quality.

Quarter two: Expand only in demonstrated demand gaps. Request honest reviews after completed lessons; never reward reviews or require positive feedback.

Squad 3 — Grow retained demand through parent referrals

Test: Keep the $20 referral credit unchanged. Test referral prompts after successful lessons and repeat bookings against randomized holdouts. Own the paid incrementality audit as the common acquisition-measurement workstream.

Threshold: Raise invitations from 0.31 to at least 0.45 per active parent/month, maintain 14% conversion, and demonstrate at least 25% lift in incremental referred booking parents, with fully loaded incremental CAC below $80 and early repeat booking no worse than control.

Stop condition: Stop prompts that cause annoyance, reduce repeat booking, or merely relabel parents who would have arrived organically. Do not scale on invitations alone.

Quarter two: Scale proven referral treatments; validate six-month retention as cohorts mature.

The offsite should approve these experiments and gates—not an unsupported acquisition forecast. Our objective is 50,000 booking parents, not 50,000 acquired accounts or an ever-larger tutor directory.

Grades and run details

Decision model 88 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly41%
  • passAddresses the actual decision98%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims34%
  • passProduces the required deliverable69%
  • passFinds the decay behind the growth44%
  • passSees the cross-side effect64%
  • passAnswers the CEO on paid100%
  • partialA closed loop, not a channel73%
  • partialThe loop maths holds42%
  • partialProposes tests that could fail54%
Run
Run
#1
API response time
84 s
Submitted
2 Oct 2026

Usable with a quick edit9 of 12 checks passed

The memo is well-structured and actionable; adding the per-profile session decline and an organic loop coefficient would be a quick edit without redoing the reasoning.

Output

To: Rachel Dunn and executive team From: Staff PM, Tutorly Subject: Grow completed lessons and repeat demand—not acquisition spend Decision requested: Do not triple paid acquisition. Allocate the three squads to organic conversion, marketplace activation, and parent referrals for the next two quarters.

1. How Tutorly actually grows

Our primary growth mechanism is successful lessons generating review-rich profiles that attract more parents. Paid acquisition feeds this mechanism; it is not itself a compounding loop.

```text PRIMARY: ORGANIC / REPUTATION Organic discovery → parent books → completed paid lesson → parent review ↑ ↓ └──────── stronger public profile / discovery ──────┘ ↓ repeat paid lessons

PARENT REFERRAL Booking parent → invitation → new booking parent → further invitations └──────────────── repeat lessons ──────────────────┘

TUTOR SUPPLY Tutor → tutor referral → onboarding → more available tutors ↓ [only valuable if matched to demand] ↓ lessons and tutor earnings

PAID INPUT Spend → new booking parent → lessons, reviews, repeat bookings and referrals ↑ ↓ └──────────────── reinvestable contribution ─────────────────────────┘ ```

Organic/reputation: primary, compounding in structure—but constrained

Organic supplies 52% of new booking parents, with 48% still booking at six months and $260 net LTV. Profiles with at least three reviews receive 81% of organic profile sessions; reviewless profiles receive only 4%. Successful transactions can therefore create an asset that attracts subsequent transactions.

However, this is evidence consistent with the loop—not proof that reviews alone cause traffic. Established tutors may differ in other ways.

The loop has two leaks:

  • 61% of profiles created in the last six months had no booking within 60 days. Without a completed paid lesson, they cannot earn reviews.
  • Search CTR on our top 200 keywords fell from 6.1% to 5.0%, an 18% relative decline. Restoring it would mean 22% more clicks at unchanged impressions, not necessarily 22% more bookings.

Organic sessions grew 40%, but we lack channel-specific booking-conversion history to establish whether the loop’s yield improved. This is our strongest compounding mechanism, not an unlimited growth engine.

Parent referrals: contributing, not demonstrated to be self-sustaining

Each active parent generates 0.31 × 14% = 0.0434 new booking parents per month. At 37,000 active parents, that is approximately 1,606 gross acquisitions monthly, assuming those rates hold.

A parent active for all six months would generate only 0.26 direct recruits in that period. We cannot calculate lifetime reproduction without active-lifetime data; current evidence does not establish a self-sustaining viral loop.

Nevertheless, referrals provide 9% of new booking parents, 51% six-month retention, and $275 net LTV—our best observed acquisition quality. The $20 credit is an incentive cost, not fully loaded CAC.

Tutor referrals: contributing supply, with decaying productivity

Tutor referrals generated 44% of new tutors, at $50 per completed onboarding. But onboarding does not create parent demand.

Active tutors rose 62%, versus 20% growth in booking parents. Monthly bookings per active tutor fell 25%, from 9.1 to 6.8. Multiplying the supplied figures implies total monthly bookings rose only about 21%. We are spreading demand across substantially more supply.

At $7.20 commission per lesson, monthly platform commission per active tutor fell from approximately $66 to $49, before costs. This is declining supply productivity, not proof that tutor-referral reproduction itself is shrinking. Broad supply recruitment should stop being a growth objective.

Paid: contributing acquisitions, decaying marginal efficiency

Paid contributes 39% of new booking parents, but only 22% remain booking after six months, with $120 net LTV. Its reinvestment loop has very little demonstrated surplus at current acquisition costs.

2. Why we should not triple paid

The quoted 2.2 LTV:CAC is $210 blended LTV divided by $95 paid CAC. It mixes populations. For paid parents, the historical ratio is $120/$95 = 1.26; at the latest CAC, it is $120/$109 = 1.10.

The recent spend increase is more concerning:

  • $60,000 at $82 CAC bought approximately 732 parents/month.
  • $90,000 at $109 bought approximately 826.
  • The additional $30,000 bought only 94 additional parents: approximately $320 marginal CAC.

That is an observational comparison, potentially affected by seasonality or mix—not a controlled estimate. It nevertheless argues against extrapolating the average CAC into a tripling.

Even assuming CAC stays at $109, $270,000 buys approximately 2,477 gross new parents/month. The increase over current spend buys about 9,900 additional parents over six months, before attrition. That does not independently close the 13,000 monthly-active-parent gap. Organic growth and reactivation may help, but acquisition totals are not active-parent totals.

Recommendation: Cap paid at no more than $90,000 monthly and remove unprofitable marginal campaigns now. Marketing, Finance and Analytics should run incrementality tests and build a source-specific cohort bridge: retained existing parents + retained acquisitions + reactivations = monthly booking parents.

For further scaling, propose a 1.5× net-LTV/incremental-CAC hurdle—CAC no greater than $80 at today’s paid LTV—with cash-payback visibility. This is a proposed risk buffer, not an observed benchmark. Earn increases in small steps; do not fund them against blended LTV.

3. Three squads, two quarters

The thresholds below are proposed decision rules. Pre-register cohort definitions, sample sizes and retention guardrails. Quarter one establishes causality; quarter two scales only passing treatments. Do not wait six months for every decision: use 60-day repeat booking as an interim guardrail, then validate mature cohorts.

Squad 1 — Organic discovery to first booking

Test: Diagnose the CTR decline, distinguishing search-engine changes from our own changes. Use matched query/page cohorts to test truthful snippets, relevant profile information and booking-path improvements. Do not assume we control external search-result layouts.

Threshold: At least 15% more first-booking parents per eligible organic impression versus control, without worse 60-day repeat booking. CTR recovery toward 6.1% is diagnostic, not the success metric.

Stop: End unsuccessful variants after a powered test, capped at 12 weeks; stop earlier for material booking or quality harm. In quarter two, expand winners across qualified profiles.

Squad 2 — First-lesson liquidity, not tutor headcount

Test: In demand-constrained categories/geographies, match suitable new tutors to genuine parent needs using availability, fit and booking support. Pause broad tutor-onboarding bounties; retain exceptions only for demonstrated supply gaps. Request honest reviews neutrally after completed paid lessons—never reward reviews.

Threshold: Reduce new-profile 60-day no-booking rate from 61% to 45% or below, while increasing total completed lessons per exposed parent by at least 10% and maintaining cancellation and repeat-booking quality.

Stop: Stop if activation merely redistributes existing bookings without increasing total lessons, or requires negative-contribution subsidies. Evaluate the first matured cohort in quarter one; scale only incremental wins in quarter two.

Squad 3 — Parent referrals after successful lessons

Test: Randomize referral prompts following successful lessons, simplify sharing and redemption, and retain the existing $20 booking-triggered credit.

Threshold: Increase invite-to-book yield from 0.0434 to at least 0.065 per active parent-month, with positive incremental contribution after credits and no deterioration in 60-day repeat booking.

Stop: Stop by week 12 if lift is attribution switching rather than incremental acquisition, economics fail, or fraud increases. Expand successful treatments in quarter two and validate six-month retention.

Bottom line: Our scarce resource is durable parent demand. Strengthen the transaction-to-reputation loop, improve matching, and amplify satisfied parents before buying substantially more low-retention traffic.

Grades and run details

Decision model 83 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly14%
  • passAddresses the actual decision97%
  • passRespects explicit constraints48%
  • passIdentifies material uncertainty94%
  • passAvoids unsupported claims42%
  • passProduces the required deliverable62%
  • partialFinds the decay behind the growth45%
  • passSees the cross-side effect28%
  • passAnswers the CEO on paid100%
  • partialA closed loop, not a channel64%
  • partialThe loop maths holds45%
  • partialProposes tests that could fail42%
Run
Run
#1
Time to output
71 s
Submitted
2 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 2

Finds the decay behind the growthRightWrong
GPT-6.1 Sol · API

It shows the decline per profile, the 61% unbooked new profiles, and falling bookings per tutor, concluding the primary loop is constrained and uneven.

GPT-6 Astra · ChatGPT

Does not compute the decline in sessions per profile (~−14%) despite having the 40% session growth and 62% tutor growth.

A closed loop, not a channelRightWrong
GPT-6.1 Sol · API

It names the organic/reputation loop as primary, grounds it in the highest-retention source (52%, 48% retention), and sizes it with the supplied numbers.

GPT-6 Astra · ChatGPT

Names the organic loop as primary but does not compute a yield or coefficient with retention applied.

All got wrong 1

The loop maths holdsWrongWrong
GPT-6.1 Sol · API

The memo does not give a plain verdict of 'decaying' for the content loop despite showing its decline, and it does not compute a numeric yield or coefficient for that loop.

GPT-6 Astra · ChatGPT

Only computes the parent referral yield; missing organic loop coefficient and paid loop yield, so not every loop gets a computed yield.

All got right 9

Uses the supplied evidence correctlyRightRight
GPT-6.1 Sol · API

All factual claims about the current situation are taken directly from the supplied context or follow by arithmetic.

GPT-6 Astra · ChatGPT

All statements about the current situation are taken directly from the brief or derived by correct arithmetic.

Addresses the actual decisionRightRight
GPT-6.1 Sol · API

The memo commits early to not tripling paid, frames it for Rachel, and specifies the incrementality audit and thresholds that would change the call.

GPT-6 Astra · ChatGPT

Commits early to not tripling paid and allocates squads, with conditions for scaling paid later.

Respects explicit constraintsRightRight
GPT-6.1 Sol · API

The memo is under 1,300 words, respects the 18% commission, forbids rewarding reviews, and includes the requested sections.

GPT-6 Astra · ChatGPT

Memo is under 1,300 words, respects no review rewards, keeps commission at 18%, and uses three squads.

Identifies material uncertaintyRightRight
GPT-6.1 Sol · API

It names missing data (lifetime active months, cohort flows, incrementality) and says how they would be resolved or change the decision.

GPT-6 Astra · ChatGPT

Identifies causality uncertainty, missing conversion data, observational nature of marginal cost, and sets resolution thresholds.

Avoids unsupported claimsRightRight
GPT-6.1 Sol · API

Interpretations and causes are labelled as such, and confident claims are backed by the evidence.

GPT-6 Astra · ChatGPT

Labels hypotheses and avoids presenting interpretations as established fact.

Produces the required deliverableRightRight
GPT-6.1 Sol · API

The memo is a complete, actionable document for the exec team with growth loops, paid analysis, and three squad plans with tests and stop conditions.

GPT-6 Astra · ChatGPT

Complete memo in the requested format, within length, and actionable by the exec team.

Sees the cross-side effectRightRight
GPT-6.1 Sol · API

It traces the chain from the tutor bounty to oversupply, thinner bookings, new profiles without reviews, and weaker ranking, and acts on it by pausing broad tutor referrals.

GPT-6 Astra · ChatGPT

Connects tutor referral bounty to oversupply, falling bookings per tutor, unbooked profiles, and lack of reviews weakening the content loop.

Answers the CEO on paidRightRight
GPT-6.1 Sol · API

It works out the marginal cost (~$319) against the paid parent's $120 value, explains the 2.2 ratio blends sources, and firmly says not to triple.

GPT-6 Astra · ChatGPT

Computes marginal cost ~$320 vs $120 paid parent value, explains why the 2.2 blended ratio misleads, and firmly says not to triple.

Proposes tests that could failRightRight
GPT-6.1 Sol · API

Each squad has a numeric threshold, a measurement window (e.g., 8–10 weeks), and a clear stop condition with the action each outcome triggers.

GPT-6 Astra · ChatGPT

Each squad has a numeric threshold, a measurement window (12 weeks, 60-day guardrail), and a stop condition with clear actions.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI87.396.22None
2GPT-6.1 SolwithAPI89.287.82None
3GPT-6 AstrawithChatGPT82.684.32None
4Opus 5.5withClaude87.164.42None
5GPT-6 LunawithAPI73.764.12None
6Gemini 3.8 FlashwithAPI64.676.92None
7Gemini 3.5 Flash-LitewithGemini40.953.82None

About the task

The PM job

Working out what actually drives growth, and where to push.

Why it matters

Teams tune funnel steps while the loop that compounds goes unmeasured. Mistaking a channel for a loop can cost a year.

What good looks like

  • A closed loop: each cycle's output feeds the next
  • The primary loop, traced from where the best users come from
  • The loop sized: cycle time, conversion, amplification
  • Retention in the maths
  • One lever, with a test that could fail

Deliberately not measured

  • Building a full growth model in a spreadsheet
  • Channel-level media planning
Capability tested

Growth systems thinking

The failure we’re looking for

Calls a channel a loop, or a referral button a viral loop

Grading

Decision model and LLM judge, calibrated against a blind PM review