Tasks / Define

Develop product strategy

Can the model make a coherent choice grounded in the evidence, rather than list aspirations?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    It is a complete one-page strategy for the CEO/board, with staged funding, metrics, and risk controls, usable as-is with light edits.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  2. Makes a clear choice100% pass
    It chooses matching quality as the one bet and names what it forgoes: broad paid acquisition, broad tutor recruitment, and unrelated marketplace expansion.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  3. Diagnosis before prescription100% pass
    It diagnoses the mismatch between abundant supply and failed matching, and the proposed actions directly target tutor choice and repeat bookings.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace

Where it slips

  1. Proposes tests that could fail35% pass
    The months 7–12 gate says 'gains persist and contribution economics support expansion' without a numeric threshold or defined contribution target, so not every gate names a threshold.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  2. Uses the supplied evidence correctly40% pass
    The output makes unsupported claims about the current system (open-ended tutor browsing as default journey, existence of data on subject/exam-board segments) that are not in the supplied context.
    GPT-6.1 Sol · API · Supply or demand for a stalled marketplace
  3. Avoids unsupported claims48% pass
    Presents as fact that the current default journey is open-ended tutor browsing and that data on highest-volume subject/exam-board segments exists, neither of which is in the supplied evidence.
    GPT-6.1 Sol · API · Supply or demand for a stalled marketplace

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our tutoring marketplace has stalled at about 4,000 bookings a month. Write a one-page strategy for the next 12 months (about 500 words) for our CEO to take to the board. We can fund one major bet.

ScenarioTutorly connects secondary-school students with independent tutors for online lessons at £30–45 an hour; we take a 20% commission. There are 6,500 active tutors and 11,000 active students. We have 18 months of runway and about £1.2m to put behind the one bet.
MetricsTutor supply up 60% year on year; bookings up 8%. 70% of tutors get zero bookings in a typical month. Repeat booking rate is 64% for students who book twice.
ResearchStudents cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.
Board viewThe board believes paid acquisition of students is the answer.
What a strong answer does

Recognise the contradiction: supply is abundant but matching fails, so buying demand into a broken match pours water into a leaky bucket. Choose matching quality and say what is given up.

Critical failures (cap the score)
  • Recommends paid demand acquisition without addressing matching
Case

v1.9 · synthetic · marketplace, board pressure

Usable with a quick edit8 of 10 checks passed

The strategy is complete and actionable; the unsupported assumptions about the current journey and data can be removed or verified with a quick edit without redoing the reasoning.

Output

Tutorly: 12-Month Strategy for the Board

Decision requested: Back one major bet: a guided matching service that helps students find the right tutor and converts successful introductions into repeat lessons. Commit up to £1.2m, released against evidence. Do not make broad paid acquisition the primary strategy yet.

Why this bet

Tutorly’s constraint appears to be successful matching, not tutor availability. Tutor supply grew 60% year on year, while bookings grew just 8%; 70% of tutors receive no bookings in a typical month. Adding more supply will not solve this imbalance.

Students identify “finding the right tutor” as their main problem. The 64% repeat booking rate among students who book twice suggests value after a relationship forms—but excludes students who never reach a second lesson. It does not establish strong retention across the whole marketplace.

The board’s acquisition thesis may be right: insufficient student demand could explain underutilisation. But buying more traffic before testing the matching bottleneck risks paying to reproduce today’s disappointing conversion. At £6–9 commission per lesson, acquisition economics require sustained repeat bookings.

The major bet: guided matching

Replace open-ended tutor browsing as the default journey with a short intake covering subject, exam board, goals, budget, availability and learning preferences. Present three suitable, genuinely available tutors, explain each recommendation and make the first booking straightforward.

Start with human-assisted matching rather than a large algorithm build. Use this service to learn what predicts a successful relationship, then automate proven steps. After the first lesson, collect fit feedback, offer a quick rematch where necessary and prompt a recurring booking where appropriate.

Concentrate initially on the highest-volume subject and exam-board segments identified in our data. Pause broad tutor recruitment; prioritise responsive existing tutors with relevant availability. Success means productive student–tutor relationships, not equal booking distribution across all 6,500 tutors.

Execution and investment

Months 1–3: Diagnose and test. Establish funnel baselines from student arrival through first, second and subsequent lessons. Run a controlled experiment comparing guided matching with today’s journey, using comparable student cohorts. Measure tutor responsiveness and matching-service cost.

Months 4–6: Prove repeatability. Expand only if the pilot improves completed bookings and repeat behaviour—not merely enquiry volume. Refine recommendations, automate intake and scheduling, and establish contribution economics after matching support costs.

Months 7–12: Scale the winner. Roll out proven workflows across suitable segments. Increase paid student acquisition only where matched cohorts demonstrate viable acquisition payback, turning acquisition into distribution for the same bet.

Allocate £450k to product and engineering, £350k to matching operations, £150k to research and measurement, and £250k as a gated scaling reserve. Review spending quarterly against the remaining runway.

Board scorecard and stop rules

Set initial targets of 25% higher first-booking conversion and 20% higher second-booking conversion versus control, with no deterioration in cancellations or lesson satisfaction. Track 90-day bookings and contribution per student, including acquisition and service costs.

Aim to reach 6,000 monthly bookings by month 12, as a planning target rather than a forecast. If the six-month evidence shows no meaningful improvement or an uneconomic service cost, stop expansion and preserve capital.

Strategic principle: Prove that Tutorly can reliably turn student demand into lasting tutoring relationships before paying substantially more to generate that demand.

Check by check

Got wrong · 2
  • Uses the supplied evidence correctlyThe output makes unsupported claims about the current system (open-ended tutor browsing as default journey, existence of data on subject/exam-board segments) that are not in the supplied context.
  • Avoids unsupported claimsPresents as fact that the current default journey is open-ended tutor browsing and that data on highest-volume subject/exam-board segments exists, neither of which is in the supplied evidence.
Got right · 8
  • Addresses the actual decisionCommits to one bet (guided matching) early, says what is sacrificed (broad paid acquisition as primary strategy), and states the condition (no meaningful improvement or uneconomic cost) that would change the call.
  • Respects explicit constraintsThe output is a one-page strategy of about 500 words for the CEO to take to the board, commits to one major bet within the £1.2m funding, and respects all stated constraints.
  • Identifies material uncertaintyNames the uncertainty (whether matching or demand is the real bottleneck), acknowledges the board's thesis may be right, and specifies that six-month evidence of no improvement or uneconomic cost would stop expansion.
  • Produces the required deliverableDelivers a complete, actionable one-page strategy (~516 words) for the CEO to take to the board, with decision, rationale, execution plan, and scorecard.
  • Surfaces the central contradictionClearly identifies that supply grew 60% while bookings grew only 8% and 70% of tutors get zero bookings, surfacing the mismatch.
  • Makes a clear choiceChooses guided matching as the one bet and explicitly forgoes broad paid acquisition as the primary strategy.
  • Diagnosis before prescriptionStates the crux (matching, not supply, is the constraint) and each major action (guided matching, human-assisted intake, rematch) directly addresses it.
  • Proposes tests that could failProposes numeric thresholds (25% higher first-booking conversion, 20% higher second-booking conversion), a six-month window, and a clear stop rule (no meaningful improvement or uneconomic cost).

Claims the judge couldn’t find in the brief

  • The current default journey is open-ended tutor browsing
  • We have data identifying highest-volume subject and exam-board segments

Grades and run details

Decision model 80 · LLM judge 9 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly19%
  • passAddresses the actual decision100%
  • passRespects explicit constraints46%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims4%
  • passProduces the required deliverable83%
  • passSurfaces the central contradiction100%
  • passMakes a clear choice100%
  • passDiagnosis before prescription100%
  • partialProposes tests that could fail83%
Run
Run
#1
API response time
22 s
Submitted
29 Sept 2026

Usable with a quick edit8 of 10 checks passed

The output is well-structured and evidence-based; a quick edit to add a numeric threshold for the pilot kill criterion would make it fully usable without re-checking evidence or redoing reasoning.

Output

Tutorly: 12-month strategy — make every good match count

Recommendation: Put our £1.2m behind improving the first tutor match, not broad paid acquisition. Tutor supply is growing far faster than bookings, 70% of tutors get no booking in a typical month, and students say finding the right tutor is their main problem. Buying more student traffic before fixing that experience risks paying to send more families into the same frustration. The encouraging signal is that students who book again show a 64% repeat rate: when a match works, there is reason to believe students stay.

The bet: a guided matching and first-booking service

Build a faster, more trusted path from a student’s need to a first lesson. Ask about subject, exam level, goals, budget, availability and preferred teaching style; then show a short, ranked list of tutors who are available and suitable—not a directory students must search alone. Make tutor profiles clearer on relevant experience and approach, and prompt tutors to respond quickly. For students who still struggle, offer human-assisted matching.

Start with the subjects and exam levels that already have the strongest demand and enough suitable tutors. Use actual enquiries, bookings, repeat behaviour and student feedback to improve recommendations. This is one integrated liquidity bet: make existing demand easier to serve with existing supply, rather than adding supply or buying broad traffic.

How we will execute

In the first 90 days, analyse the booking funnel and pilot assisted matching in a few high-demand subject/exam segments. Compare it with the current journey on first booking, time to match, cancellations and student satisfaction. In months 4–6, automate the most successful parts and improve availability and profile information. In months 7–12, expand only where the pilot improves bookings without harming quality. Help tutors without bookings improve their profiles or availability; stop actively promoting tutors who are unavailable or consistently unresponsive.

Allocate the £1.2m across the year: £450k product and data, £300k matching and student support, £250k tutor onboarding and quality, £100k testing and measurement, and £100k contingency. Release funding in stages, with a formal review after the pilot.1

What success looks like

By month 12, target 6,000 monthly bookings (up from 4,000) and at least 45% of active tutors receiving a booking in a typical month (up from 30%). Also improve first-booking conversion and time to first lesson, while maintaining or improving repeat booking, cancellation and satisfaction rates. Report results by subject and exam level so growth cannot hide weak or overloaded segments.

Hold broad paid acquisition for now. Run only small, measurable acquisition tests in segments where matching is working; scale spend only if the resulting students book and repeat at acceptable acquisition cost. If the pilot fails to improve first bookings, stop and reassess before committing the remaining funds. This gives the board a measurable growth bet while protecting our 18-month runway.

What a PM had to fix

  1. 1Invented evidenceVerify or remove the claimTargeted repair

    What we’d changeTreat £1.2m as a funding ceiling, not a spending commitment. Assuming one-hour lessons, reaching 6,000 monthly bookings adds only £12k–18k in monthly commission revenue before costs. That target alone does not justify the investment. Cost the 90-day matching pilot first; release further funding only if incremental completed bookings, repeat behaviour and service costs support a credible path to covering the investment within our financing horizon. Assess that path against actual cash burn and remaining runway before scaling.

Check by check

Got wrong · 1
  • Proposes tests that could failThe pilot kill criterion lacks a numeric threshold; it only says 'fails to improve first bookings' without a specific measurable target.
Mixed · 1
  • Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the supplied context with no invented facts.The two graders disagreed on this one.
Got right · 8
  • Addresses the actual decisionCommits early to one choice (improving matching, not paid acquisition), framed for the CEO, and says the pilot failing to improve first bookings would trigger reassessment.
  • Respects explicit constraintsDelivers a one-page strategy of about 500 words for the CEO to take to the board, respecting the form, reader and length.
  • Identifies material uncertaintyIdentifies that the effect of matching improvement on bookings is uncertain, and specifies that a pilot will resolve it, with a clear stop condition if it fails.
  • Avoids unsupported claimsInterpretations and forecasts are presented as risks or hypotheses, not as established facts; factual claims are all supported.
  • Produces the required deliverableThe output is a complete, actionable strategy memo that the CEO could take to the board with minimal edits.
  • Surfaces the central contradictionClearly surfaces the contradiction: supply grew 60% but bookings only 8%, 70% of tutors idle, and students can't find the right tutor.
  • Makes a clear choiceChooses one bet (matching quality) and explicitly forgoes broad paid acquisition, naming the sacrifice.
  • Diagnosis before prescriptionDiagnoses the crux (broken matching despite abundant supply) and all major actions directly address that crux.

Grades and run details

Decision model 80 · LLM judge 10 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly42%
  • passAddresses the actual decision100%
  • passRespects explicit constraints66%
  • passIdentifies material uncertainty95%
  • partialAvoids unsupported claims11%
  • passProduces the required deliverable96%
  • passSurfaces the central contradiction100%
  • passMakes a clear choice100%
  • passDiagnosis before prescription100%
  • partialProposes tests that could fail80%
Run
Run
#1
API response time
21 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 LunawithAPI87.290.52None
2GPT-6.1 SolwithAPI87.285.92None
3Sonnet 5.5withAPI79.285.92None
4GPT-6 AstrawithChatGPT89.775.52None
5Opus 5.5withClaude78.971.42None
6Gemini 3.5 Flash-LitewithGemini58.357.72None

About the task

The PM job

Writing a strategy memo that leadership can act on.

Why it matters

Strategy that doesn't choose isn't strategy. Models write fluent strategic prose easily; making a choice the evidence supports and naming what it gives up is harder.

What good looks like

  • Diagnoses the one obstacle that matters before prescribing
  • Makes one clear choice and names what is sacrificed
  • Grounds the choice in the supplied evidence
  • Surfaces the tension or contradiction in the data
  • Defines how we would know it is working

Deliberately not measured

  • Market sizing accuracy beyond the supplied data
  • Financial modelling
Capability tested

Choosing where to play and what not to do, from supplied evidence

The failure we’re looking for

Strategic-sounding aspirations without a diagnosis or a choice

Grading

Decision model and LLM judge, calibrated against a blind PM review