Tasks / Define

Develop product strategy

Can the model make a coherent choice grounded in the evidence, rather than list aspirations?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    It is a complete one-page strategy for the CEO/board, with staged funding, metrics, and risk controls, usable as-is with light edits.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  2. Makes a clear choice100% pass
    It chooses matching quality as the one bet and names what it forgoes: broad paid acquisition, broad tutor recruitment, and unrelated marketplace expansion.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  3. Diagnosis before prescription100% pass
    It diagnoses the mismatch between abundant supply and failed matching, and the proposed actions directly target tutor choice and repeat bookings.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace

Where it slips

  1. Proposes tests that could fail35% pass
    The months 7–12 gate says 'gains persist and contribution economics support expansion' without a numeric threshold or defined contribution target, so not every gate names a threshold.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  2. Uses the supplied evidence correctly40% pass
    The output makes unsupported claims about the current system (open-ended tutor browsing as default journey, existence of data on subject/exam-board segments) that are not in the supplied context.
    GPT-6.1 Sol · API · Supply or demand for a stalled marketplace
  3. Avoids unsupported claims48% pass
    Presents as fact that the current default journey is open-ended tutor browsing and that data on highest-volume subject/exam-board segments exists, neither of which is in the supplied evidence.
    GPT-6.1 Sol · API · Supply or demand for a stalled marketplace

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our tutoring marketplace has stalled at about 4,000 bookings a month. Write a one-page strategy for the next 12 months (about 500 words) for our CEO to take to the board. We can fund one major bet.

ScenarioTutorly connects secondary-school students with independent tutors for online lessons at £30–45 an hour; we take a 20% commission. There are 6,500 active tutors and 11,000 active students. We have 18 months of runway and about £1.2m to put behind the one bet.
MetricsTutor supply up 60% year on year; bookings up 8%. 70% of tutors get zero bookings in a typical month. Repeat booking rate is 64% for students who book twice.
ResearchStudents cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.
Board viewThe board believes paid acquisition of students is the answer.
What a strong answer does

Recognise the contradiction: supply is abundant but matching fails, so buying demand into a broken match pours water into a leaky bucket. Choose matching quality and say what is given up.

Critical failures (cap the score)
  • Recommends paid demand acquisition without addressing matching
Case

v1.9 · synthetic · marketplace, board pressure

Usable with a quick edit8 of 10 checks passed

The strategy is clear and actionable; minor unsupported claims (e.g., £50 acquisition cost) can be removed or labelled as assumptions with a quick edit, without redoing the reasoning.

Output

Tutorly 12-Month Strategy: Fix Matching, Then Scale

Recommendation

Put the £1.2m behind one bet: guided matching with a first-lesson guarantee. Paid acquisition comes second, released in stages once the funnel converts.

Why not paid acquisition first

Tutor supply is up 60% and bookings only 8%, so the constraint is not volume. Seventy percent of tutors get no bookings, and students say their biggest problem is finding the right tutor. Tutors say "not enough students," but that is a symptom of poor discovery, not a shortage of demand alone.

The economics are also tight. At an average £37.50 lesson and 20% commission, we earn about £7.50 per booking. A £50 acquisition cost needs roughly seven bookings from a student just to break even. Our best evidence of retention is that 64% of students who book twice book again, but we don't yet know how many get to a second booking. Paying to send more students into a search experience they find frustrating risks buying churn.

The bet

1. Guided matching (about £600k). Replace browse-and-search with a short intake (subject, exam board, grade target, learning needs, availability, budget) that returns three ranked tutors. Add structured tutor profiles and quality signals such as response time, outcomes and reviews. Route new demand to under-booked tutors who fit well, rather than only to the top few.

2. First-lesson guarantee (about £300k). If the first lesson isn't right, the student is rematched and the lesson is free, funded by us. This targets the drop-off before the second booking, which is where the strong repeat behaviour begins.

3. Supply curation (about £100k). Stop passive tutor onboarding. Admit tutors by subject-level demand, and give inactive tutors coaching or profile support, or deprioritise them.

4. Gated acquisition reserve (£200k). Held back and released only when the gates below are met.

Gates for releasing acquisition spend

Two consecutive months of: - first-to-second booking conversion up at least 10 points from baseline; - median time to first booking under 48 hours; - modelled payback on acquisition inside 6 months.

Until then, test small paid channels only to measure cost per matched student.

Targets (month 12)

  • Bookings: 4,000 → 6,000+ per month
  • Tutors with at least one booking a month: 30% → 50%
  • Students reaching a second booking: +10 points
  • Student satisfaction after first lesson: tracked from launch

Timeline

  • Months 1–2: Instrument the funnel (search → enquiry → first lesson → second lesson). Establish the baseline. Interview 30 students and 30 tutors.
  • Months 3–6: Launch matching to one or two high-demand subjects, such as GCSE maths and sciences. Run against a control group.
  • Months 7–9: Roll out across all subjects, and launch the guarantee if the pilot shows it pays back.
  • Months 10–12: Release acquisition spend if the gates are met. Report to the board.

Risks

  • Matching doesn't lift conversion. Pilots with control groups will show this by month 6, leaving 12 months of runway to change course.
  • Guarantee costs exceed benefit. We cap exposure per student and monitor abuse.
  • Tutor churn from curation. We prioritise communication and coaching before any deprioritisation.

Ask of the board

Approve the matching bet and the staged acquisition gate. If the board is not willing to defer scaled paid acquisition, the fallback is a capped £150k test with cost per second booking as the success measure.

Check by check

Got wrong · 2
  • Uses the supplied evidence correctlyOutput states that students find the search experience frustrating, that we don't know how many get to a second booking, and implies a £50 acquisition cost, none of which are in the supplied context.
  • Avoids unsupported claimsPresents 'search experience they find frustrating' and a £50 acquisition cost as facts without support, and states 'we don't yet know how many get to a second booking' as if known.
Got right · 8
  • Addresses the actual decisionCommits to guided matching with first-lesson guarantee as the one bet, defers paid acquisition, and provides a fallback if the board disagrees.
  • Respects explicit constraintsRespects the one-major-bet constraint, delivers a one-page strategy of about 500 words for the CEO to take to the board.
  • Identifies material uncertaintyNames risks (matching not lifting conversion, guarantee costs, tutor churn) and sets gates with numeric thresholds and a fallback plan.
  • Produces the required deliverableDelivers a complete one-page strategy with recommendation, rationale, plan, timeline, risks, and ask, usable by the CEO.
  • Surfaces the central contradictionHighlights that supply is up 60% but bookings only 8% and 70% of tutors get no bookings, surfacing the matching failure.
  • Makes a clear choiceChooses matching as the one bet and explicitly forgoes immediate paid acquisition, naming what is sacrificed.
  • Diagnosis before prescriptionDiagnoses poor discovery as the crux, and each major action (guided matching, guarantee, curation) directly addresses it.
  • Proposes tests that could failGates have numeric thresholds (10-point conversion lift, under 48 hours, 6-month payback), measurement windows (two consecutive months, by month 6), and actions (release spend or change course).

Claims the judge couldn’t find in the brief

  • we don't yet know how many get to a second booking
  • search experience they find frustrating
  • A £50 acquisition cost

Grades and run details

Decision model 75 · LLM judge 9 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly12%
  • passAddresses the actual decision100%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty100%
  • failAvoids unsupported claims14%
  • passProduces the required deliverable78%
  • passSurfaces the central contradiction100%
  • passMakes a clear choice99%
  • passDiagnosis before prescription100%
  • partialProposes tests that could fail62%
Run
Run
#1
API response time
20 s
Submitted
29 Sept 2026

Usable with a quick edit8 of 10 checks passed

The strategy is complete, evidence-based, and actionable; the only small gap is a numeric threshold for the final expansion gate, which is a quick edit.

Output

Tutorly: turn tutor choice into sustained bookings

12-month strategy | Board decision requested

Recommendation

Commit our one major bet to helping students find a suitable tutor and establish a recurring lesson habit. Authorise up to £1.2m in staged funding. Defer broad paid student acquisition until we demonstrate that we can reliably convert demand into repeat bookings at sustainable economics.

Why this bet

Tutor supply has grown 60%, while bookings have grown just 8%. With 70% of tutors receiving no bookings in a typical month, recruiting more supply is unlikely to unlock growth. Students identify finding the right tutor as their biggest problem. Our working hypothesis is that Tutorly offers access without enough confidence to choose.

The board’s acquisition thesis remains plausible: tutors report insufficient students. But that describes their experience, not necessarily the underlying constraint. More traffic could help, or could send more students into an ineffective matching journey.

The reported 64% repeat rate among students who book twice suggests potential value after an initial relationship forms. It does not establish retention across all students or explain losses before the second booking. We must validate its definition and measurement window.

What we will build

Start with the subject and year-group segment showing the strongest combination of existing demand, available tutors and repeat behaviour. Offer students a short intake covering learning goals, syllabus, budget and availability, followed by three suitable, available tutors with clear reasons for each recommendation.

After the first lesson, check fit, offer a straightforward rematch where needed and make recurring booking easy. Begin with manually assisted matching to learn what predicts successful relationships before automating it. Broad tutor recruitment and unrelated marketplace expansion will be deferred.

Execution and investment

In months 1–2, release £150,000 to establish the funnel from qualified enquiry through first, second and subsequent lessons, interview students who abandoned booking and run a controlled matching pilot. Audit “active” user definitions: the headline counts do not establish usable demand or supply.

In months 3–6, release a further £350,000 only if the pilot improves second completed bookings within 30 days by at least 20% relative to the existing journey, without materially increasing refunds or complaints. This is a proposed investment threshold, not a forecast; sample size and measurement rules will be agreed before testing.

In months 7–12, release the remaining £700,000 only if gains persist and contribution economics support expansion. Automate proven matching steps and extend into adjacent segments. Small paid-acquisition tests can then establish whether additional demand scales profitably.

Success and downside protection

Target 8,000 completed monthly bookings by month 121, with term-time and holiday comparisons reported separately. Track contribution after payment costs, support, incentives and matching operations alongside booking growth.

Assuming one-hour bookings, current commission revenue is only £24,000–£36,000 monthly. Doubling bookings would not itself establish sustainability. Acquisition spending must therefore recover its cost from observed contribution within six months, rather than speculative lifetime value.

If matching fails the month-six gate, stop further funding and reassess the acquisition thesis. With 18 months of runway, preserving capital when evidence disappoints is part of the strategy.

What a PM had to fix

  1. 1Test or gate too weakTighten the testQuick edit

    What we’d changeGive the target a basis, a short bridge showing the conversion, booking-frequency or demand change it needs, and say exactly when the month-6 gate is judged.

Check by check

Got wrong · 1
  • Proposes tests that could failThe months 7–12 gate says 'gains persist and contribution economics support expansion' without a numeric threshold or defined contribution target, so not every gate names a threshold.
Mixed · 1
  • Uses the supplied evidence correctlyAll current-situation facts used are from the metrics, research, scenario, or arithmetic from them, with assumptions clearly labelled.The two graders disagreed on this one.
Got right · 8
  • Addresses the actual decisionIt commits early to matching quality as the one bet and says the month-six failure would stop funding and reassess acquisition.
  • Respects explicit constraintsIt respects the one-bet funding constraint, 12-month horizon, CEO/board audience, and approximate 500-word length.
  • Identifies material uncertaintyIt names the unproven matching hypothesis and repeat-rate measurement as uncertainties, with a numeric month-six gate to resolve them.
  • Avoids unsupported claimsCausal and forward-looking statements are framed as hypotheses, possibilities, or assumptions rather than established fact.
  • Produces the required deliverableIt is a complete one-page strategy for the CEO/board, with staged funding, metrics, and risk controls, usable as-is with light edits.
  • Surfaces the central contradictionIt explicitly states supply grew 60% while bookings grew only 8%, and that 70% of tutors get no bookings.
  • Makes a clear choiceIt chooses matching quality as the one bet and names what it forgoes: broad paid acquisition, broad tutor recruitment, and unrelated marketplace expansion.
  • Diagnosis before prescriptionIt diagnoses the mismatch between abundant supply and failed matching, and the proposed actions directly target tutor choice and repeat bookings.

Grades and run details

Decision model 85 · LLM judge 10 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly11%
  • passAddresses the actual decision100%
  • passRespects explicit constraints58%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims76%
  • passProduces the required deliverable71%
  • passSurfaces the central contradiction100%
  • passMakes a clear choice100%
  • passDiagnosis before prescription100%
  • partialProposes tests that could fail50%
Run
Run
#1
Time to output
39 s
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 LunawithAPI87.290.52None
2GPT-6.1 SolwithAPI87.285.92None
3Sonnet 5.5withAPI79.285.92None
4GPT-6 AstrawithChatGPT89.775.52None
5Opus 5.5withClaude78.971.42None
6Gemini 3.5 Flash-LitewithGemini58.357.72None

About the task

The PM job

Writing a strategy memo that leadership can act on.

Why it matters

Strategy that doesn't choose isn't strategy. Models write fluent strategic prose easily; making a choice the evidence supports and naming what it gives up is harder.

What good looks like

  • Diagnoses the one obstacle that matters before prescribing
  • Makes one clear choice and names what is sacrificed
  • Grounds the choice in the supplied evidence
  • Surfaces the tension or contradiction in the data
  • Defines how we would know it is working

Deliberately not measured

  • Market sizing accuracy beyond the supplied data
  • Financial modelling
Capability tested

Choosing where to play and what not to do, from supplied evidence

The failure we’re looking for

Strategic-sounding aspirations without a diagnosis or a choice

Grading

Decision model and LLM judge, calibrated against a blind PM review