Tasks / Define

Develop product strategy

Can the model make a coherent choice grounded in the evidence, rather than list aspirations?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    It is a complete one-page strategy for the CEO/board, with staged funding, metrics, and risk controls, usable as-is with light edits.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  2. Diagnosis before prescription100% pass
    It diagnoses the mismatch between abundant supply and failed matching, and the proposed actions directly target tutor choice and repeat bookings.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  3. Surfaces the central contradiction100% pass
    It explicitly states supply grew 60% while bookings grew only 8%, and that 70% of tutors get no bookings.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace

Where it slips

  1. Proposes tests that could fail32% pass
    The months 7–12 gate says 'gains persist and contribution economics support expansion' without a numeric threshold or defined contribution target, so not every gate names a threshold.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  2. Uses the supplied evidence correctly39% pass
    The output makes unsupported claims about the current system (open-ended tutor browsing as default journey, existence of data on subject/exam-board segments) that are not in the supplied context.
    GPT-6.1 Sol · API · Supply or demand for a stalled marketplace
  3. Avoids unsupported claims41% pass
    Presents as fact that the current default journey is open-ended tutor browsing and that data on highest-volume subject/exam-board segments exists, neither of which is in the supplied evidence.
    GPT-6.1 Sol · API · Supply or demand for a stalled marketplace

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our tutoring marketplace has stalled at about 4,000 bookings a month. Write a one-page strategy for the next 12 months (about 500 words) for our CEO to take to the board. We can fund one major bet.

What the model was given4 items: Scenario, Metrics, Research, Board view
ScenarioTutorly connects secondary-school students with independent tutors for online lessons at £30–45 an hour; we take a 20% commission. There are 6,500 active tutors and 11,000 active students. We have 18 months of runway and about £1.2m to put behind the one bet.
MetricsTutor supply up 60% year on year; bookings up 8%. 70% of tutors get zero bookings in a typical month. Repeat booking rate is 64% for students who book twice.
ResearchStudents cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.
Board viewThe board believes paid acquisition of students is the answer.
What a strong answer doesThe answer key the graders mark against

Recognise the contradiction: supply is abundant but matching fails, so buying demand into a broken match pours water into a leaky bucket. Choose matching quality and say what is given up.

Critical failures (cap the score at 40)
  • Recommends paid demand acquisition without addressing matching
Case

Staff-level · v1.9 · synthetic · marketplace, board pressure

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit3 of 10 checks passed

A PM would need to remove invented current-state claims and add a hard decision gate and explicit sacrifice before using it. The reasoning is directionally strong but not evidence-safe as written.

Output

MEMORANDUM

TO: Board of Directors, Tutorly FROM: Chief Executive Officer DATE: October 2023 SUBJECT: 12-Month Strategic Plan: Unlocking Growth via Managed Matching

---

Executive Summary: Diagnosing the Growth Plateau

Our marketplace has stalled at 4,000 bookings/month (~£30k monthly net revenue) despite healthy user numbers (11k students, 6.5k tutors).

The Board’s instinct to deploy our £1.2m reserve into paid student acquisition is understandable, but premature. The data reveals a liquidity and conversion bottleneck, not an awareness problem: * Supply is bloated and disengaged: Tutor supply grew 60% YoY, yet 70% receive zero bookings monthly. * Friction kills discovery: Students cite “finding the right tutor” as their primary obstacle. Our open-directory model forces parents to sift through 6,500 profiles, causing decision paralysis and drop-off. * Underlying product-market fit is proven: When students reach a second booking, 64% retain.

Pouring paid traffic into an unguided directory will burn our £1.2m on high CAC and low conversion. Our single bet for the next 12 months is transitioning from an open directory to an intelligent, Managed Matching Engine.

---

The Single Bet: "The Perfect Match" Engine

Instead of browsing endless profiles, parents will complete a 60-second diagnostic intake (curriculum tier, personality, learning barriers, availability). Our matching engine pairs them with the single best tutor—backed by a 100% “First Lesson Satisfaction Guarantee” (rebooked free if not a fit).

By actively directing demand to a curated, high-performing top quartile of our supply (~1,500 tutors), we solve both sides of the marketplace: students find immediate trust, and active tutors earn dependable income.

#### Allocation of the £1.2m Capital: 1. Product & Matching Tech (£500k): Algorithmic intake flow, automated calendar sync, and in-app instant booking to eliminate messaging back-and-forth. 2. Onboarding & Operations (£300k): Quality vetting, tiering supply, and a proactive concierge team to manually rescue failed first sessions. 3. Targeted Performance Marketing (£400k): Phased acquisition deployed only once the match-to-book conversion rate improves.

---

12-Month Execution Roadmap

  • Q1–Q2 (Fix the Engine):
  • Replace public browse with the guided diagnostic matching flow.
  • Restrict active routing to the top 25% highest-rated tutors to concentrate liquidity.
  • Launch the "Fit Guarantee" to de-risk the first booking.
  • Q3 (Solve the 2nd Booking):
  • Automate recurring weekly reservation slots right after lesson one (optimizing for the 64% retention threshold).
  • Introduce 4-lesson starter packs to mathematically guarantee lesson two.
  • Q4 (Scale What Works):
  • Deploy the remaining £400k into intent-driven paid channels (Google Search for GCSE/A-Level revision), scaling into a proven, high-converting funnel.

---

Targets & Milestones

MetricCurrent12-Month Target
:---:---:---
Monthly Bookings4,00010,000+
Search-to-First-Booking CVRBaseline+40%
% Tutors with ≥1 Booking/mo30%65% (of active tier)
Annualized Net Revenue (20%)£360k£900k+

Conclusion: Fixing search friction activates the high-retention engine we already possess. By funding matching instead of brute-force acquisition, we reach sustainable unit economics and double our run-rate within our existing 18-month runway.

Claims the judge couldn’t find in the brief

  • Tutorly currently operates an open-directory model in which parents sift through 6,500 profiles.
  • The current model causes decision paralysis and drop-off.
  • The problem is a liquidity and conversion bottleneck, not an awareness problem.
  • Paid traffic into an unguided directory will burn £1.2m on high CAC and low conversion.
  • There is a high-performing top quartile of supply of about 1,500 tutors.

Grades and run details

Decision model 65 · LLM judge 3 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly76%
  • passAddresses the actual decision69%
  • passRespects explicit constraints40%
  • partialIdentifies material uncertainty42%
  • failAvoids unsupported claims85%
  • passProduces the required deliverable67%
  • passSurfaces the central contradiction100%
  • passMakes a clear choice95%
  • passDiagnosis before prescription100%
  • failProposes tests that could fail28%
Run
Run
#1
API response time
16 s
Submitted
1 Oct 2026

Needs more than a quick edit8 of 10 checks passed

The output invents a fact about not tracking first-to-second conversion, requiring re-checking of evidence. A PM would need to verify data systems before using it.

Output

Tutorly: 12-Month Strategy

Sep 24, 2026 · @Martin Slaney

The problem is matching, not demand1

Tutor supply grew 60% last year; bookings grew 8%. 70% of our 6,500 tutors get no bookings in a typical month, and our 11,000 active students generate only \~4,000 bookings a month. Both sides report the same failure from opposite ends: students can't find the right tutor; tutors can't find students. Adding supply has made the catalogue bigger and choosing harder.

The bright spot: students who book twice repeat at 64%. The relationship works once it starts.3 The leak is between arriving and the second lesson.

Why paid acquisition first is the wrong bet

Paid acquisition sends more students into the same funnel that already fails to convert most of the ones we have. We'd be buying students at full cost and losing them at the step before the relationship forms. At a 20% take on £30–45 lessons (\~£7.50 per booking), we only recover acquisition cost from students who stick; today too few get there. Fix conversion first, and every pound of acquisition afterwards is worth more.

The bet: managed matching

Replace browse-and-hope with matching we take responsibility for.

  • Guided intake. Subject, level, exam board, goals, learning style, budget, availability. Two minutes, not twenty profiles.
  • We recommend three tutors, ranked on fit and track record, not on who signed up most recently.
  • Low-risk first lesson. A discounted trial plus a "right tutor or your next lesson free" guarantee.
  • Curate supply. Stop broad tutor recruitment. Recruit only where demand outstrips supply (by subject and exam level); rank and surface tutors on conversion and repeat rate; move inactive tutors out of results.
  • Start with human matchers, then automate what works. The concierge data trains the algorithm.

Budget and milestones

Allocation£
Matching product, data and ranking (team of \~4)500k
Human matching concierge (pilot, then scale down)200k
Trial-lesson subsidy and guarantee150k
Gated acquisition tranche (released at month 6 only if targets hit)350k
Total1.2m

Month 3: matching live for new students in top 3 subjects. Baseline set for first-to-second-lesson conversion (we don't track this cleanly today2; it's the number that matters most).

Month 6 gate: first-to-second conversion up 50% on baseline; share of tutors with zero bookings below 55%. Hit it, and we release £350k into paid acquisition, now into a funnel that converts. Miss it, and we hold the cash as runway.

Month 12: 6,000 bookings a month (+50%), with growth coming from bookings per student, not student count alone.

Ask of the board

Back matching as the one bet, with acquisition as its reward, not its substitute. We're not against growth spend; we're against spending it before the product can hold the students it buys. With 18 months of runway, the month-6 gate gives us a clear read while two-thirds of the budget and a year of runway remain.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisTargeted repair

    Source: ResearchStudents cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.

    What we’d changeFrame matching as the working hypothesis the research points to, and say what the first months will test, rather than stating it as the diagnosis.

  2. 2Invented evidenceVerify or remove the claimQuick edit

    What we’d changeRemove it: the brief doesn't say what's tracked. Say the baseline will be set in month 3.

  3. 3Hypothesis stated as factReframe it as a hypothesisQuick edit

    Source: MetricsRepeat booking rate is 64% for students who book twice.

    What we’d changeThe 64% applies to students who book twice. It doesn't show the relationship works for everyone who starts.

Claims the judge couldn’t find in the brief

  • we don't track this cleanly today (first-to-second-lesson conversion).
  • Adding supply has made the catalogue bigger and choosing harder.
  • The leak is between arriving and the second lesson.

Grades and run details

Decision model 80 · LLM judge 8 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly69%
  • passAddresses the actual decision100%
  • passRespects explicit constraints58%
  • passIdentifies material uncertainty99%
  • failAvoids unsupported claims49%
  • passProduces the required deliverable90%
  • passSurfaces the central contradiction100%
  • passMakes a clear choice99%
  • passDiagnosis before prescription100%
  • passProposes tests that could fail50%
Artefacts
Run
Run
#1
Time to output
62 s
Submitted
24 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Addresses the actual decisionMixedRight
Gemini 3.8 Flash · API

It commits to managed matching, but does not clearly state what result or condition would change that decision.

Opus 5.5 · Claude

Commits to managed matching as the one bet, with acquisition gated on matching improvement, and states the condition that would change the call.

Respects explicit constraintsMixedRight
Gemini 3.8 Flash · API

It allocates £400k to paid acquisition as part of the plan without a concrete gate, so it does not fully enforce the one-major-bet constraint.

Opus 5.5 · Claude

One-page strategy, ~500 words, for CEO to board, one major bet funded within £1.2m.

Identifies material uncertaintyWrongRight
Gemini 3.8 Flash · API

It does not name material unknowns such as whether matching improves conversion, whether top tutors have capacity, or whether the guarantee is affordable.

Opus 5.5 · Claude

Identifies uncertainty about matching effectiveness and sets a month-6 gate with numeric thresholds to resolve it.

Makes a clear choiceMixedRight
Gemini 3.8 Flash · API

It chooses matching, but does not clearly name what is sacrificed, and it still funds paid acquisition without a hard threshold.

Opus 5.5 · Claude

Chooses managed matching as the one bet and explicitly forgoes immediate paid acquisition.

Proposes tests that could failWrongRight
Gemini 3.8 Flash · API

The only gate, deploying paid marketing once match-to-book conversion improves, lacks a numeric threshold, measurement window, and explicit action for failure.

Opus 5.5 · Claude

Month-6 gate has numeric thresholds (conversion up 50%, zero-booking tutors below 55%), a 3-month window, and clear actions for pass/fail.

All got wrong 2

Uses the supplied evidence correctlyWrongWrong
Gemini 3.8 Flash · API

It invents current-state facts such as an open-directory model, parents sifting through profiles, decision paralysis, and a high-performing top quartile of about 1,500 tutors.

Opus 5.5 · Claude

Output claims 'we don't track this cleanly today' about first-to-second conversion, which is not in the supplied context and is an invented fact about current systems.

Avoids unsupported claimsWrongWrong
Gemini 3.8 Flash · API

It presents several interpretations and forecasts as established fact, including high CAC, low conversion, decision paralysis, and guaranteed lesson two.

Opus 5.5 · Claude

Presents interpretations like 'Adding supply has made the catalogue bigger and choosing harder' and 'The leak is between arriving and the second lesson' as established facts without labelling them as hypotheses.

All got right 3

Produces the required deliverableRightRight
Gemini 3.8 Flash · API

It is a board-ready one-page strategy memo for the CEO, roughly within the requested length and usable with light edits.

Opus 5.5 · Claude

Complete strategy memo with diagnosis, bet, budget, milestones, and ask, usable by CEO for board.

Surfaces the central contradictionRightRight
Gemini 3.8 Flash · API

It explicitly identifies that tutor supply grew 60% while 70% of tutors receive zero bookings and bookings grew only 8%.

Opus 5.5 · Claude

Highlights that supply grew 60% but bookings only 8% and 70% of tutors get no bookings, surfacing the matching failure.

Diagnosis before prescriptionRightRight
Gemini 3.8 Flash · API

It diagnoses matching/search friction as the crux and proposes intake, routing, guarantee, and recurring booking actions aimed at that crux.

Opus 5.5 · Claude

Diagnoses matching as the crux, and all actions (guided intake, recommendations, trial, curation) address it.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 LunawithAPI87.290.52None
2GPT-6.1 SolwithAPI87.285.92None
3Sonnet 5.5withAPI79.285.92None
4GPT-6 AstrawithChatGPT89.775.52None
5Opus 5.5withClaude78.971.42None
6Gemini 3.5 Flash-LitewithGemini58.357.72None
7Gemini 3.8 FlashwithAPI71.438.62None

About the task

The PM job

Writing a strategy memo that leadership can act on.

Why it matters

Strategy that doesn't choose isn't strategy. Models write fluent strategic prose easily; making a choice the evidence supports and naming what it gives up is harder.

What good looks like

  • Diagnoses the one obstacle that matters before prescribing
  • Makes one clear choice and names what is sacrificed
  • Grounds the choice in the supplied evidence
  • Surfaces the tension or contradiction in the data
  • Defines how we would know it is working

Deliberately not measured

  • Market sizing accuracy beyond the supplied data
  • Financial modelling
Capability tested

Choosing where to play and what not to do, from supplied evidence

The failure we’re looking for

Strategic-sounding aspirations without a diagnosis or a choice

Grading

Decision model and LLM judge, calibrated against a blind PM review