Tasks / Define

Develop product strategy

Can the model make a coherent choice grounded in the evidence, rather than list aspirations?

Measures the modelTask v1.1 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    It is a complete one-page strategy for the CEO/board, with staged funding, metrics, and risk controls, usable as-is with light edits.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  2. Makes a clear choice100% pass
    It chooses matching quality as the one bet and names what it forgoes: broad paid acquisition, broad tutor recruitment, and unrelated marketplace expansion.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  3. Diagnosis before prescription100% pass
    It diagnoses the mismatch between abundant supply and failed matching, and the proposed actions directly target tutor choice and repeat bookings.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace

Where it slips

  1. Proposes tests that could fail35% pass
    The months 7–12 gate says 'gains persist and contribution economics support expansion' without a numeric threshold or defined contribution target, so not every gate names a threshold.
    GPT-6 Astra · ChatGPT · Supply or demand for a stalled marketplace
  2. Uses the supplied evidence correctly40% pass
    The output makes unsupported claims about the current system (open-ended tutor browsing as default journey, existence of data on subject/exam-board segments) that are not in the supplied context.
    GPT-6.1 Sol · API · Supply or demand for a stalled marketplace
  3. Avoids unsupported claims48% pass
    Presents as fact that the current default journey is open-ended tutor browsing and that data on highest-volume subject/exam-board segments exists, neither of which is in the supplied evidence.
    GPT-6.1 Sol · API · Supply or demand for a stalled marketplace

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

Our tutoring marketplace has stalled at about 4,000 bookings a month. Write a one-page strategy for the next 12 months (about 500 words) for our CEO to take to the board. We can fund one major bet.

ScenarioTutorly connects secondary-school students with independent tutors for online lessons at £30–45 an hour; we take a 20% commission. There are 6,500 active tutors and 11,000 active students. We have 18 months of runway and about £1.2m to put behind the one bet.
MetricsTutor supply up 60% year on year; bookings up 8%. 70% of tutors get zero bookings in a typical month. Repeat booking rate is 64% for students who book twice.
ResearchStudents cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.
Board viewThe board believes paid acquisition of students is the answer.
What a strong answer does

Recognise the contradiction: supply is abundant but matching fails, so buying demand into a broken match pours water into a leaky bucket. Choose matching quality and say what is given up.

Critical failures (cap the score)
  • Recommends paid demand acquisition without addressing matching
Case

v1.9 · synthetic · marketplace, board pressure

Needs more than a quick edit5 of 10 checks passed

The output contains unsupported claims about the current platform and lacks uncertainty handling and test criteria, so a PM would need to re-check evidence and add missing elements before use.

Output

Tutorly: 12-Month Growth Strategy (Board Briefing)

Executive Summary

Tutorly has reached a critical inflection point, plateauing at roughly 4,000 bookings per month. While tutor supply has surged by 60% year-on-year, demand has lagged behind, with bookings up only 8%. Consequently, 70% of our active tutors receive zero bookings in a typical month, threatening supply quality and platform retention.

With 18 months of runway and £1.2m in capital, we face a strategic choice. While the board traditionally advocates for paid student acquisition, qualitative research reveals that our primary bottleneck is not awareness, but friction in matching. Students cite “finding the right tutor” as their top hurdle, while tutors report “not enough students.” Throwing capital at top-of-funnel paid acquisition will merely leak budget into an inefficient discovery funnel.

Instead, our single major bet for the next 12 months is to re-engineer the marketplace discovery and matching engine, transforming Tutorly from a static directory into an intelligent, high-conversion matching platform.

---

The Core Problem & Diagnosis

Our core metrics tell a clear story of imbalance: * The Supply Glut: 6,500 active tutors vs. 11,000 active students, resulting in 70% of tutors experiencing zero-booking months. * The Demand Efficiency: Repeat booking is healthy at 64% for students who book twice, proving that product-market fit exists once a match is made4. The failure is entirely in initial activation and discovery.1

---

The 12-Month Strategic Bet: Intelligent Matchmaking & Activation

We will deploy our £1.2m budget across three synchronized pillars over the next 12 months to drive bookings from 4,000 to 10,000+ monthly transactions.

1. Algorithmic Matching & Guided Discovery (Months 1–4)

  • Action: Shift from manual search filters to an AI-driven, intent-based matching questionnaire for incoming students. Students specify learning style, goals, budget (£30–45/hr), and schedule, and receive top-3 instant tutor recommendations.
  • Investment: £400k in product engineering, UX redesign, and recommendation algorithms.

2. Supply-Side Activation & Tiered Visibility (Months 5–8)

  • Action: Address the 70% zero-booking tutor rate by introducing performance-backed visibility tiers, response-time incentives, and automated profile optimization tools. Tutors who maintain high responsiveness and trial conversion rates are prioritized in the new matching engine.
  • Investment: £300k in tutor success tooling, automated coaching, and CRM workflows.2

3. Targeted Demand Amplification (Months 9–12)

  • Action: With conversion rates optimized, deploy the remaining £500k into targeted performance marketing and SEO. Rather than broad, leaky acquisition campaigns, we will target high-intent secondary-school parents with hyper-specific landing pages tied directly to instant matching funnels.
  • Investment: £500k in performance channels, referral loops, and parent trust initiatives.

---

Financial Impact & Board Recommendation

By fixing the matching bottleneck first, we project a rise in trial-to-booking conversion rates from current lows to over 25%, lifting monthly bookings past 10,0003 and securing a sustainable, high-margin marketplace model well within our 18-month runway.

We request board approval to allocate the £1.2m capital to this intelligent matching strategy, prioritizing structural efficiency over superficial top-of-funnel spend.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisStart again

    Source: ResearchStudents cite 'finding the right tutor' as their top problem. Tutors cite 'not enough students'.

    What we’d changePresent matching friction as the hypothesis the research suggests, not a proven diagnosis, and stage the spending so the board learns before it commits the full £1.2m.

  2. 2Constraint missedRestore the constraintSubstantial rework

    Source: BriefWe can fund one major bet.

    What we’d changeCut back to one bet. Tutor tooling needs its own case when supply already far outstrips bookings.

  3. 3Invented evidenceVerify or remove the claimTargeted repair

    What we’d changeRemove the forecast or show the model: no conversion baseline is given, and nothing connects 25% to 10,000 bookings.

  4. 4Hypothesis stated as factReframe it as a hypothesisQuick edit

    Source: MetricsRepeat booking rate is 64% for students who book twice.

    What we’d changeA 64% repeat rate among two-time bookers suggests value after a match. It doesn't prove product-market fit.

Check by check

Got wrong · 5
  • Uses the supplied evidence correctlyThe output invents facts not in the supplied context, such as the platform being a 'static directory' and current conversion rates being 'low', with no supporting data.
  • Addresses the actual decisionThe output commits to one bet but never states what result or condition would change that decision, as required.
  • Identifies material uncertaintyNo unknowns that could change the decision are named, and no resolution methods or conditions are given.
  • Avoids unsupported claimsInterpretations like 'the failure is entirely in initial activation and discovery' and 'static directory' are presented as established fact without qualification.
  • Proposes tests that could failNo tests, gates, or kill criteria with numeric thresholds, measurement windows, or triggered actions are proposed.
Got right · 5
  • Respects explicit constraintsThe output is a one-page strategy of about 500 words for the CEO to take to the board, and it funds one major bet.
  • Produces the required deliverableThe output is a complete, actionable one-page strategy within the word limit, suitable for the CEO to present to the board.
  • Surfaces the central contradictionIt clearly identifies that tutor supply grew 60% while bookings grew only 8%, surfacing the mismatch.
  • Makes a clear choiceIt chooses re-engineering the matching engine as the single bet and implicitly forgoes paid acquisition as the primary strategy.
  • Diagnosis before prescriptionThe crux is friction in matching, grounded in student and tutor feedback, and all three pillars directly address that crux.

Claims the judge couldn’t find in the brief

  • The failure is entirely in initial activation and discovery.
  • The current platform is a static directory.
  • Trial-to-booking conversion rates are currently low.

Grades and run details

Decision model 50 · LLM judge 5 of 11 checks
Decision model checks
  • failUses the supplied evidence correctly34%
  • partialAddresses the actual decision29%
  • partialRespects explicit constraints30%
  • failIdentifies material uncertainty84%
  • failAvoids unsupported claims92%
  • passProduces the required deliverable46%
  • passSurfaces the central contradiction100%
  • passMakes a clear choice70%
  • passDiagnosis before prescription99%
  • failProposes tests that could fail87%
Artefacts
Run
Run
#1
Time to output
48 s
Submitted
24 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 84% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 LunawithAPI87.290.52None
2GPT-6.1 SolwithAPI87.285.92None
3Sonnet 5.5withAPI79.285.92None
4GPT-6 AstrawithChatGPT89.775.52None
5Opus 5.5withClaude78.971.42None
6Gemini 3.5 Flash-LitewithGemini58.357.72None

About the task

The PM job

Writing a strategy memo that leadership can act on.

Why it matters

Strategy that doesn't choose isn't strategy. Models write fluent strategic prose easily; making a choice the evidence supports and naming what it gives up is harder.

What good looks like

  • Diagnoses the one obstacle that matters before prescribing
  • Makes one clear choice and names what is sacrificed
  • Grounds the choice in the supplied evidence
  • Surfaces the tension or contradiction in the data
  • Defines how we would know it is working

Deliberately not measured

  • Market sizing accuracy beyond the supplied data
  • Financial modelling
Capability tested

Choosing where to play and what not to do, from supplied evidence

The failure we’re looking for

Strategic-sounding aspirations without a diagnosis or a choice

Grading

Decision model and LLM judge, calibrated against a blind PM review