Usable with a quick edit8 of 10 checks passed
The output is well-structured and evidence-based; a quick edit to add a numeric threshold for the pilot kill criterion would make it fully usable without re-checking evidence or redoing reasoning.
Output
Tutorly: 12-month strategy — make every good match count
Recommendation: Put our £1.2m behind improving the first tutor match, not broad paid acquisition. Tutor supply is growing far faster than bookings, 70% of tutors get no booking in a typical month, and students say finding the right tutor is their main problem. Buying more student traffic before fixing that experience risks paying to send more families into the same frustration. The encouraging signal is that students who book again show a 64% repeat rate: when a match works, there is reason to believe students stay.
The bet: a guided matching and first-booking service
Build a faster, more trusted path from a student’s need to a first lesson. Ask about subject, exam level, goals, budget, availability and preferred teaching style; then show a short, ranked list of tutors who are available and suitable—not a directory students must search alone. Make tutor profiles clearer on relevant experience and approach, and prompt tutors to respond quickly. For students who still struggle, offer human-assisted matching.
Start with the subjects and exam levels that already have the strongest demand and enough suitable tutors. Use actual enquiries, bookings, repeat behaviour and student feedback to improve recommendations. This is one integrated liquidity bet: make existing demand easier to serve with existing supply, rather than adding supply or buying broad traffic.
How we will execute
In the first 90 days, analyse the booking funnel and pilot assisted matching in a few high-demand subject/exam segments. Compare it with the current journey on first booking, time to match, cancellations and student satisfaction. In months 4–6, automate the most successful parts and improve availability and profile information. In months 7–12, expand only where the pilot improves bookings without harming quality. Help tutors without bookings improve their profiles or availability; stop actively promoting tutors who are unavailable or consistently unresponsive.
Allocate the £1.2m across the year: £450k product and data, £300k matching and student support, £250k tutor onboarding and quality, £100k testing and measurement, and £100k contingency. Release funding in stages, with a formal review after the pilot.1
What success looks like
By month 12, target 6,000 monthly bookings (up from 4,000) and at least 45% of active tutors receiving a booking in a typical month (up from 30%). Also improve first-booking conversion and time to first lesson, while maintaining or improving repeat booking, cancellation and satisfaction rates. Report results by subject and exam level so growth cannot hide weak or overloaded segments.
Hold broad paid acquisition for now. Run only small, measurable acquisition tests in segments where matching is working; scale spend only if the resulting students book and repeat at acceptable acquisition cost. If the pilot fails to improve first bookings, stop and reassess before committing the remaining funds. This gives the board a measurable growth bet while protecting our 18-month runway.
What a PM had to fix
1Invented evidenceVerify or remove the claimTargeted repair
What we’d changeTreat £1.2m as a funding ceiling, not a spending commitment. Assuming one-hour lessons, reaching 6,000 monthly bookings adds only £12k–18k in monthly commission revenue before costs. That target alone does not justify the investment. Cost the 90-day matching pilot first; release further funding only if incremental completed bookings, repeat behaviour and service costs support a credible path to covering the investment within our financing horizon. Assess that path against actual cash burn and remaining runway before scaling.
Check by check
Got wrong · 1
- Proposes tests that could failThe pilot kill criterion lacks a numeric threshold; it only says 'fails to improve first bookings' without a specific measurable target.
Mixed · 1
- Uses the supplied evidence correctlyAll statements about the current situation are taken directly from the supplied context with no invented facts.The two graders disagreed on this one.
Got right · 8
- Addresses the actual decisionCommits early to one choice (improving matching, not paid acquisition), framed for the CEO, and says the pilot failing to improve first bookings would trigger reassessment.
- Respects explicit constraintsDelivers a one-page strategy of about 500 words for the CEO to take to the board, respecting the form, reader and length.
- Identifies material uncertaintyIdentifies that the effect of matching improvement on bookings is uncertain, and specifies that a pilot will resolve it, with a clear stop condition if it fails.
- Avoids unsupported claimsInterpretations and forecasts are presented as risks or hypotheses, not as established facts; factual claims are all supported.
- Produces the required deliverableThe output is a complete, actionable strategy memo that the CEO could take to the board with minimal edits.
- Surfaces the central contradictionClearly surfaces the contradiction: supply grew 60% but bookings only 8%, 70% of tutors idle, and students can't find the right tutor.
- Makes a clear choiceChooses one bet (matching quality) and explicitly forgoes broad paid acquisition, naming the sacrifice.
- Diagnosis before prescriptionDiagnoses the crux (broken matching despite abundant supply) and all major actions directly address that crux.
Grades and run details
Decision model 80 · LLM judge 10 of 11 checks
Decision model checks
- failUses the supplied evidence correctly42%
- passAddresses the actual decision100%
- passRespects explicit constraints66%
- passIdentifies material uncertainty95%
- partialAvoids unsupported claims11%
- passProduces the required deliverable96%
- passSurfaces the central contradiction100%
- passMakes a clear choice100%
- passDiagnosis before prescription100%
- partialProposes tests that could fail80%
Run
- Run
- #1
- API response time
- 21 s
- Submitted
- 29 Sept 2026