Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 6 graded outputs by 3 models. 67% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits to a clear two-quarter roadmap and named deferrals, and states the evidence or conditions that would change later choices.
    GPT-6.1 Sol · API · Two squads, eight asks, one half
  2. Identifies material uncertainty100% pass
    It names material unknowns such as reminder causation, recovered support time, dashboard adoption, timely cancellations, and Physitrack demand, and says how they would be resolved.
    GPT-6.1 Sol · API · Two squads, eight asks, one half
  3. Outcomes, with certainty that falls with distance100% pass
    Items are framed by outcomes, near-term work is specific, and later work is deliberately looser pending evidence.
    GPT-6.1 Sol · API · Two squads, eight asks, one half

Where it slips

  1. Fits the stated capacity67% pass
    Committed Q1/Q2 Approvals & Policy work exceeds 11 squad-weeks per quarter, and the conditional procurement plan displaces multi-entity and spend-controls work.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  2. Respects explicit constraints67% pass
    It exceeds the 1,200-word limit for the case and proposes procurement work that would consume the Approvals & Policy squad's multi-entity and spend-controls capacity.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Every commitment serves the goals75% pass
    Physitrack integration is committed in Q1 but is not tied to either CEO goal (no-shows or multi-site churn) or to freeing capacity, and the criterion requires such items to be deferred with a reason.
    Opus 5.5 · Claude · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Tidewell's booking product. Write the roadmap for the next two quarters (Q4 2026 and Q1 2027) and what comes after, for Dana Okafor, our CPO, and the two squad leads. It will also be shared with Sales and Design, so it needs to say what we're not doing and why. Keep it under 800 words. Everything we know is below.

About TidewellOnline booking and scheduling for independent physiotherapy clinics. 1,450 clinics, $3.9M ARR.
Goals for the next two quarters (set by the CEO)1. Cut no-shows for our clinics: it's the main thing we promise them. 2. Cut churn among multi-site clinics.
DataNo-shows average 11% of appointments across all clinics. The 40 clinics that text their own patients a reminder the day before average 6%. Multi-site clinics are 16% of clinics and 38% of ARR; their monthly logo churn is 2.9%, against 1.3% for single-site clinics. In exit surveys, 23 of the 31 multi-site clinics that left in the last year cited 'can't see all our locations in one place'. Booking page: 64% of visits end in a booking. In 18 customer interviews this quarter, nobody mentioned the booking page.
CapacityTwo squads. Each has 13 weeks a quarter, but about a quarter of each squad's time goes on support and bugs, so we plan on 10 squad-weeks of roadmap work per squad per quarter.
Candidate work (estimates in squad-weeks, from the squad leads)1. Calendar sync fix, 5. Calendar sync failures cause 38% of support tickets; engineering expects fixing them to free about 3 squad-weeks a quarter of support time, starting the quarter after it ships. 2. Messaging service (SMS provider, patient consent records, templates), 8. Nothing patient-facing on its own. 3. SMS reminders with confirm or cancel, 6. Needs the messaging service. 4. Multi-site dashboard (every location's bookings and utilisation in one view), 10. 5. Waitlist auto-fill: texts waitlisted patients when a slot frees up, 7. Needs the messaging service. It only fills slots cancelled with some notice, which is rare today: patients don't cancel, they just don't turn up. 6. Deposits at booking, 14. Needs a payments partner contract; Legal says the earliest it can be signed is mid-February 2027. 7. Booking page redesign, 9. Requested by the Head of Design. 8. Physitrack integration (sends patients' exercise plans to the Physitrack app), 8. See the Sales note.
Sales noteFrom Jordan, Head of Sales: “FitPhysio Group (30 sites, $64,800 ARR) is our biggest deal this year. I told them Physitrack would be live by Christmas. It's what closes the deal.”
FitPhysio's evaluation notes (shared by their operations director)“Must have: see all 30 sites' bookings and utilisation in one place. Nice to have: Physitrack integration.”
What a strong answer does

A two-quarter roadmap that fits about 40 squad-weeks (10 per squad per quarter), with items framed as the outcome they serve. Puts the calendar sync fix first because it frees capacity for Q1, builds the messaging service before reminders and the waitlist, and sequences reminders before the waitlist since slots only free up once patients can cancel. Commits the multi-site dashboard in the half: it serves the churn goal and is FitPhysio's own must-have, so it tells Sales to lead with the dashboard rather than promising Physitrack by Christmas. Leaves deposits for after the contract (it can't be built in the half once signing is mid-February) and doesn't commit the booking page redesign without evidence. A committed plan of about 36 squad-weeks (sync 5, messaging 8, dashboard 10, reminders 6, waitlist 7) leaves a little slack; Physitrack only fits if something else goes. Near-term items are specific, later ones looser, and it names what's not being done and why.

Critical failures (cap the score)
  • Schedules an item before something it depends on: SMS reminders or waitlist auto-fill before the messaging service, or deposits shipping in the half
  • Commits clearly more than about 40 squad-weeks of work to the two quarters
Case

v1.0 · synthetic · B2B SaaS, healthcare, capacity, sales promise

Usable with a quick edit11 of 11 checks passed

A PM could use this roadmap with light edits because the commitments, capacity, dependencies, deferrals, and Sales message are clear. The main gap is that Q1 reserved capacity is not yet sized into specific follow-up work.

Output

Tidewell booking roadmap: Q4 2026–Q1 2027

For Dana Okafor, CPO; squad leads; Sales and Design

Direction

Prioritise fewer no-shows and lower multi-site churn, not a broader feature catalogue.

  • No-shows average 11%. The 40 clinics sending their own reminders average 6%: promising evidence for reminders, but not proof of causation.
  • Multi-site clinics represent 38% of ARR. Their monthly logo churn is 2.9%, versus 1.3% for single-site clinics. Of 31 multi-site exits, 23 cited the lack of an all-location view.

These findings support two primary investments: multi-site visibility and SMS reminders. Fixing calendar sync makes both more sustainable by reducing support demand.

Q4 2026: retain multi-site clinics and build the reminder foundation

SquadRoadmap allocationDeliverable
------:---
ACalendar sync fix — 5 squad-weeksShip first; address failures responsible for 38% of support tickets.
AMessaging service — 5 of 8 squad-weeksBegin provider integration, consent records and templates. Not a customer-facing launch.
BMulti-site dashboard — 10 squad-weeksDeliver bookings and utilisation across locations in one view.

Total: 20 squad-weeks, matching available roadmap capacity. Messaging carries three weeks into Q1.

Pilot the dashboard with multi-site operators, including FitPhysio if they agree, before broad release. Design should prioritise cross-location workflows and dashboard usability rather than a booking-page refresh.

Calendar sync is a capacity investment as well as a reliability fix. Engineering forecasts three additional squad-weeks per quarter from Q1; we will verify that reduction rather than spend it in advance.

Q1 2027: launch reminders and prove outcomes

SquadRoadmap allocationDeliverable
------:---
AFinish messaging — 3 squad-weeksComplete the prerequisite service.
ASMS reminders with confirm/cancel — 6 squad-weeksPilot, then expand based on results and consent readiness.
A1 squad-week reservedLaunch hardening and fixes.
B10 squad-weeks reservedDashboard adoption and retention improvements, reminder rollout support, and evidence-led follow-up work.

Committed feature build: nine squad-weeks; baseline capacity: 20. The remaining capacity is an explicit investment envelope, not an estimate for unspecified features. Squad leads will size follow-up work after dashboard feedback and reminder pilot results. Any recovered support capacity is additional headroom.

Do not fill this envelope with lower-priority commitments merely to reach full utilisation.

Success measures

  • Reminders: compare no-show rates for pilot clinics with their baseline and a comparable holdout where practical. Track delivery, confirmations, cancellations and cancellation notice. Set rollout thresholds before launch; do not promise the observed 6% rate.
  • Dashboard: track active multi-site usage, ability to complete cross-location tasks, and location-visibility complaints.
  • Retention: monitor multi-site monthly logo churn against the 2.9% baseline. Expect adoption signals before a reliable churn verdict.
  • Calendar sync: track sync-related tickets and actual support time recovered.

Not committed—and why

  • Physitrack integration (8 weeks): not on the Christmas roadmap. FitPhysio’s documented must-have is the dashboard; Physitrack is a nice-to-have. Jordan and Product must promptly correct the promise, confirm the buying criteria with FitPhysio, and offer a dashboard-led evaluation. Do not replace an unsupported date with another.
  • Booking page redesign (9 weeks): 64% conversion alone does not establish a problem, and none of 18 interviews raised it. Revisit with funnel evidence or usability findings.
  • Deposits (14 weeks): the partner contract cannot be signed before mid-February 2027. Too large and dependency-constrained for a credible Q1 commitment.
  • Waitlist auto-fill (7 weeks): today, cancellations with notice are rare. First establish whether reminders create enough recoverable slots. Consider Q1 only if evidence and remaining delivery time support it.

After Q1

Prioritise waitlist auto-fill if timely cancellations increase; otherwise improve reminder effectiveness and dashboard adoption. Reassess deposits after the contract and an operational/business case are ready. Reconsider Physitrack against verified demand and opportunity cost. No dates beyond Q1 until outcome evidence and estimates justify them.

Check by check

Got right · 11
  • Uses the supplied evidence correctlyThe output’s current-situation facts, figures, customer statements, estimates, dates, and capacity numbers all come from the supplied context or follow directly from it.
  • Addresses the actual decisionIt commits to a clear two-quarter roadmap and named deferrals, and states the evidence or conditions that would change later choices.
  • Respects explicit constraintsIt is addressed to the named readers, stays under 800 words, covers Q4/Q1 and after, says what is not being done and why, and its allocations fit the stated capacity.
  • Identifies material uncertaintyIt names material unknowns such as reminder causation, recovered support time, dashboard adoption, timely cancellations, and Physitrack demand, and says how they would be resolved.
  • Avoids unsupported claimsInterpretations and forecasts are labelled as evidence, assumptions, or verification items rather than presented as established facts.
  • Produces the required deliverableIt is a usable roadmap for the CPO, squad leads, Sales, and Design, with near-term commitments, success measures, deferrals, and post-Q1 direction.
  • Every commitment serves the goalsEvery committed item serves no-show reduction, multi-site churn reduction, or capacity freed for those goals, and non-serving items are deferred with reasons.
  • Handles the Sales promise with evidenceIt uses FitPhysio’s own must-have/nice-to-have notes to tell Sales to correct the Physitrack promise and lead with the dashboard.
  • Outcomes, with certainty that falls with distanceItems are framed by outcomes, near-term work is specific, and later work is deliberately looser pending evidence.
  • Fits the stated capacityThe committed work totals 29 squad-weeks against 40 available, with checkable quarter sums and explicit slack/reserved capacity.
  • Sequences around dependenciesMessaging precedes reminders, deposits are deferred until after the contract, and the dashboard is committed before Physitrack is reconsidered.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly78%
  • passAddresses the actual decision98%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims55%
  • passProduces the required deliverable68%
  • passEvery commitment serves the goals96%
  • passHandles the Sales promise with evidence100%
  • passOutcomes, with certainty that falls with distance83%
  • passFits the stated capacity70%
  • passSequences around dependencies69%
Run
Run
#1
API response time
40 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI84.7100.02None
2GPT-6.1 SolwithAPI97.984.62None
3Opus 5.5withClaude82.465.12None

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review