Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 6 graded outputs by 3 models. 67% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits to a clear two-quarter roadmap and named deferrals, and states the evidence or conditions that would change later choices.
    GPT-6.1 Sol · API · Two squads, eight asks, one half
  2. Identifies material uncertainty100% pass
    It names material unknowns such as reminder causation, recovered support time, dashboard adoption, timely cancellations, and Physitrack demand, and says how they would be resolved.
    GPT-6.1 Sol · API · Two squads, eight asks, one half
  3. Outcomes, with certainty that falls with distance100% pass
    Items are framed by outcomes, near-term work is specific, and later work is deliberately looser pending evidence.
    GPT-6.1 Sol · API · Two squads, eight asks, one half

Where it slips

  1. Fits the stated capacity67% pass
    Committed Q1/Q2 Approvals & Policy work exceeds 11 squad-weeks per quarter, and the conditional procurement plan displaces multi-entity and spend-controls work.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  2. Respects explicit constraints67% pass
    It exceeds the 1,200-word limit for the case and proposes procurement work that would consume the Approvals & Policy squad's multi-entity and spend-controls capacity.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Every commitment serves the goals75% pass
    Physitrack integration is committed in Q1 but is not tied to either CEO goal (no-shows or multi-site churn) or to freeing capacity, and the criterion requires such items to be deferred with a reason.
    Opus 5.5 · Claude · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Tidewell's booking product. Write the roadmap for the next two quarters (Q4 2026 and Q1 2027) and what comes after, for Dana Okafor, our CPO, and the two squad leads. It will also be shared with Sales and Design, so it needs to say what we're not doing and why. Keep it under 800 words. Everything we know is below.

About TidewellOnline booking and scheduling for independent physiotherapy clinics. 1,450 clinics, $3.9M ARR.
Goals for the next two quarters (set by the CEO)1. Cut no-shows for our clinics: it's the main thing we promise them. 2. Cut churn among multi-site clinics.
DataNo-shows average 11% of appointments across all clinics. The 40 clinics that text their own patients a reminder the day before average 6%. Multi-site clinics are 16% of clinics and 38% of ARR; their monthly logo churn is 2.9%, against 1.3% for single-site clinics. In exit surveys, 23 of the 31 multi-site clinics that left in the last year cited 'can't see all our locations in one place'. Booking page: 64% of visits end in a booking. In 18 customer interviews this quarter, nobody mentioned the booking page.
CapacityTwo squads. Each has 13 weeks a quarter, but about a quarter of each squad's time goes on support and bugs, so we plan on 10 squad-weeks of roadmap work per squad per quarter.
Candidate work (estimates in squad-weeks, from the squad leads)1. Calendar sync fix, 5. Calendar sync failures cause 38% of support tickets; engineering expects fixing them to free about 3 squad-weeks a quarter of support time, starting the quarter after it ships. 2. Messaging service (SMS provider, patient consent records, templates), 8. Nothing patient-facing on its own. 3. SMS reminders with confirm or cancel, 6. Needs the messaging service. 4. Multi-site dashboard (every location's bookings and utilisation in one view), 10. 5. Waitlist auto-fill: texts waitlisted patients when a slot frees up, 7. Needs the messaging service. It only fills slots cancelled with some notice, which is rare today: patients don't cancel, they just don't turn up. 6. Deposits at booking, 14. Needs a payments partner contract; Legal says the earliest it can be signed is mid-February 2027. 7. Booking page redesign, 9. Requested by the Head of Design. 8. Physitrack integration (sends patients' exercise plans to the Physitrack app), 8. See the Sales note.
Sales noteFrom Jordan, Head of Sales: “FitPhysio Group (30 sites, $64,800 ARR) is our biggest deal this year. I told them Physitrack would be live by Christmas. It's what closes the deal.”
FitPhysio's evaluation notes (shared by their operations director)“Must have: see all 30 sites' bookings and utilisation in one place. Nice to have: Physitrack integration.”
What a strong answer does

A two-quarter roadmap that fits about 40 squad-weeks (10 per squad per quarter), with items framed as the outcome they serve. Puts the calendar sync fix first because it frees capacity for Q1, builds the messaging service before reminders and the waitlist, and sequences reminders before the waitlist since slots only free up once patients can cancel. Commits the multi-site dashboard in the half: it serves the churn goal and is FitPhysio's own must-have, so it tells Sales to lead with the dashboard rather than promising Physitrack by Christmas. Leaves deposits for after the contract (it can't be built in the half once signing is mid-February) and doesn't commit the booking page redesign without evidence. A committed plan of about 36 squad-weeks (sync 5, messaging 8, dashboard 10, reminders 6, waitlist 7) leaves a little slack; Physitrack only fits if something else goes. Near-term items are specific, later ones looser, and it names what's not being done and why.

Critical failures (cap the score)
  • Schedules an item before something it depends on: SMS reminders or waitlist auto-fill before the messaging service, or deposits shipping in the half
  • Commits clearly more than about 40 squad-weeks of work to the two quarters
Case

v1.0 · synthetic · B2B SaaS, healthcare, capacity, sales promise

Usable with a quick edit9 of 11 checks passed

The output is well-structured and evidence-based; the only flaw is a missing explicit tie of Physitrack to the churn goal, which a PM could fix with a one-sentence edit without redoing any reasoning.

Output

Tidewell Booking: Roadmap, Q4 2026 to Q1 2027

For: Dana Okafor (CPO), Squad Leads · Shared with: Sales, Design

The short version

We have two goals: cut no-shows, and cut multi-site churn. The evidence points clearly to one fix for each.

  • No-shows: Clinics that text patients the day before see 6% no-shows against 11% overall. SMS reminders are the fix.
  • Multi-site churn: Multi-site clinics churn at 2.9% a month against 1.3% for single-site, and they are 38% of ARR. 23 of 31 who left cited "can't see all our locations in one place." The multi-site dashboard is the fix.

Everything else is either groundwork for these two, or waits.

Capacity

Each squad has 10 squad-weeks of roadmap work a quarter, so 20 in total. If the calendar sync fix ships in Q4, it frees about 3 squad-weeks from Q1 onward, giving roughly 23.

Q4 2026 (20 of 20 squad-weeks)

SquadWorkWeeksWhy
AMulti-site dashboard10Goal 2. Top exit reason. FitPhysio's must-have.
BCalendar sync fix538% of support tickets. Frees about 3 weeks a quarter from Q1.
BMessaging service (part 1)5Foundation for reminders.

Risk: The dashboard takes all of Squad A's quarter, so it lands late December. Any slip pushes it into January. We will cut scope before we miss the date, for example by shipping bookings first and utilisation second.

Q1 2027 (17 of about 23 squad-weeks)

SquadWorkWeeksWhy
BMessaging service (part 2)3Finishes the foundation.
BSMS reminders with confirm/cancel6Goal 1. Target launch mid-March.
APhysitrack integration8Commitment to FitPhysio (see below).
Buffer~6These are first estimates on new work. Covers slippage, then early waitlist work.

What success looks like

  • No-shows: Clinics using reminders move toward the 6% seen by self-texting clinics. We'll measure from launch.
  • Multi-site churn: Monthly churn falls from 2.9% toward the single-site rate. Exit-survey mentions of "one place" drop.
  • Support load: Calendar sync tickets fall by most of their current 38% share.

FitPhysio and Physitrack: a date change is needed

Jordan told FitPhysio that Physitrack would be live by Christmas. We can't deliver that without dropping either the dashboard or reminders, and those are the CEO's two goals.

FitPhysio's own evaluation notes list the dashboard as their must-have and Physitrack as nice-to-have. The plan above gives them the thing they said they need by year-end, with Physitrack following in Q1.

Ask for Jordan: Please reset the Physitrack date with FitPhysio to Q1 2027, leading with the dashboard landing in December. I'm happy to join that call.

Going forward, please check dates with product before they go to customers. That way Sales can commit with confidence.

What we're not doing, and why

Booking page redesign (9 weeks): not planned. - 64% of visits already end in a booking. - Across 18 customer interviews this quarter, nobody raised it. - It doesn't serve either goal.

Design: we'd value your help on the dashboard and on the reminder messages patients receive. Those are where design effort changes outcomes this half. We'll revisit the redesign if conversion drops or customers start asking for it.

Waitlist auto-fill (7 weeks): after Q1. - It only fills slots cancelled with notice, and today patients rarely cancel. They just don't turn up. - Reminders with a cancel button should change that. Once reminders are live, waitlist becomes worth building. It is first in line for the Q1 buffer or for Q2.

Deposits at booking (14 weeks): not before Q2, and not yet committed. - The payments contract can't be signed before mid-February 2027, so building can't start this half. - It also adds friction for patients. We'll decide after we see what reminders do to no-shows. If reminders get us close to 6%, deposits may not be needed.

After Q1 2027 (indicative)

  1. Waitlist auto-fill: Builds on reminders and the messaging service.
  2. Deposits decision: Based on reminder results and the signed contract.
  3. Multi-site follow-ons: Based on dashboard usage and churn data, such as cross-site reporting and permissions.
  4. Booking page: Only if the data changes.

Decisions needed

  • Dana: Approve the plan, and the Physitrack date reset with FitPhysio.
  • Jordan: Own the FitPhysio conversation this week.
  • Squad leads: Confirm estimates, and the dashboard scope-cut plan in case of slippage.

Check by check

Got wrong · 1
  • Every commitment serves the goalsPhysitrack integration is committed in Q1 but is not tied to either CEO goal (no-shows or multi-site churn) or to freeing capacity, and the criterion requires such items to be deferred with a reason.
Mixed · 1
  • Fits the stated capacityCommitted work sums to 37 squad-weeks (20 in Q4, 17 in Q1) against 40-43 available, with a buffer, and it names what was cut or deferred.The two graders disagreed on this one.
Got right · 9
  • Uses the supplied evidence correctlyEvery statement about the current situation is directly from the brief or follows from arithmetic, with no invented facts.
  • Addresses the actual decisionThe output commits to a clear two-quarter plan, addresses the Physitrack promise, and says what is not being done and why, framed for the named readers.
  • Respects explicit constraintsThe output is under 800 words, addresses the required readers, and says what is not being done and why, respecting all stated constraints.
  • Identifies material uncertaintyIt names risks (dashboard slip, estimate uncertainty), uses a buffer, and ties later decisions (deposits, waitlist, booking page) to results from earlier work.
  • Avoids unsupported claimsForecasts and interpretations are clearly labeled as such (e.g., 'should change that'), and no unsupported facts are presented as established.
  • Produces the required deliverableThe roadmap is complete, in the right form, under 800 words, and usable by Dana, squad leads, Sales, and Design with minimal edits.
  • Handles the Sales promise with evidenceIt uses FitPhysio's own notes to show the dashboard is the must-have, gives Sales a clear line to reset the Physitrack date to Q1, and asks Jordan to own the conversation.
  • Outcomes, with certainty that falls with distanceEvery item names its outcome or problem, near-term items are specific, and later items (after Q1) are deliberately looser and indicative.
  • Sequences around dependenciesMessaging service is built before SMS reminders, calendar sync fix precedes the freed capacity, deposits are not in the half, and waitlist is after reminders.

Grades and run details

Decision model 77 · LLM judge 11 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly55%
  • passAddresses the actual decision92%
  • partialRespects explicit constraints13%
  • passIdentifies material uncertainty79%
  • partialAvoids unsupported claims46%
  • passProduces the required deliverable64%
  • partialEvery commitment serves the goals59%
  • passHandles the Sales promise with evidence99%
  • passOutcomes, with certainty that falls with distance85%
  • failFits the stated capacity18%
  • passSequences around dependencies68%
Run
Run
#1
Time to output
47 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI84.7100.02None
2GPT-6.1 SolwithAPI97.984.62None
3Opus 5.5withClaude82.465.12None

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review