Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Outcomes, with certainty that falls with distance95% pass
    Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).
    Sonnet 5.5 · API · Two squads, eight asks, one half
  2. Sequences around dependencies89% pass
    Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.
    Sonnet 5.5 · API · Two squads, eight asks, one half
  3. Plans on the squads we actually have89% pass
    It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.
    GPT-6 Astra · ChatGPT · A year of spend management, with a hard deadline

Where it slips

  1. Makes the call on procurement50% pass
    It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.
    GPT-6 Luna · API · A year of spend management, with a hard deadline
  2. Produces the required deliverable50% pass
    The roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Uses the supplied evidence correctly61% pass
    Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.
    GPT-6 Astra · ChatGPT · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Tidewell's booking product. Write the roadmap for the next two quarters (Q4 2026 and Q1 2027) and what comes after, for Dana Okafor, our CPO, and the two squad leads. It will also be shared with Sales and Design, so it needs to say what we're not doing and why. Keep it under 800 words. Everything we know is below.

What the model was given7 items: About Tidewell, Goals for the next two quarters (set by the CEO), Data, Capacity, Candidate work (estimates in squad-weeks, from the squad leads), Sales note, FitPhysio's evaluation notes (shared by their operations director)
About TidewellOnline booking and scheduling for independent physiotherapy clinics. 1,450 clinics, $3.9M ARR.
Goals for the next two quarters (set by the CEO)1. Cut no-shows for our clinics: it's the main thing we promise them. 2. Cut churn among multi-site clinics.
DataNo-shows average 11% of appointments across all clinics. The 40 clinics that text their own patients a reminder the day before average 6%. Multi-site clinics are 16% of clinics and 38% of ARR; their monthly logo churn is 2.9%, against 1.3% for single-site clinics. In exit surveys, 23 of the 31 multi-site clinics that left in the last year cited 'can't see all our locations in one place'. Booking page: 64% of visits end in a booking. In 18 customer interviews this quarter, nobody mentioned the booking page.
CapacityTwo squads. Each has 13 weeks a quarter, but about a quarter of each squad's time goes on support and bugs, so we plan on 10 squad-weeks of roadmap work per squad per quarter.
Candidate work (estimates in squad-weeks, from the squad leads)1. Calendar sync fix, 5. Calendar sync failures cause 38% of support tickets; engineering expects fixing them to free about 3 squad-weeks a quarter of support time, starting the quarter after it ships. 2. Messaging service (SMS provider, patient consent records, templates), 8. Nothing patient-facing on its own. 3. SMS reminders with confirm or cancel, 6. Needs the messaging service. 4. Multi-site dashboard (every location's bookings and utilisation in one view), 10. 5. Waitlist auto-fill: texts waitlisted patients when a slot frees up, 7. Needs the messaging service. It only fills slots cancelled with some notice, which is rare today: patients don't cancel, they just don't turn up. 6. Deposits at booking, 14. Needs a payments partner contract; Legal says the earliest it can be signed is mid-February 2027. 7. Booking page redesign, 9. Requested by the Head of Design. 8. Physitrack integration (sends patients' exercise plans to the Physitrack app), 8. See the Sales note.
Sales noteFrom Jordan, Head of Sales: “FitPhysio Group (30 sites, $64,800 ARR) is our biggest deal this year. I told them Physitrack would be live by Christmas. It's what closes the deal.”
FitPhysio's evaluation notes (shared by their operations director)“Must have: see all 30 sites' bookings and utilisation in one place. Nice to have: Physitrack integration.”
What a strong answer doesThe answer key the graders mark against

A two-quarter roadmap that fits about 40 squad-weeks (10 per squad per quarter), with items framed as the outcome they serve. Puts the calendar sync fix first because it frees capacity for Q1, builds the messaging service before reminders and the waitlist, and sequences reminders before the waitlist since slots only free up once patients can cancel. Commits the multi-site dashboard in the half: it serves the churn goal and is FitPhysio's own must-have, so it tells Sales to lead with the dashboard rather than promising Physitrack by Christmas. Leaves deposits for after the contract (it can't be built in the half once signing is mid-February) and doesn't commit the booking page redesign without evidence. A committed plan of about 36 squad-weeks (sync 5, messaging 8, dashboard 10, reminders 6, waitlist 7) leaves a little slack; Physitrack only fits if something else goes. Near-term items are specific, later ones looser, and it names what's not being done and why.

Critical failures (cap the score at 40)
  • Schedules an item before something it depends on: SMS reminders or waitlist auto-fill before the messaging service, or deposits shipping in the half
  • Commits clearly more than about 40 squad-weeks of work to the two quarters
Case

Mid-level · v1.1 · synthetic · B2B SaaS, healthcare, capacity, sales promise

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit5 of 11 checks passed

The output contains an unsupported benchmark claim, lacks slack in the capacity plan, and commits deposits without tying it to the goals, requiring rework of evidence and reasoning before use.

Output

To: Dana Okafor (CPO), Squad Leads CC: Sales, Design From: Product Manager, Booking & Scheduling Subject: Product Roadmap: Q4 2026 – Q1 2027 (and Beyond)

---

Executive Summary

Our mandate is clear: cut no-shows and halt multi-site clinic churn.

Multi-site clinics represent 38% of our ARR ($1.48M) but churn at more than double the rate of single-site clinics (2.9% vs. 1.3%/month). Concurrently, no-shows cost our clinics 11% of their billable volume, despite evidence that pre-visit SMS reminders cut this to 6%.

This roadmap directly tackles these two threats while paying down technical debt to unlock 15% more product development capacity starting in Q1 2027.

---

Capacity Model

  • Standard Capacity: 2 squads × 10 roadmap weeks/quarter = 20 squad-weeks/quarter (3 weeks/squad reserved for support/maintenance).
  • Q4 2026: 20 squad-weeks available.
  • Q1 2027: 23 squad-weeks available (the Q4 calendar sync fix recovers 3 squad-weeks of support overhead per quarter).

---

Q4 2026: Multi-Site Retention & Core Platform

Squad 1: Clinic Operations (10 squad-weeks) * Multi-site Dashboard (10 weeks): Builds unified multi-location visibility into bookings and practitioner utilisation. * Rationale: 74% of churned multi-site clinics (23 of 31) cited the lack of a cross-location view as their reason for leaving. It is also the sole “Must Have” listed by FitPhysio Group ($64.8k ARR).

Squad 2: Platform & Messaging (10 squad-weeks) * Calendar Sync Fix (5 weeks): Resolves root causes of sync failures. * Rationale: Drives 38% of all support tickets. Resolving this unlocks 3 squad-weeks of engineering capacity every quarter starting Q1 2027. * Messaging Service Foundation (5 of 8 weeks): SMS gateway provider integration, consent management, and compliance architecture. * Rationale: Mandatory architectural foundation for automated patient communications in Q1.

---

Q1 2027: Slashing No-Shows & Revenue Recovery

Total Capacity: 23 squad-weeks (includes +3 weeks unlocked from Calendar Sync).

Squad 1: Patient Engagement (11 squad-weeks) * Complete Messaging Service (3 weeks): Finalise templates and delivery logging. * SMS Reminders with Confirm/Cancel (6 weeks): Automated 24-hour reminders with bidirectional text response. * Rationale: Targets our #1 clinic promise. The 40 clinics currently texting reminders manually see no-shows drop from 11% to 6%. * Deposit Discovery & Technical Spec (2 weeks): Lay groundwork for payment gateway integration once legal clears contracts in mid-February.

Squad 2: Clinic Automation (12 squad-weeks) * Waitlist Auto-Fill (7 weeks): Automatically texts waitlisted patients when slots open. * Rationale: SMS confirm/cancel in Squad 1 will generate advance cancellations. Auto-fill instantly monetises those open slots. * Deposits at Booking – Phase 1 (5 of 14 weeks): Begin implementation mid-February once payment contracts execute.

---

What Comes After (Q2 2027+)

  • Deposits at Booking – Phase 2 (9 weeks remaining): Complete deposit workflows to reduce late cancellations and guarantee revenue for booked slots.
  • Integration Ecosystem (Physitrack / EHRs): Standardised webhooks and partner app marketplace based on clinic demand.

---

What We Are Not Doing (And Why)

#### 1. Physitrack Integration (8 weeks) * The Context: Sales requested this to close FitPhysio Group ($64.8k ARR) before Christmas. * Why We Are Deferring: FitPhysio’s operations team explicitly evaluated Physitrack as a “Nice to have,” whereas the Multi-site Dashboard is their sole “Must have.” Building Physitrack in Q4 would force us to drop the Multi-site Dashboard, risking both the FitPhysio close and the broader 38% of our ARR currently vulnerable to churn. We will re-evaluate partner integrations in Q2 2027.

#### 2. Booking Page Redesign (9 weeks) * The Context: Proposed by Design to refresh the booking interface. * Why We Are Deferring: Our current page converts at 64%—well above healthcare industry benchmarks. In 18 clinic interviews this past quarter, not one customer raised booking page friction. Spending 9 weeks here does not move our core metrics of reducing no-shows or stopping churn.

#### 3. Deposits at Booking in Q4 2026 * Why We Are Waiting: Legal confirmed the payments partner contract cannot be signed until mid-February 2027. We will not commit engineering resources ahead of executed legal agreements.

#### 4. Waitlist Auto-Fill in Q4 2026 * Why We Are Sequencing in Q1: Waitlist auto-fill requires cancellations with advance notice to work. Today, patients simply no-show. We must launch SMS confirm/cancel first to create the advance openings that waitlist auto-fill recovers.

Claims the judge couldn’t find in the brief

  • Booking page conversion is well above healthcare industry benchmarks

Grades and run details

Decision model 82 · LLM judge 6 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly17%
  • passAddresses the actual decision70%
  • passRespects explicit constraints11%
  • partialIdentifies material uncertainty32%
  • partialAvoids unsupported claims52%
  • failProduces the required deliverable14%
  • passEvery commitment serves the goals56%
  • passHandles the Sales promise with evidence82%
  • passOutcomes, with certainty that falls with distance36%
  • passFits the stated capacity23%
  • passSequences around dependencies70%
Run
Run
#1
API response time
38 s
Submitted
1 Oct 2026

Usable with a quick edit11 of 11 checks passed

The roadmap is clear, actionable, and addresses all required elements; a PM could use it with only minor wording edits.

Output

Tidewell Booking Roadmap: Q4 2026, Q1 2027 and beyond

For: Dana Okafor, squad leads. Shared with Sales and Design.

Goals and what the data says

  1. Cut no-shows. The average is 11%. The 40 clinics that text their own patients a reminder average 6%. That group is small and self-selected, so this suggests reminders help but doesn't prove it. We'll measure with a holdout group.
  2. Cut multi-site churn. Multi-site clinics are 16% of clinics but 38% of ARR (about $1.5M). Their monthly churn is 2.9%, against 1.3% for single-site. Of the 31 that left last year, 23 cited "can't see all our locations in one place."

Capacity

There are 10 squad-weeks per squad per quarter, so 20 a quarter. Calendar sync ships in Q1, so its freed support time (about 3 squad-weeks) only adds capacity from Q2, bringing Q2 to about 23.

The plan

Squad ASquad B
Q4 2026 (20)Messaging service (8), then start SMS reminders (2)Multi-site dashboard (10)
Q1 2027 (20)Finish SMS reminders (4), calendar sync fix (5), 1 sparePhysitrack integration (8, conditional, see below), 2 for dashboard follow-up
Q2 2027 (~23)Deposits (14), start once the contract is signedWaitlist auto-fill (7), if the trigger below is met

Q4: build both foundations. - Multi-site dashboard (10). This directly targets the top churn reason, and it is FitPhysio's stated must-have. Target: live mid-December. It fills Squad B's quarter with no buffer, so if it slips, we roll out to FitPhysio's sites in phases rather than wait. - Messaging service (8). SMS provider, consent records and templates. Nothing patient-facing ships from it alone, but every no-show feature depends on it.

Q1: ship reminders and pay down support load. - SMS reminders with confirm or cancel (6). This is our main no-show feature. It should ship in January. Success is measured against a holdout group, with the target being movement toward the 6% the self-reminding clinics achieve. - Calendar sync fix (5). Sync failures cause 38% of support tickets. The fix pays for itself in under two quarters. It goes after reminders because reminders serve the CEO's goals directly. - Physitrack integration (8). See the FitPhysio section below.

Q2 and after. - Deposits (14). Deposits should cut no-shows, but Legal can't sign the payments partner contract before mid-February, and 14 weeks of work doesn't fit in what remains of Q1. Action: start the contract process now so the build begins the day it's signed. - Waitlist auto-fill (7). Only build it if reminders prove the case. It fills slots cancelled with notice, and today patients mostly just don't show up. Confirm or cancel reminders should turn some no-shows into early cancellations. Trigger: the share of appointments cancelled with notice rises meaningfully in the first 6–8 weeks after reminders launch.

FitPhysio and Physitrack

FitPhysio (30 sites, $64,800 ARR) is our biggest deal this year. Sales promised Physitrack by Christmas. FitPhysio's own evaluation notes say the dashboard is a must-have and Physitrack is nice-to-have.

  • The dashboard is on track for Christmas, and it is the thing that closes the deal.
  • Physitrack by Christmas would cost us. Squad B is full with the dashboard, so the 8 weeks would have to come from messaging. That delays reminders, our main no-show feature, by a full quarter, for a nice-to-have for one account.
  • Recommendation: Jordan goes back to FitPhysio now. The message is that the dashboard, their must-have, is live before Christmas, and Physitrack follows in Q1. I'd like Jordan and me to have that conversation this week, before the expectation hardens.
  • Physitrack is conditional. We build it in Q1 if FitPhysio signs, or if a second account asks for it. Otherwise that capacity goes to the dashboard follow-up and to buffer.

What we're not doing, and why

  • Booking page redesign (9 weeks, not scheduled). 64% of booking page visits already end in a booking, and none of our 18 customer interviews mentioned the page. It serves neither goal. I'd revisit it if we see evidence of drop-off or patient complaints. Design isn't sidelined, though: the dashboard and the reminder and confirm flows need heavy design input in Q4 and Q1, and I'd like Design embedded with both squads.
  • Waitlist auto-fill now. It would fill few slots today (see the trigger above).
  • Deposits before Q2. Blocked by the contract and too large to fit anyway.
  • Physitrack before Christmas. It would push back our main no-show feature for a nice-to-have.

Risks

  • Dashboard timing. Squad B has no slack in Q4. The mitigation is a phased FitPhysio rollout.
  • Reminder effect. The 11% vs 6% gap may partly reflect which clinics choose to send reminders. The holdout group will give us the real number.
  • Support load. Support and bugs are planned at 25% of time. If they run higher, Q1 calendar sync and Physitrack slip first, before reminders.
  • Legal. If the payments contract slips past mid-February, deposits slip with it one-for-one.

Decisions needed

  1. Dana: approve the Q4 plan and the no-go on the booking page redesign.
  2. Jordan: take the revised Physitrack timing to FitPhysio.
  3. Legal: start the payments partner contract now.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly47%
  • passAddresses the actual decision96%
  • passRespects explicit constraints25%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims29%
  • passProduces the required deliverable17%
  • passEvery commitment serves the goals28%
  • passHandles the Sales promise with evidence100%
  • passOutcomes, with certainty that falls with distance88%
  • passFits the stated capacity45%
  • passSequences around dependencies70%
Run
Run
#1
API response time
39 s
Submitted
30 Sept 2026
Critical failure: Commits clearly more than about 40 squad-weeks of work to the two quarters. Score capped at 40.

Needs more than a quick edit2 of 11 checks passed

A PM would need to rework the capacity math and decide whether Physitrack is committed or deferred. The evidence and sequencing reasoning also need correction before it can be used.

Output

To: Dana Okafor (CPO), Squad Leads, Sales, Design From: PM, Tidewell Booking Subject: Booking Product Roadmap: Q4 2026 – Q1 2027 & Beyond

At $3.9M ARR across 1,450 independent physio clinics, our immediate focus is clear: ruthlessly align our capacity with the CEO's goals of cutting no-shows and reducing multi-site churn.

With two squads and 10 productive weeks per squad-quarter (20 total squad-weeks per quarter), here is our execution plan.

---

Q4 2026: Foundation & The Multi-Site Anchor

Goal Alignment: Multi-site churn, Support capacity.

  • Multi-Site Dashboard (10 squad-weeks – Squad A): A single view of every location’s bookings and utilization.
  • Why: Multi-site clinics drive 38% of our ARR ($1.48M) but churn at 2.9% (vs. 1.3% single-site). "Can't see all locations in one place" was cited in 74% of multi-site exit surveys. This is also the explicit "Must Have" for FitPhysio Group (30 sites, $64.8K ARR).
  • Messaging Service (8 squad-weeks – Squad B): Infrastructure for SMS provider, consent, and templates.
  • Why: Prerequisite for our Q1 no-show reduction engine.
  • Calendar Sync Fix (5 squad-weeks – Squad B, spilling into Q1): Resolving sync failures that drive 38% of support tickets.
  • Why: Engineering estimates this will permanently reclaim ~3 squad-weeks/quarter of engineering time starting Q1.

---

Q1 2027: Solving No-Shows

Goal Alignment: No-show reduction.

  • Calendar Sync Fix completion (0 extra weeks, Squad B finishing up).
  • SMS Reminders with Confirm/Cancel (6 squad-weeks – Squad B): Leveraging the Q4 Messaging Service.
  • Why: Our data proves this works: clinics manually texting reminders drop no-shows from 11% to 6%. Automated confirm/cancel turns passive ghosting into active rescheduling, opening slots early.
  • Physitrack Integration (8 squad-weeks – Squad A): Sends exercise plans directly to the Physitrack app.
  • Why: Sales commitment to close FitPhysio Group by Christmas. (Note: While FitPhysio listed this as "nice-to-have" vs. the dashboard's "must-have," it is a commercial blocker per Sales).
  • Waitlist Auto-Fill (7 squad-weeks – Squad A, utilizing reclaimed support capacity): Texts waitlisted patients when slots free up.
  • Why: Directly targets no-shows, but depends on SMS reminders and cancellations working first.

---

What Comes After (Q2 2027 & Beyond)

  • Deposits at Booking (14 squad-weeks): Paused until Legal secures the payments partner contract (earliest mid-February 2027). This will be our primary Q2/Q3 bet to further crush no-shows once infrastructure is live.

---

What We Are NOT Doing & Why

1. Booking Page Redesign (9 squad-weeks): * Why Not: 64% of visits already convert to bookings, and 18 customer interviews this quarter yielded zero mentions of the booking page. Spending design capacity here is a distraction from our retention and no-show crises. 2. Deposits in Q4/Q1: * Why Not: Blocked by external legal/contract timelines until mid-Q1 at best. 3. Prioritizing Physitrack over Multi-Site in Q4: * Why Not: While Sales needs Physitrack for FitPhysio, FitPhysio’s actual operational blocker is the Multi-Site Dashboard. Building the dashboard first secures our largest pipeline deal and addresses the core churn driver for our highest-value segment.

Claims the judge couldn’t find in the brief

  • Fixing calendar sync will permanently reclaim about 3 squad-weeks per quarter of engineering time starting Q1.
  • Clinics manually texting reminders drop no-shows from 11% to 6%.

Grades and run details

Decision model 32 · LLM judge 2 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly13%
  • partialAddresses the actual decision30%
  • failRespects explicit constraints44%
  • failIdentifies material uncertainty25%
  • partialAvoids unsupported claims43%
  • failProduces the required deliverable27%
  • partialEvery commitment serves the goals44%
  • failHandles the Sales promise with evidence23%
  • passOutcomes, with certainty that falls with distance34%
  • failFits the stated capacity57%
  • passSequences around dependencies38%
Run
Run
#1
Time to output
4 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 9

Uses the supplied evidence correctlyMixedRightWrong
Gemini 3.8 Flash · API

The output claims the booking page conversion is 'well above healthcare industry benchmarks,' which is not in the supplied context and is an unsupported statement about the current situation.

Sonnet 5.5 · API

Every statement about the current situation is taken directly from the brief or derived by arithmetic, with no invented facts.

Gemini 3.5 Flash-Lite · Gemini

It treats correlation as causation for SMS reminders and adds unsupported 'permanently reclaim' language about calendar sync capacity.

Addresses the actual decisionRightRightWrong
Gemini 3.8 Flash · API

The output commits to a clear two-quarter roadmap with specific items, says what is not being done and why, and is framed for the CPO and squad leads.

Sonnet 5.5 · API

The roadmap commits to a clear two-quarter plan with conditional items, names what is not being done and why, and states what would change the Physitrack and waitlist decisions.

Gemini 3.5 Flash-Lite · Gemini

It commits to a roadmap but does not clearly state what condition would change the answer, especially around the Physitrack commitment and capacity trade-off.

Respects explicit constraintsRightRightWrong
Gemini 3.8 Flash · API

The output is under 800 words, in memo form, addressed to the named readers, and includes what is not being done and why.

Sonnet 5.5 · API

The output is under 800 words, addresses the named readers, includes what is not being done and why, and is a roadmap for the requested period.

Gemini 3.5 Flash-Lite · Gemini

It violates the stated capacity constraint by committing 44 squad-weeks to two quarters despite saying capacity is 40.

Identifies material uncertaintyWrongRightWrong
Gemini 3.8 Flash · API

The output does not name specific unknowns that could change the decision, such as delays in the messaging service or the calendar sync fix not freeing the expected capacity, nor does it say how they would be resolved.

Sonnet 5.5 · API

It identifies the self-selection bias in the 6% no-show figure, the need for a holdout group, the trigger for waitlist auto-fill, and the conditional nature of Physitrack, along with risks like dashboard timing and legal delays.

Gemini 3.5 Flash-Lite · Gemini

It does not name the material unknowns that could change the roadmap, such as whether FitPhysio will accept the dashboard or whether reclaimed support capacity will materialise.

Avoids unsupported claimsWrongRightWrong
Gemini 3.8 Flash · API

The output presents the booking page conversion as 'well above healthcare industry benchmarks' as an established fact, but no benchmark data is provided in the context.

Sonnet 5.5 · API

Hypotheses about reminders and cancellations are clearly labelled as suggestions or expectations, and no confident claim goes beyond the supplied evidence.

Gemini 3.5 Flash-Lite · Gemini

It presents unestablished causal and capacity claims as fact, including that texting 'drops' no-shows and that sync fixes 'permanently' reclaim engineering time.

Produces the required deliverableMixedRightWrong
Gemini 3.8 Flash · API

The roadmap memo is complete, under 800 words, addressed to the required readers, and could be acted on with light edits.

Sonnet 5.5 · API

The document is a complete roadmap with a plan, what's not being done, risks, and decisions needed, all in a form the CPO and squad leads can act on.

Gemini 3.5 Flash-Lite · Gemini

Although it is a roadmap memo under 800 words, it is not usable as-is because the plan exceeds capacity and leaves the Physitrack trade-off unresolved.

Every commitment serves the goalsMixedRightWrong
Gemini 3.8 Flash · API

Deposits at Booking Phase 1 is committed in Q1 but is not tied to either cutting no-shows or reducing multi-site churn, nor is it presented as freeing capacity.

Sonnet 5.5 · API

Every firmly committed item (messaging, reminders, dashboard, calendar sync) directly serves the no-show or multi-site churn goals or frees capacity; the booking page redesign, waitlist auto-fill, deposits, and Physitrack are deferred with reasons given.

Gemini 3.5 Flash-Lite · Gemini

Physitrack is committed on its own without being tied to the two CEO goals or to freeing capacity, and the dashboard is not clearly positioned as the substitute for the Sales promise.

Handles the Sales promise with evidenceRightRightWrong
Gemini 3.8 Flash · API

The roadmap uses FitPhysio's own notes to show the multi-site dashboard is their must-have, defers Physitrack, and gives Sales a clear line to lead with the dashboard.

Sonnet 5.5 · API

It uses FitPhysio's own must-have/nice-to-have notes to show the dashboard closes the deal, recommends Sales lead with the dashboard and delay Physitrack, and gives a clear line for the conversation.

Gemini 3.5 Flash-Lite · Gemini

It notes FitPhysio's must-have is the dashboard but does not give Sales a clear line on Physitrack, such as later or only if something else is cut.

Fits the stated capacityMixedRightWrong
Gemini 3.8 Flash · API

The committed work sums to exactly the available capacity (20 weeks in Q4, 23 in Q1) with no slack, while the brief expects some slack in the plan.

Sonnet 5.5 · API

The committed work sums to 20 squad-weeks in Q4 and 12 firm weeks in Q1 (plus conditional 8), well within the 40 total, with explicit slack and a list of what was cut or deferred.

Gemini 3.5 Flash-Lite · Gemini

The committed items sum to 44 squad-weeks against 40 available, and the claimed reclaimed capacity does not make the plan fit.

All got right 2

Outcomes, with certainty that falls with distanceRightRightRight
Gemini 3.8 Flash · API

Every item names its outcome or problem, near-term items are specific with weeks and scope, and later items like 'Integration Ecosystem' are deliberately looser.

Sonnet 5.5 · API

Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).

Gemini 3.5 Flash-Lite · Gemini

Each item names an outcome or problem, and later items are looser than near-term items.

Sequences around dependenciesRightRightRight
Gemini 3.8 Flash · API

All dependencies are respected: messaging service before SMS reminders and waitlist, SMS reminders before waitlist auto-fill, deposits after contract signing, and calendar sync fix before capacity increase.

Sonnet 5.5 · API

Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.

Gemini 3.5 Flash-Lite · Gemini

Messaging is scheduled before SMS reminders and waitlist, and deposits are deferred until after the contract date.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI89.6100.02None
2GPT-6 AstrawithChatGPT91.787.52None
3GPT-6.1 SolwithAPI91.784.62None
4GPT-6 LunawithAPI85.076.32None
5Opus 5.5withClaude80.565.12None
6Gemini 3.8 FlashwithAPI70.151.92None
7Gemini 3.5 Flash-LitewithGemini28.48.322 capped

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review