Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Outcomes, with certainty that falls with distance95% pass
    Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).
    Sonnet 5.5 · API · Two squads, eight asks, one half
  2. Sequences around dependencies89% pass
    Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.
    Sonnet 5.5 · API · Two squads, eight asks, one half
  3. Plans on the squads we actually have89% pass
    It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.
    GPT-6 Astra · ChatGPT · A year of spend management, with a hard deadline

Where it slips

  1. Makes the call on procurement50% pass
    It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.
    GPT-6 Luna · API · A year of spend management, with a hard deadline
  2. Produces the required deliverable50% pass
    The roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Uses the supplied evidence correctly61% pass
    Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.
    GPT-6 Astra · ChatGPT · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Tidewell's booking product. Write the roadmap for the next two quarters (Q4 2026 and Q1 2027) and what comes after, for Dana Okafor, our CPO, and the two squad leads. It will also be shared with Sales and Design, so it needs to say what we're not doing and why. Keep it under 800 words. Everything we know is below.

What the model was given7 items: About Tidewell, Goals for the next two quarters (set by the CEO), Data, Capacity, Candidate work (estimates in squad-weeks, from the squad leads), Sales note, FitPhysio's evaluation notes (shared by their operations director)
About TidewellOnline booking and scheduling for independent physiotherapy clinics. 1,450 clinics, $3.9M ARR.
Goals for the next two quarters (set by the CEO)1. Cut no-shows for our clinics: it's the main thing we promise them. 2. Cut churn among multi-site clinics.
DataNo-shows average 11% of appointments across all clinics. The 40 clinics that text their own patients a reminder the day before average 6%. Multi-site clinics are 16% of clinics and 38% of ARR; their monthly logo churn is 2.9%, against 1.3% for single-site clinics. In exit surveys, 23 of the 31 multi-site clinics that left in the last year cited 'can't see all our locations in one place'. Booking page: 64% of visits end in a booking. In 18 customer interviews this quarter, nobody mentioned the booking page.
CapacityTwo squads. Each has 13 weeks a quarter, but about a quarter of each squad's time goes on support and bugs, so we plan on 10 squad-weeks of roadmap work per squad per quarter.
Candidate work (estimates in squad-weeks, from the squad leads)1. Calendar sync fix, 5. Calendar sync failures cause 38% of support tickets; engineering expects fixing them to free about 3 squad-weeks a quarter of support time, starting the quarter after it ships. 2. Messaging service (SMS provider, patient consent records, templates), 8. Nothing patient-facing on its own. 3. SMS reminders with confirm or cancel, 6. Needs the messaging service. 4. Multi-site dashboard (every location's bookings and utilisation in one view), 10. 5. Waitlist auto-fill: texts waitlisted patients when a slot frees up, 7. Needs the messaging service. It only fills slots cancelled with some notice, which is rare today: patients don't cancel, they just don't turn up. 6. Deposits at booking, 14. Needs a payments partner contract; Legal says the earliest it can be signed is mid-February 2027. 7. Booking page redesign, 9. Requested by the Head of Design. 8. Physitrack integration (sends patients' exercise plans to the Physitrack app), 8. See the Sales note.
Sales noteFrom Jordan, Head of Sales: “FitPhysio Group (30 sites, $64,800 ARR) is our biggest deal this year. I told them Physitrack would be live by Christmas. It's what closes the deal.”
FitPhysio's evaluation notes (shared by their operations director)“Must have: see all 30 sites' bookings and utilisation in one place. Nice to have: Physitrack integration.”
What a strong answer doesThe answer key the graders mark against

A two-quarter roadmap that fits about 40 squad-weeks (10 per squad per quarter), with items framed as the outcome they serve. Puts the calendar sync fix first because it frees capacity for Q1, builds the messaging service before reminders and the waitlist, and sequences reminders before the waitlist since slots only free up once patients can cancel. Commits the multi-site dashboard in the half: it serves the churn goal and is FitPhysio's own must-have, so it tells Sales to lead with the dashboard rather than promising Physitrack by Christmas. Leaves deposits for after the contract (it can't be built in the half once signing is mid-February) and doesn't commit the booking page redesign without evidence. A committed plan of about 36 squad-weeks (sync 5, messaging 8, dashboard 10, reminders 6, waitlist 7) leaves a little slack; Physitrack only fits if something else goes. Near-term items are specific, later ones looser, and it names what's not being done and why.

Critical failures (cap the score at 40)
  • Schedules an item before something it depends on: SMS reminders or waitlist auto-fill before the messaging service, or deposits shipping in the half
  • Commits clearly more than about 40 squad-weeks of work to the two quarters
Case

Mid-level · v1.1 · synthetic · B2B SaaS, healthcare, capacity, sales promise

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit9 of 11 checks passed

The roadmap is strong and usable, but the unsupported 'undermines booking reliability' claim and the slightly confusing Q1 '20/20' allocation that includes uncommitted ceiling work need a quick correction before it is fully reliable.

Output

Tidewell booking roadmap: Q4 2026–Q1 2027

For Dana Okafor, CPO; both squad leads; Sales and Design

Strategy and trade-offs

Prioritise fewer missed appointments and better multi-site retention, not a broader feature set.

  • No-shows average 11%. Clinics sending their own reminders average 6%—a promising signal, not proof that reminders alone cause the difference.
  • Multi-site clinics represent 38% of ARR and churn at 2.9% monthly, versus 1.3% for single-site clinics. Of 31 multi-site departures, 23 cited the lack of a consolidated view.
  • Fix calendar sync early: it causes 38% of support tickets, undermines booking reliability, and could release capacity for subsequent work.

Q4 2026: retain multi-site clinics and establish the reminder foundation

SquadCommitted workSquad-weeks
---------:
AShip multi-site dashboard: bookings and utilisation across locations10
BShip calendar sync fix5
BBegin messaging service: provider, consent records and templates5 of 8
Total20 / 20

Sequencing: Squad B tackles calendar sync first, then messaging. Messaging remains unfinished at quarter-end; we are not promising patient-facing SMS in Q4.

Dashboard rollout: Validate workflows with multi-site operators, including FitPhysio, before broad release. Track activation, repeated cross-location use and whether operators can answer their core bookings/utilisation questions without separate location views.

Calendar acceptance: Confirm failures and related tickets decline. Engineering’s expected three squad-weeks of quarterly savings are a forecast, not guaranteed capacity.

Q1 2027: launch reminders and validate outcomes

SquadPlanned workSquad-weeks
---------:
BComplete messaging service3
BBuild and roll out SMS reminders with confirm/cancel6
ADashboard adoption, measured follow-up improvements and outcome evaluationUp to 10
BReminder rollout contingency1
Total baseline allocation20 / 20

Squad A’s allocation is a budget ceiling, not ten weeks of pre-approved features. Dana and the leads will approve specific improvements from dashboard usage and customer evidence; unused capacity stays uncommitted rather than funding lower-priority scope.

The calendar fix could increase Q1 capacity to 23 squad-weeks across both squads. Confirm actual support savings before allocating the extra three weeks; use them first for rollout risks and demonstrated retention blockers.

Reminder sequencing: Finish messaging before reminder implementation. With nine squad-weeks on this dependency chain, plan for a late-Q1 rollout, not a January launch.

Measurement: - Pilot reminders with a comparison group where practical. Measure no-shows, advance cancellations, delivery failures and opt-outs; do not promise the observational 6% result. - Track multi-site dashboard adoption and monthly churn against the 2.9% baseline. Review churn reasons and rolling trends; the small population makes short-term attribution uncertain. - Establish baselines and experiment design before rollout. Expect fuller reminder-impact evidence in Q2.

Not doing in these two quarters

  • Physitrack integration: Not committed. FitPhysio calls it nice to have; the dashboard is its must-have. Dana and Jordan should immediately correct the Christmas promise with FitPhysio, confirm whether its buying criteria have changed, and offer the dashboard delivery plan. An integration commitment requires an explicit roadmap trade-off—not hidden extra work.
  • Waitlist auto-fill: Defer until reminders demonstrably create cancellations with enough notice to refill. Today, its prerequisite behaviour is rare.
  • Deposits: No Q1 delivery promise. The contract cannot be signed before mid-February, and the 14-week build is a substantial commitment before testing the simpler reminder intervention.
  • Booking page redesign: Defer. The 64% conversion rate alone does not establish a problem, and none of 18 interviews raised it. Design should support dashboard and reminder usability first.

What comes after

Q2 2027 onward: Scale effective reminders and address remaining multi-site retention gaps. Prioritise waitlist auto-fill if advance cancellations create sufficient refill opportunities. Reassess deposits if no-shows remain high and the partner contract is ready. Revisit Physitrack and booking redesign only with stronger customer or commercial evidence.

Claims the judge couldn’t find in the brief

  • Calendar sync failures undermine booking reliability.

Grades and run details

Decision model 100 · LLM judge 9 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly85%
  • passAddresses the actual decision96%
  • passRespects explicit constraints56%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims75%
  • passProduces the required deliverable34%
  • passEvery commitment serves the goals97%
  • passHandles the Sales promise with evidence100%
  • passOutcomes, with certainty that falls with distance78%
  • passFits the stated capacity47%
  • passSequences around dependencies93%
Run
Run
#1
Time to output
47 s
Submitted
1 Oct 2026

Usable with a quick edit11 of 11 checks passed

The roadmap is clear, actionable, and addresses all required elements; a PM could use it with only minor wording edits.

Output

Tidewell Booking Roadmap: Q4 2026, Q1 2027 and beyond

For: Dana Okafor, squad leads. Shared with Sales and Design.

Goals and what the data says

  1. Cut no-shows. The average is 11%. The 40 clinics that text their own patients a reminder average 6%. That group is small and self-selected, so this suggests reminders help but doesn't prove it. We'll measure with a holdout group.
  2. Cut multi-site churn. Multi-site clinics are 16% of clinics but 38% of ARR (about $1.5M). Their monthly churn is 2.9%, against 1.3% for single-site. Of the 31 that left last year, 23 cited "can't see all our locations in one place."

Capacity

There are 10 squad-weeks per squad per quarter, so 20 a quarter. Calendar sync ships in Q1, so its freed support time (about 3 squad-weeks) only adds capacity from Q2, bringing Q2 to about 23.

The plan

Squad ASquad B
Q4 2026 (20)Messaging service (8), then start SMS reminders (2)Multi-site dashboard (10)
Q1 2027 (20)Finish SMS reminders (4), calendar sync fix (5), 1 sparePhysitrack integration (8, conditional, see below), 2 for dashboard follow-up
Q2 2027 (~23)Deposits (14), start once the contract is signedWaitlist auto-fill (7), if the trigger below is met

Q4: build both foundations. - Multi-site dashboard (10). This directly targets the top churn reason, and it is FitPhysio's stated must-have. Target: live mid-December. It fills Squad B's quarter with no buffer, so if it slips, we roll out to FitPhysio's sites in phases rather than wait. - Messaging service (8). SMS provider, consent records and templates. Nothing patient-facing ships from it alone, but every no-show feature depends on it.

Q1: ship reminders and pay down support load. - SMS reminders with confirm or cancel (6). This is our main no-show feature. It should ship in January. Success is measured against a holdout group, with the target being movement toward the 6% the self-reminding clinics achieve. - Calendar sync fix (5). Sync failures cause 38% of support tickets. The fix pays for itself in under two quarters. It goes after reminders because reminders serve the CEO's goals directly. - Physitrack integration (8). See the FitPhysio section below.

Q2 and after. - Deposits (14). Deposits should cut no-shows, but Legal can't sign the payments partner contract before mid-February, and 14 weeks of work doesn't fit in what remains of Q1. Action: start the contract process now so the build begins the day it's signed. - Waitlist auto-fill (7). Only build it if reminders prove the case. It fills slots cancelled with notice, and today patients mostly just don't show up. Confirm or cancel reminders should turn some no-shows into early cancellations. Trigger: the share of appointments cancelled with notice rises meaningfully in the first 6–8 weeks after reminders launch.

FitPhysio and Physitrack

FitPhysio (30 sites, $64,800 ARR) is our biggest deal this year. Sales promised Physitrack by Christmas. FitPhysio's own evaluation notes say the dashboard is a must-have and Physitrack is nice-to-have.

  • The dashboard is on track for Christmas, and it is the thing that closes the deal.
  • Physitrack by Christmas would cost us. Squad B is full with the dashboard, so the 8 weeks would have to come from messaging. That delays reminders, our main no-show feature, by a full quarter, for a nice-to-have for one account.
  • Recommendation: Jordan goes back to FitPhysio now. The message is that the dashboard, their must-have, is live before Christmas, and Physitrack follows in Q1. I'd like Jordan and me to have that conversation this week, before the expectation hardens.
  • Physitrack is conditional. We build it in Q1 if FitPhysio signs, or if a second account asks for it. Otherwise that capacity goes to the dashboard follow-up and to buffer.

What we're not doing, and why

  • Booking page redesign (9 weeks, not scheduled). 64% of booking page visits already end in a booking, and none of our 18 customer interviews mentioned the page. It serves neither goal. I'd revisit it if we see evidence of drop-off or patient complaints. Design isn't sidelined, though: the dashboard and the reminder and confirm flows need heavy design input in Q4 and Q1, and I'd like Design embedded with both squads.
  • Waitlist auto-fill now. It would fill few slots today (see the trigger above).
  • Deposits before Q2. Blocked by the contract and too large to fit anyway.
  • Physitrack before Christmas. It would push back our main no-show feature for a nice-to-have.

Risks

  • Dashboard timing. Squad B has no slack in Q4. The mitigation is a phased FitPhysio rollout.
  • Reminder effect. The 11% vs 6% gap may partly reflect which clinics choose to send reminders. The holdout group will give us the real number.
  • Support load. Support and bugs are planned at 25% of time. If they run higher, Q1 calendar sync and Physitrack slip first, before reminders.
  • Legal. If the payments contract slips past mid-February, deposits slip with it one-for-one.

Decisions needed

  1. Dana: approve the Q4 plan and the no-go on the booking page redesign.
  2. Jordan: take the revised Physitrack timing to FitPhysio.
  3. Legal: start the payments partner contract now.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly47%
  • passAddresses the actual decision96%
  • passRespects explicit constraints25%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims29%
  • passProduces the required deliverable17%
  • passEvery commitment serves the goals28%
  • passHandles the Sales promise with evidence100%
  • passOutcomes, with certainty that falls with distance88%
  • passFits the stated capacity45%
  • passSequences around dependencies70%
Run
Run
#1
API response time
39 s
Submitted
30 Sept 2026

Usable with a quick edit9 of 11 checks passed

The output is well-structured and evidence-based; the only flaw is a missing explicit tie of Physitrack to the churn goal, which a PM could fix with a one-sentence edit without redoing any reasoning.

Output

Tidewell Booking: Roadmap, Q4 2026 to Q1 2027

For: Dana Okafor (CPO), Squad Leads · Shared with: Sales, Design

The short version

We have two goals: cut no-shows, and cut multi-site churn. The evidence points clearly to one fix for each.

  • No-shows: Clinics that text patients the day before see 6% no-shows against 11% overall. SMS reminders are the fix.
  • Multi-site churn: Multi-site clinics churn at 2.9% a month against 1.3% for single-site, and they are 38% of ARR. 23 of 31 who left cited "can't see all our locations in one place." The multi-site dashboard is the fix.

Everything else is either groundwork for these two, or waits.

Capacity

Each squad has 10 squad-weeks of roadmap work a quarter, so 20 in total. If the calendar sync fix ships in Q4, it frees about 3 squad-weeks from Q1 onward, giving roughly 23.

Q4 2026 (20 of 20 squad-weeks)

SquadWorkWeeksWhy
AMulti-site dashboard10Goal 2. Top exit reason. FitPhysio's must-have.
BCalendar sync fix538% of support tickets. Frees about 3 weeks a quarter from Q1.
BMessaging service (part 1)5Foundation for reminders.

Risk: The dashboard takes all of Squad A's quarter, so it lands late December. Any slip pushes it into January. We will cut scope before we miss the date, for example by shipping bookings first and utilisation second.

Q1 2027 (17 of about 23 squad-weeks)

SquadWorkWeeksWhy
BMessaging service (part 2)3Finishes the foundation.
BSMS reminders with confirm/cancel6Goal 1. Target launch mid-March.
APhysitrack integration8Commitment to FitPhysio (see below).
Buffer~6These are first estimates on new work. Covers slippage, then early waitlist work.

What success looks like

  • No-shows: Clinics using reminders move toward the 6% seen by self-texting clinics. We'll measure from launch.
  • Multi-site churn: Monthly churn falls from 2.9% toward the single-site rate. Exit-survey mentions of "one place" drop.
  • Support load: Calendar sync tickets fall by most of their current 38% share.

FitPhysio and Physitrack: a date change is needed

Jordan told FitPhysio that Physitrack would be live by Christmas. We can't deliver that without dropping either the dashboard or reminders, and those are the CEO's two goals.

FitPhysio's own evaluation notes list the dashboard as their must-have and Physitrack as nice-to-have. The plan above gives them the thing they said they need by year-end, with Physitrack following in Q1.

Ask for Jordan: Please reset the Physitrack date with FitPhysio to Q1 2027, leading with the dashboard landing in December. I'm happy to join that call.

Going forward, please check dates with product before they go to customers. That way Sales can commit with confidence.

What we're not doing, and why

Booking page redesign (9 weeks): not planned. - 64% of visits already end in a booking. - Across 18 customer interviews this quarter, nobody raised it. - It doesn't serve either goal.

Design: we'd value your help on the dashboard and on the reminder messages patients receive. Those are where design effort changes outcomes this half. We'll revisit the redesign if conversion drops or customers start asking for it.

Waitlist auto-fill (7 weeks): after Q1. - It only fills slots cancelled with notice, and today patients rarely cancel. They just don't turn up. - Reminders with a cancel button should change that. Once reminders are live, waitlist becomes worth building. It is first in line for the Q1 buffer or for Q2.

Deposits at booking (14 weeks): not before Q2, and not yet committed. - The payments contract can't be signed before mid-February 2027, so building can't start this half. - It also adds friction for patients. We'll decide after we see what reminders do to no-shows. If reminders get us close to 6%, deposits may not be needed.

After Q1 2027 (indicative)

  1. Waitlist auto-fill: Builds on reminders and the messaging service.
  2. Deposits decision: Based on reminder results and the signed contract.
  3. Multi-site follow-ons: Based on dashboard usage and churn data, such as cross-site reporting and permissions.
  4. Booking page: Only if the data changes.

Decisions needed

  • Dana: Approve the plan, and the Physitrack date reset with FitPhysio.
  • Jordan: Own the FitPhysio conversation this week.
  • Squad leads: Confirm estimates, and the dashboard scope-cut plan in case of slippage.

Grades and run details

Decision model 82 · LLM judge 11 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly56%
  • passAddresses the actual decision91%
  • passRespects explicit constraints14%
  • passIdentifies material uncertainty82%
  • partialAvoids unsupported claims17%
  • failProduces the required deliverable11%
  • partialEvery commitment serves the goals65%
  • passHandles the Sales promise with evidence99%
  • passOutcomes, with certainty that falls with distance83%
  • passFits the stated capacity39%
  • passSequences around dependencies69%
Run
Run
#1
Time to output
47 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyMixedRightRight
GPT-6 Astra · ChatGPT

Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.

Sonnet 5.5 · API

Every statement about the current situation is taken directly from the brief or derived by arithmetic, with no invented facts.

Opus 5.5 · Claude

Every statement about the current situation is directly from the brief or follows from arithmetic, with no invented facts.

Avoids unsupported claimsMixedRightRight
GPT-6 Astra · ChatGPT

It labels many hypotheses carefully, but still asserts that calendar sync failures undermine booking reliability without supplied evidence.

Sonnet 5.5 · API

Hypotheses about reminders and cancellations are clearly labelled as suggestions or expectations, and no confident claim goes beyond the supplied evidence.

Opus 5.5 · Claude

Forecasts and interpretations are clearly labeled as such (e.g., 'should change that'), and no unsupported facts are presented as established.

Produces the required deliverableRightRightMixed
GPT-6 Astra · ChatGPT

The roadmap is present for Q4 2026, Q1 2027, and after, with squad-week tables, deferrals, and measurement notes that the CPO and squad leads could act on.

Sonnet 5.5 · API

The document is a complete roadmap with a plan, what's not being done, risks, and decisions needed, all in a form the CPO and squad leads can act on.

Opus 5.5 · Claude

The roadmap is complete, in the right form, under 800 words, and usable by Dana, squad leads, Sales, and Design with minimal edits.

Every commitment serves the goalsRightRightWrong
GPT-6 Astra · ChatGPT

Committed items serve no-show reduction, multi-site churn, or capacity release; Physitrack, waitlist, deposits, and booking redesign are deferred with reasons.

Sonnet 5.5 · API

Every firmly committed item (messaging, reminders, dashboard, calendar sync) directly serves the no-show or multi-site churn goals or frees capacity; the booking page redesign, waitlist auto-fill, deposits, and Physitrack are deferred with reasons given.

Opus 5.5 · Claude

Physitrack integration is committed in Q1 but is not tied to either CEO goal (no-shows or multi-site churn) or to freeing capacity, and the criterion requires such items to be deferred with a reason.

All got right 7

Addresses the actual decisionRightRightRight
GPT-6 Astra · ChatGPT

It commits early to prioritising no-show reduction and multi-site retention, with a clear two-quarter plan and explicit deferrals, and states conditions that would change later choices.

Sonnet 5.5 · API

The roadmap commits to a clear two-quarter plan with conditional items, names what is not being done and why, and states what would change the Physitrack and waitlist decisions.

Opus 5.5 · Claude

The output commits to a clear two-quarter plan, addresses the Physitrack promise, and says what is not being done and why, framed for the named readers.

Respects explicit constraintsRightRightRight
GPT-6 Astra · ChatGPT

It is addressed to the named readers, stays under 800 words, includes what is not being done and why, and respects the capacity and dependency constraints in its proposed plan.

Sonnet 5.5 · API

The output is under 800 words, addresses the named readers, includes what is not being done and why, and is a roadmap for the requested period.

Opus 5.5 · Claude

The output is under 800 words, addresses the required readers, and says what is not being done and why, respecting all stated constraints.

Identifies material uncertaintyRightRightRight
GPT-6 Astra · ChatGPT

It names material unknowns such as whether reminders cause the 6% result, whether calendar savings materialise, dashboard adoption, churn attribution, and FitPhysio's buying criteria, and says how they would be checked.

Sonnet 5.5 · API

It identifies the self-selection bias in the 6% no-show figure, the need for a holdout group, the trigger for waitlist auto-fill, and the conditional nature of Physitrack, along with risks like dashboard timing and legal delays.

Opus 5.5 · Claude

It names risks (dashboard slip, estimate uncertainty), uses a buffer, and ties later decisions (deposits, waitlist, booking page) to results from earlier work.

Handles the Sales promise with evidenceRightRightRight
GPT-6 Astra · ChatGPT

It uses FitPhysio's own must-have/nice-to-have notes, commits the dashboard, and tells Sales to correct the Christmas Physitrack promise rather than hide extra work.

Sonnet 5.5 · API

It uses FitPhysio's own must-have/nice-to-have notes to show the dashboard closes the deal, recommends Sales lead with the dashboard and delay Physitrack, and gives a clear line for the conversation.

Opus 5.5 · Claude

It uses FitPhysio's own notes to show the dashboard is the must-have, gives Sales a clear line to reset the Physitrack date to Q1, and asks Jordan to own the conversation.

Outcomes, with certainty that falls with distanceRightRightRight
GPT-6 Astra · ChatGPT

Items are framed by outcomes, near-term work is specific, and later work is deliberately looser and conditional.

Sonnet 5.5 · API

Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).

Opus 5.5 · Claude

Every item names its outcome or problem, near-term items are specific, and later items (after Q1) are deliberately looser and indicative.

Fits the stated capacityRightRightRight
GPT-6 Astra · ChatGPT

The quarter tables sum to 20 squad-weeks each, the known committed work is within capacity, and it names deferred items and uncommitted slack.

Sonnet 5.5 · API

The committed work sums to 20 squad-weeks in Q4 and 12 firm weeks in Q1 (plus conditional 8), well within the 40 total, with explicit slack and a list of what was cut or deferred.

Opus 5.5 · Claude

Committed work sums to 37 squad-weeks (20 in Q4, 17 in Q1) against 40-43 available, with a buffer, and it names what was cut or deferred.

Sequences around dependenciesRightRightRight
GPT-6 Astra · ChatGPT

Messaging is sequenced before reminders, waitlist is deferred until reminders create cancellations, deposits are deferred past the contract date, and Physitrack is not committed.

Sonnet 5.5 · API

Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.

Opus 5.5 · Claude

Messaging service is built before SMS reminders, calendar sync fix precedes the freed capacity, deposits are not in the half, and waitlist is after reminders.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI89.6100.02None
2GPT-6 AstrawithChatGPT91.787.52None
3GPT-6.1 SolwithAPI91.784.62None
4GPT-6 LunawithAPI85.076.32None
5Opus 5.5withClaude80.565.12None
6Gemini 3.8 FlashwithAPI70.151.92None
7Gemini 3.5 Flash-LitewithGemini28.48.322 capped

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review