Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Outcomes, with certainty that falls with distance95% pass
    Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).
    Sonnet 5.5 · API · Two squads, eight asks, one half
  2. Sequences around dependencies89% pass
    Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.
    Sonnet 5.5 · API · Two squads, eight asks, one half
  3. Plans on the squads we actually have89% pass
    It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.
    GPT-6 Astra · ChatGPT · A year of spend management, with a hard deadline

Where it slips

  1. Makes the call on procurement50% pass
    It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.
    GPT-6 Luna · API · A year of spend management, with a hard deadline
  2. Produces the required deliverable50% pass
    The roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Uses the supplied evidence correctly61% pass
    Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.
    GPT-6 Astra · ChatGPT · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Tidewell's booking product. Write the roadmap for the next two quarters (Q4 2026 and Q1 2027) and what comes after, for Dana Okafor, our CPO, and the two squad leads. It will also be shared with Sales and Design, so it needs to say what we're not doing and why. Keep it under 800 words. Everything we know is below.

What the model was given7 items: About Tidewell, Goals for the next two quarters (set by the CEO), Data, Capacity, Candidate work (estimates in squad-weeks, from the squad leads), Sales note, FitPhysio's evaluation notes (shared by their operations director)
About TidewellOnline booking and scheduling for independent physiotherapy clinics. 1,450 clinics, $3.9M ARR.
Goals for the next two quarters (set by the CEO)1. Cut no-shows for our clinics: it's the main thing we promise them. 2. Cut churn among multi-site clinics.
DataNo-shows average 11% of appointments across all clinics. The 40 clinics that text their own patients a reminder the day before average 6%. Multi-site clinics are 16% of clinics and 38% of ARR; their monthly logo churn is 2.9%, against 1.3% for single-site clinics. In exit surveys, 23 of the 31 multi-site clinics that left in the last year cited 'can't see all our locations in one place'. Booking page: 64% of visits end in a booking. In 18 customer interviews this quarter, nobody mentioned the booking page.
CapacityTwo squads. Each has 13 weeks a quarter, but about a quarter of each squad's time goes on support and bugs, so we plan on 10 squad-weeks of roadmap work per squad per quarter.
Candidate work (estimates in squad-weeks, from the squad leads)1. Calendar sync fix, 5. Calendar sync failures cause 38% of support tickets; engineering expects fixing them to free about 3 squad-weeks a quarter of support time, starting the quarter after it ships. 2. Messaging service (SMS provider, patient consent records, templates), 8. Nothing patient-facing on its own. 3. SMS reminders with confirm or cancel, 6. Needs the messaging service. 4. Multi-site dashboard (every location's bookings and utilisation in one view), 10. 5. Waitlist auto-fill: texts waitlisted patients when a slot frees up, 7. Needs the messaging service. It only fills slots cancelled with some notice, which is rare today: patients don't cancel, they just don't turn up. 6. Deposits at booking, 14. Needs a payments partner contract; Legal says the earliest it can be signed is mid-February 2027. 7. Booking page redesign, 9. Requested by the Head of Design. 8. Physitrack integration (sends patients' exercise plans to the Physitrack app), 8. See the Sales note.
Sales noteFrom Jordan, Head of Sales: “FitPhysio Group (30 sites, $64,800 ARR) is our biggest deal this year. I told them Physitrack would be live by Christmas. It's what closes the deal.”
FitPhysio's evaluation notes (shared by their operations director)“Must have: see all 30 sites' bookings and utilisation in one place. Nice to have: Physitrack integration.”
What a strong answer doesThe answer key the graders mark against

A two-quarter roadmap that fits about 40 squad-weeks (10 per squad per quarter), with items framed as the outcome they serve. Puts the calendar sync fix first because it frees capacity for Q1, builds the messaging service before reminders and the waitlist, and sequences reminders before the waitlist since slots only free up once patients can cancel. Commits the multi-site dashboard in the half: it serves the churn goal and is FitPhysio's own must-have, so it tells Sales to lead with the dashboard rather than promising Physitrack by Christmas. Leaves deposits for after the contract (it can't be built in the half once signing is mid-February) and doesn't commit the booking page redesign without evidence. A committed plan of about 36 squad-weeks (sync 5, messaging 8, dashboard 10, reminders 6, waitlist 7) leaves a little slack; Physitrack only fits if something else goes. Near-term items are specific, later ones looser, and it names what's not being done and why.

Critical failures (cap the score at 40)
  • Schedules an item before something it depends on: SMS reminders or waitlist auto-fill before the messaging service, or deposits shipping in the half
  • Commits clearly more than about 40 squad-weeks of work to the two quarters
Case

Mid-level · v1.1 · synthetic · B2B SaaS, healthcare, capacity, sales promise

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit9 of 11 checks passed

A PM could use most of it, but the Q1 calendar sync placement undermines the capacity logic and would need resequencing before it is fully decision-ready.

Output

Tidewell Booking Roadmap | Q4 2026–Q1 2027

For: Dana Okafor, CPO; squad leads; Sales and Design North star: Reduce clinic no-shows and multi-site churn. We have 20 roadmap squad-weeks per quarter across two squads; support and bugs are planned separately.

Q4 2026 — Build the multi-site view and messaging foundation

  • Multi-site dashboard — 10 squad-weeks. Show bookings and utilisation across every location in one view. This directly addresses the most-cited multi-site exit reason (23 of 31 surveyed departures) and FitPhysio’s stated must-have.
  • Messaging service — 8 squad-weeks. Build the SMS provider, patient consent records and templates needed for reminders. This is foundational work, not a patient-facing launch; reminders follow in Q1.

Capacity: One squad takes the dashboard; the other builds messaging, with two squad-weeks uncommitted.

Sales: Jordan’s Christmas commitment for Physitrack is not in this roadmap. FitPhysio called it “nice to have,” while the multi-site view is a must-have. Please correct the commitment with FitPhysio now; do not sell a Christmas delivery date.

Q1 2027 — Launch reminders and fix calendar sync

  • SMS reminders with confirm/cancel — 6 squad-weeks. Launch using the Q4 messaging foundation, with rollout and measurement. Forty clinics that send their own day-before texts average 6% no-shows versus 11% overall. That’s a promising signal, not proof of causation; measure results against a comparison group and track cancellations as well as no-shows.
  • Calendar sync fix — 5 squad-weeks. Address a cause of 38% of support tickets. Engineering expects it to free about three squad-weeks of support time per quarter, starting Q2.

The remaining Q1 roadmap capacity is reserved for rollout learning and iteration rather than adding another large commitment.

Success measures: No-show rate for reminder-enabled appointments versus comparison; reminder delivery, confirmations and cancellations; dashboard adoption across multi-site locations; and multi-site support and churn trends.

After Q1 — Reassess before committing

  • Waitlist auto-fill: Revisit once reminders give patients a way to cancel and we know whether cancellations happen with enough notice. Today, patients mostly don’t cancel, so auto-fill has few openings to act on.
  • Deposits: Reassess after the payments contract can be signed (earliest mid-February) and we have evidence on reminder impact and clinic appetite. The 14-week estimate and partner dependency make it unsuitable for this committed plan.
  • Physitrack: Reassess against customer evidence and capacity; it is not committed for Q4 or Q1.
  • Booking page redesign: Revisit if research identifies a meaningful problem. Booking-page conversion is 64%, and nobody in 18 recent interviews raised it; it is not the best use of capacity against current goals.

Explicitly not doing now

We are not building the booking redesign, waitlist auto-fill, deposits or Physitrack in these two quarters. This is a focus decision: prioritize the evidenced churn driver and test the clearest no-show intervention before investing in lower-evidence or dependency-heavy work.

Grades and run details

Decision model 91 · LLM judge 10 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly27%
  • passAddresses the actual decision81%
  • failRespects explicit constraints11%
  • passIdentifies material uncertainty80%
  • passAvoids unsupported claims53%
  • passProduces the required deliverable15%
  • passEvery commitment serves the goals91%
  • passHandles the Sales promise with evidence100%
  • passOutcomes, with certainty that falls with distance85%
  • passFits the stated capacity32%
  • passSequences around dependencies81%
Run
Run
#1
API response time
31 s
Submitted
1 Oct 2026

Usable with a quick edit9 of 11 checks passed

The output is well-structured and evidence-based; the only flaw is a missing explicit tie of Physitrack to the churn goal, which a PM could fix with a one-sentence edit without redoing any reasoning.

Output

Tidewell Booking: Roadmap, Q4 2026 to Q1 2027

For: Dana Okafor (CPO), Squad Leads · Shared with: Sales, Design

The short version

We have two goals: cut no-shows, and cut multi-site churn. The evidence points clearly to one fix for each.

  • No-shows: Clinics that text patients the day before see 6% no-shows against 11% overall. SMS reminders are the fix.
  • Multi-site churn: Multi-site clinics churn at 2.9% a month against 1.3% for single-site, and they are 38% of ARR. 23 of 31 who left cited "can't see all our locations in one place." The multi-site dashboard is the fix.

Everything else is either groundwork for these two, or waits.

Capacity

Each squad has 10 squad-weeks of roadmap work a quarter, so 20 in total. If the calendar sync fix ships in Q4, it frees about 3 squad-weeks from Q1 onward, giving roughly 23.

Q4 2026 (20 of 20 squad-weeks)

SquadWorkWeeksWhy
AMulti-site dashboard10Goal 2. Top exit reason. FitPhysio's must-have.
BCalendar sync fix538% of support tickets. Frees about 3 weeks a quarter from Q1.
BMessaging service (part 1)5Foundation for reminders.

Risk: The dashboard takes all of Squad A's quarter, so it lands late December. Any slip pushes it into January. We will cut scope before we miss the date, for example by shipping bookings first and utilisation second.

Q1 2027 (17 of about 23 squad-weeks)

SquadWorkWeeksWhy
BMessaging service (part 2)3Finishes the foundation.
BSMS reminders with confirm/cancel6Goal 1. Target launch mid-March.
APhysitrack integration8Commitment to FitPhysio (see below).
Buffer~6These are first estimates on new work. Covers slippage, then early waitlist work.

What success looks like

  • No-shows: Clinics using reminders move toward the 6% seen by self-texting clinics. We'll measure from launch.
  • Multi-site churn: Monthly churn falls from 2.9% toward the single-site rate. Exit-survey mentions of "one place" drop.
  • Support load: Calendar sync tickets fall by most of their current 38% share.

FitPhysio and Physitrack: a date change is needed

Jordan told FitPhysio that Physitrack would be live by Christmas. We can't deliver that without dropping either the dashboard or reminders, and those are the CEO's two goals.

FitPhysio's own evaluation notes list the dashboard as their must-have and Physitrack as nice-to-have. The plan above gives them the thing they said they need by year-end, with Physitrack following in Q1.

Ask for Jordan: Please reset the Physitrack date with FitPhysio to Q1 2027, leading with the dashboard landing in December. I'm happy to join that call.

Going forward, please check dates with product before they go to customers. That way Sales can commit with confidence.

What we're not doing, and why

Booking page redesign (9 weeks): not planned. - 64% of visits already end in a booking. - Across 18 customer interviews this quarter, nobody raised it. - It doesn't serve either goal.

Design: we'd value your help on the dashboard and on the reminder messages patients receive. Those are where design effort changes outcomes this half. We'll revisit the redesign if conversion drops or customers start asking for it.

Waitlist auto-fill (7 weeks): after Q1. - It only fills slots cancelled with notice, and today patients rarely cancel. They just don't turn up. - Reminders with a cancel button should change that. Once reminders are live, waitlist becomes worth building. It is first in line for the Q1 buffer or for Q2.

Deposits at booking (14 weeks): not before Q2, and not yet committed. - The payments contract can't be signed before mid-February 2027, so building can't start this half. - It also adds friction for patients. We'll decide after we see what reminders do to no-shows. If reminders get us close to 6%, deposits may not be needed.

After Q1 2027 (indicative)

  1. Waitlist auto-fill: Builds on reminders and the messaging service.
  2. Deposits decision: Based on reminder results and the signed contract.
  3. Multi-site follow-ons: Based on dashboard usage and churn data, such as cross-site reporting and permissions.
  4. Booking page: Only if the data changes.

Decisions needed

  • Dana: Approve the plan, and the Physitrack date reset with FitPhysio.
  • Jordan: Own the FitPhysio conversation this week.
  • Squad leads: Confirm estimates, and the dashboard scope-cut plan in case of slippage.

Grades and run details

Decision model 82 · LLM judge 11 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly56%
  • passAddresses the actual decision91%
  • passRespects explicit constraints14%
  • passIdentifies material uncertainty82%
  • partialAvoids unsupported claims17%
  • failProduces the required deliverable11%
  • partialEvery commitment serves the goals65%
  • passHandles the Sales promise with evidence99%
  • passOutcomes, with certainty that falls with distance83%
  • passFits the stated capacity39%
  • passSequences around dependencies69%
Run
Run
#1
Time to output
47 s
Submitted
30 Sept 2026
Critical failure: Commits clearly more than about 40 squad-weeks of work to the two quarters. Score capped at 40.

Needs more than a quick edit2 of 11 checks passed

A PM would need to rework the capacity math and decide whether Physitrack is committed or deferred. The evidence and sequencing reasoning also need correction before it can be used.

Output

To: Dana Okafor (CPO), Squad Leads, Sales, Design From: PM, Tidewell Booking Subject: Booking Product Roadmap: Q4 2026 – Q1 2027 & Beyond

At $3.9M ARR across 1,450 independent physio clinics, our immediate focus is clear: ruthlessly align our capacity with the CEO's goals of cutting no-shows and reducing multi-site churn.

With two squads and 10 productive weeks per squad-quarter (20 total squad-weeks per quarter), here is our execution plan.

---

Q4 2026: Foundation & The Multi-Site Anchor

Goal Alignment: Multi-site churn, Support capacity.

  • Multi-Site Dashboard (10 squad-weeks – Squad A): A single view of every location’s bookings and utilization.
  • Why: Multi-site clinics drive 38% of our ARR ($1.48M) but churn at 2.9% (vs. 1.3% single-site). "Can't see all locations in one place" was cited in 74% of multi-site exit surveys. This is also the explicit "Must Have" for FitPhysio Group (30 sites, $64.8K ARR).
  • Messaging Service (8 squad-weeks – Squad B): Infrastructure for SMS provider, consent, and templates.
  • Why: Prerequisite for our Q1 no-show reduction engine.
  • Calendar Sync Fix (5 squad-weeks – Squad B, spilling into Q1): Resolving sync failures that drive 38% of support tickets.
  • Why: Engineering estimates this will permanently reclaim ~3 squad-weeks/quarter of engineering time starting Q1.

---

Q1 2027: Solving No-Shows

Goal Alignment: No-show reduction.

  • Calendar Sync Fix completion (0 extra weeks, Squad B finishing up).
  • SMS Reminders with Confirm/Cancel (6 squad-weeks – Squad B): Leveraging the Q4 Messaging Service.
  • Why: Our data proves this works: clinics manually texting reminders drop no-shows from 11% to 6%. Automated confirm/cancel turns passive ghosting into active rescheduling, opening slots early.
  • Physitrack Integration (8 squad-weeks – Squad A): Sends exercise plans directly to the Physitrack app.
  • Why: Sales commitment to close FitPhysio Group by Christmas. (Note: While FitPhysio listed this as "nice-to-have" vs. the dashboard's "must-have," it is a commercial blocker per Sales).
  • Waitlist Auto-Fill (7 squad-weeks – Squad A, utilizing reclaimed support capacity): Texts waitlisted patients when slots free up.
  • Why: Directly targets no-shows, but depends on SMS reminders and cancellations working first.

---

What Comes After (Q2 2027 & Beyond)

  • Deposits at Booking (14 squad-weeks): Paused until Legal secures the payments partner contract (earliest mid-February 2027). This will be our primary Q2/Q3 bet to further crush no-shows once infrastructure is live.

---

What We Are NOT Doing & Why

1. Booking Page Redesign (9 squad-weeks): * Why Not: 64% of visits already convert to bookings, and 18 customer interviews this quarter yielded zero mentions of the booking page. Spending design capacity here is a distraction from our retention and no-show crises. 2. Deposits in Q4/Q1: * Why Not: Blocked by external legal/contract timelines until mid-Q1 at best. 3. Prioritizing Physitrack over Multi-Site in Q4: * Why Not: While Sales needs Physitrack for FitPhysio, FitPhysio’s actual operational blocker is the Multi-Site Dashboard. Building the dashboard first secures our largest pipeline deal and addresses the core churn driver for our highest-value segment.

Claims the judge couldn’t find in the brief

  • Fixing calendar sync will permanently reclaim about 3 squad-weeks per quarter of engineering time starting Q1.
  • Clinics manually texting reminders drop no-shows from 11% to 6%.

Grades and run details

Decision model 32 · LLM judge 2 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly13%
  • partialAddresses the actual decision30%
  • failRespects explicit constraints44%
  • failIdentifies material uncertainty25%
  • partialAvoids unsupported claims43%
  • failProduces the required deliverable27%
  • partialEvery commitment serves the goals44%
  • failHandles the Sales promise with evidence23%
  • passOutcomes, with certainty that falls with distance34%
  • failFits the stated capacity57%
  • passSequences around dependencies38%
Run
Run
#1
Time to output
4 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 9

Uses the supplied evidence correctlyRightRightWrong
GPT-6 Luna · API

The output uses the supplied numbers, dates, customer statements, and capacity facts correctly, and its few inferences follow from the context.

Opus 5.5 · Claude

Every statement about the current situation is directly from the brief or follows from arithmetic, with no invented facts.

Gemini 3.5 Flash-Lite · Gemini

It treats correlation as causation for SMS reminders and adds unsupported 'permanently reclaim' language about calendar sync capacity.

Addresses the actual decisionRightRightWrong
GPT-6 Luna · API

It commits to a clear two-quarter roadmap for the named readers and states conditions that would change later decisions, though it could be stronger on what would change the committed plan itself.

Opus 5.5 · Claude

The output commits to a clear two-quarter plan, addresses the Physitrack promise, and says what is not being done and why, framed for the named readers.

Gemini 3.5 Flash-Lite · Gemini

It commits to a roadmap but does not clearly state what condition would change the answer, especially around the Physitrack commitment and capacity trade-off.

Respects explicit constraintsMixedRightWrong
GPT-6 Luna · API

It is under 800 words, addressed to the named readers, includes what is not being done and why, and respects the stated capacity and dependencies.

Opus 5.5 · Claude

The output is under 800 words, addresses the required readers, and says what is not being done and why, respecting all stated constraints.

Gemini 3.5 Flash-Lite · Gemini

It violates the stated capacity constraint by committing 44 squad-weeks to two quarters despite saying capacity is 40.

Identifies material uncertaintyRightRightWrong
GPT-6 Luna · API

It names material unknowns such as reminder causation, cancellation notice, contract timing, and booking-page evidence, and says how they would be resolved.

Opus 5.5 · Claude

It names risks (dashboard slip, estimate uncertainty), uses a buffer, and ties later decisions (deposits, waitlist, booking page) to results from earlier work.

Gemini 3.5 Flash-Lite · Gemini

It does not name the material unknowns that could change the roadmap, such as whether FitPhysio will accept the dashboard or whether reclaimed support capacity will materialise.

Avoids unsupported claimsRightRightWrong
GPT-6 Luna · API

It labels the 6% versus 11% signal as not proof of causation and frames later items as reassessments rather than committed facts.

Opus 5.5 · Claude

Forecasts and interpretations are clearly labeled as such (e.g., 'should change that'), and no unsupported facts are presented as established.

Gemini 3.5 Flash-Lite · Gemini

It presents unestablished causal and capacity claims as fact, including that texting 'drops' no-shows and that sync fixes 'permanently' reclaim engineering time.

Produces the required deliverableRightMixedWrong
GPT-6 Luna · API

It is a usable roadmap for the CPO, squad leads, Sales, and Design, with committed work, deferred work, success measures, and a sales message.

Opus 5.5 · Claude

The roadmap is complete, in the right form, under 800 words, and usable by Dana, squad leads, Sales, and Design with minimal edits.

Gemini 3.5 Flash-Lite · Gemini

Although it is a roadmap memo under 800 words, it is not usable as-is because the plan exceeds capacity and leaves the Physitrack trade-off unresolved.

Every commitment serves the goalsMixedWrongWrong
GPT-6 Luna · API

The calendar sync fix is committed in Q1 but only frees support capacity starting Q2, so it does not serve the two-quarter goals or free capacity within the half as required.

Opus 5.5 · Claude

Physitrack integration is committed in Q1 but is not tied to either CEO goal (no-shows or multi-site churn) or to freeing capacity, and the criterion requires such items to be deferred with a reason.

Gemini 3.5 Flash-Lite · Gemini

Physitrack is committed on its own without being tied to the two CEO goals or to freeing capacity, and the dashboard is not clearly positioned as the substitute for the Sales promise.

Handles the Sales promise with evidenceRightRightWrong
GPT-6 Luna · API

It uses FitPhysio's own must-have/nice-to-have notes to tell Sales to correct the Christmas Physitrack promise and lead with the committed dashboard.

Opus 5.5 · Claude

It uses FitPhysio's own notes to show the dashboard is the must-have, gives Sales a clear line to reset the Physitrack date to Q1, and asks Jordan to own the conversation.

Gemini 3.5 Flash-Lite · Gemini

It notes FitPhysio's must-have is the dashboard but does not give Sales a clear line on Physitrack, such as later or only if something else is cut.

Fits the stated capacityRightRightWrong
GPT-6 Luna · API

The committed work totals 29 squad-weeks against 40 available, with explicit slack and deferrals to make it fit.

Opus 5.5 · Claude

Committed work sums to 37 squad-weeks (20 in Q4, 17 in Q1) against 40-43 available, with a buffer, and it names what was cut or deferred.

Gemini 3.5 Flash-Lite · Gemini

The committed items sum to 44 squad-weeks against 40 available, and the claimed reclaimed capacity does not make the plan fit.

All got right 2

Outcomes, with certainty that falls with distanceRightRightRight
GPT-6 Luna · API

Each item names the problem or outcome it serves, and later items are deliberately looser than near-term commitments.

Opus 5.5 · Claude

Every item names its outcome or problem, near-term items are specific, and later items (after Q1) are deliberately looser and indicative.

Gemini 3.5 Flash-Lite · Gemini

Each item names an outcome or problem, and later items are looser than near-term items.

Sequences around dependenciesRightRightRight
GPT-6 Luna · API

It places reminders after the messaging service, defers waitlist until reminders create cancellations, and defers deposits until the contract can be signed.

Opus 5.5 · Claude

Messaging service is built before SMS reminders, calendar sync fix precedes the freed capacity, deposits are not in the half, and waitlist is after reminders.

Gemini 3.5 Flash-Lite · Gemini

Messaging is scheduled before SMS reminders and waitlist, and deposits are deferred until after the contract date.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI89.6100.02None
2GPT-6 AstrawithChatGPT91.787.52None
3GPT-6.1 SolwithAPI91.784.62None
4GPT-6 LunawithAPI85.076.32None
5Opus 5.5withClaude80.565.12None
6Gemini 3.8 FlashwithAPI70.151.92None
7Gemini 3.5 Flash-LitewithGemini28.48.322 capped

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review