Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Outcomes, with certainty that falls with distance95% pass
    Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).
    Sonnet 5.5 · API · Two squads, eight asks, one half
  2. Sequences around dependencies89% pass
    Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.
    Sonnet 5.5 · API · Two squads, eight asks, one half
  3. Plans on the squads we actually have89% pass
    It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.
    GPT-6 Astra · ChatGPT · A year of spend management, with a hard deadline

Where it slips

  1. Makes the call on procurement50% pass
    It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.
    GPT-6 Luna · API · A year of spend management, with a hard deadline
  2. Produces the required deliverable50% pass
    The roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Uses the supplied evidence correctly61% pass
    Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.
    GPT-6 Astra · ChatGPT · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the PM for Tidewell's booking product. Write the roadmap for the next two quarters (Q4 2026 and Q1 2027) and what comes after, for Dana Okafor, our CPO, and the two squad leads. It will also be shared with Sales and Design, so it needs to say what we're not doing and why. Keep it under 800 words. Everything we know is below.

What the model was given7 items: About Tidewell, Goals for the next two quarters (set by the CEO), Data, Capacity, Candidate work (estimates in squad-weeks, from the squad leads), Sales note, FitPhysio's evaluation notes (shared by their operations director)
About TidewellOnline booking and scheduling for independent physiotherapy clinics. 1,450 clinics, $3.9M ARR.
Goals for the next two quarters (set by the CEO)1. Cut no-shows for our clinics: it's the main thing we promise them. 2. Cut churn among multi-site clinics.
DataNo-shows average 11% of appointments across all clinics. The 40 clinics that text their own patients a reminder the day before average 6%. Multi-site clinics are 16% of clinics and 38% of ARR; their monthly logo churn is 2.9%, against 1.3% for single-site clinics. In exit surveys, 23 of the 31 multi-site clinics that left in the last year cited 'can't see all our locations in one place'. Booking page: 64% of visits end in a booking. In 18 customer interviews this quarter, nobody mentioned the booking page.
CapacityTwo squads. Each has 13 weeks a quarter, but about a quarter of each squad's time goes on support and bugs, so we plan on 10 squad-weeks of roadmap work per squad per quarter.
Candidate work (estimates in squad-weeks, from the squad leads)1. Calendar sync fix, 5. Calendar sync failures cause 38% of support tickets; engineering expects fixing them to free about 3 squad-weeks a quarter of support time, starting the quarter after it ships. 2. Messaging service (SMS provider, patient consent records, templates), 8. Nothing patient-facing on its own. 3. SMS reminders with confirm or cancel, 6. Needs the messaging service. 4. Multi-site dashboard (every location's bookings and utilisation in one view), 10. 5. Waitlist auto-fill: texts waitlisted patients when a slot frees up, 7. Needs the messaging service. It only fills slots cancelled with some notice, which is rare today: patients don't cancel, they just don't turn up. 6. Deposits at booking, 14. Needs a payments partner contract; Legal says the earliest it can be signed is mid-February 2027. 7. Booking page redesign, 9. Requested by the Head of Design. 8. Physitrack integration (sends patients' exercise plans to the Physitrack app), 8. See the Sales note.
Sales noteFrom Jordan, Head of Sales: “FitPhysio Group (30 sites, $64,800 ARR) is our biggest deal this year. I told them Physitrack would be live by Christmas. It's what closes the deal.”
FitPhysio's evaluation notes (shared by their operations director)“Must have: see all 30 sites' bookings and utilisation in one place. Nice to have: Physitrack integration.”
What a strong answer doesThe answer key the graders mark against

A two-quarter roadmap that fits about 40 squad-weeks (10 per squad per quarter), with items framed as the outcome they serve. Puts the calendar sync fix first because it frees capacity for Q1, builds the messaging service before reminders and the waitlist, and sequences reminders before the waitlist since slots only free up once patients can cancel. Commits the multi-site dashboard in the half: it serves the churn goal and is FitPhysio's own must-have, so it tells Sales to lead with the dashboard rather than promising Physitrack by Christmas. Leaves deposits for after the contract (it can't be built in the half once signing is mid-February) and doesn't commit the booking page redesign without evidence. A committed plan of about 36 squad-weeks (sync 5, messaging 8, dashboard 10, reminders 6, waitlist 7) leaves a little slack; Physitrack only fits if something else goes. Near-term items are specific, later ones looser, and it names what's not being done and why.

Critical failures (cap the score at 40)
  • Schedules an item before something it depends on: SMS reminders or waitlist auto-fill before the messaging service, or deposits shipping in the half
  • Commits clearly more than about 40 squad-weeks of work to the two quarters
Case

Mid-level · v1.1 · synthetic · B2B SaaS, healthcare, capacity, sales promise

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 11 checks passed

A PM could use this roadmap with light edits because the commitments, capacity, dependencies, deferrals, and Sales message are clear. The main gap is that Q1 reserved capacity is not yet sized into specific follow-up work.

Output

Tidewell booking roadmap: Q4 2026–Q1 2027

For Dana Okafor, CPO; squad leads; Sales and Design

Direction

Prioritise fewer no-shows and lower multi-site churn, not a broader feature catalogue.

  • No-shows average 11%. The 40 clinics sending their own reminders average 6%: promising evidence for reminders, but not proof of causation.
  • Multi-site clinics represent 38% of ARR. Their monthly logo churn is 2.9%, versus 1.3% for single-site clinics. Of 31 multi-site exits, 23 cited the lack of an all-location view.

These findings support two primary investments: multi-site visibility and SMS reminders. Fixing calendar sync makes both more sustainable by reducing support demand.

Q4 2026: retain multi-site clinics and build the reminder foundation

SquadRoadmap allocationDeliverable
------:---
ACalendar sync fix — 5 squad-weeksShip first; address failures responsible for 38% of support tickets.
AMessaging service — 5 of 8 squad-weeksBegin provider integration, consent records and templates. Not a customer-facing launch.
BMulti-site dashboard — 10 squad-weeksDeliver bookings and utilisation across locations in one view.

Total: 20 squad-weeks, matching available roadmap capacity. Messaging carries three weeks into Q1.

Pilot the dashboard with multi-site operators, including FitPhysio if they agree, before broad release. Design should prioritise cross-location workflows and dashboard usability rather than a booking-page refresh.

Calendar sync is a capacity investment as well as a reliability fix. Engineering forecasts three additional squad-weeks per quarter from Q1; we will verify that reduction rather than spend it in advance.

Q1 2027: launch reminders and prove outcomes

SquadRoadmap allocationDeliverable
------:---
AFinish messaging — 3 squad-weeksComplete the prerequisite service.
ASMS reminders with confirm/cancel — 6 squad-weeksPilot, then expand based on results and consent readiness.
A1 squad-week reservedLaunch hardening and fixes.
B10 squad-weeks reservedDashboard adoption and retention improvements, reminder rollout support, and evidence-led follow-up work.

Committed feature build: nine squad-weeks; baseline capacity: 20. The remaining capacity is an explicit investment envelope, not an estimate for unspecified features. Squad leads will size follow-up work after dashboard feedback and reminder pilot results. Any recovered support capacity is additional headroom.

Do not fill this envelope with lower-priority commitments merely to reach full utilisation.

Success measures

  • Reminders: compare no-show rates for pilot clinics with their baseline and a comparable holdout where practical. Track delivery, confirmations, cancellations and cancellation notice. Set rollout thresholds before launch; do not promise the observed 6% rate.
  • Dashboard: track active multi-site usage, ability to complete cross-location tasks, and location-visibility complaints.
  • Retention: monitor multi-site monthly logo churn against the 2.9% baseline. Expect adoption signals before a reliable churn verdict.
  • Calendar sync: track sync-related tickets and actual support time recovered.

Not committed—and why

  • Physitrack integration (8 weeks): not on the Christmas roadmap. FitPhysio’s documented must-have is the dashboard; Physitrack is a nice-to-have. Jordan and Product must promptly correct the promise, confirm the buying criteria with FitPhysio, and offer a dashboard-led evaluation. Do not replace an unsupported date with another.
  • Booking page redesign (9 weeks): 64% conversion alone does not establish a problem, and none of 18 interviews raised it. Revisit with funnel evidence or usability findings.
  • Deposits (14 weeks): the partner contract cannot be signed before mid-February 2027. Too large and dependency-constrained for a credible Q1 commitment.
  • Waitlist auto-fill (7 weeks): today, cancellations with notice are rare. First establish whether reminders create enough recoverable slots. Consider Q1 only if evidence and remaining delivery time support it.

After Q1

Prioritise waitlist auto-fill if timely cancellations increase; otherwise improve reminder effectiveness and dashboard adoption. Reassess deposits after the contract and an operational/business case are ready. Reconsider Physitrack against verified demand and opportunity cost. No dates beyond Q1 until outcome evidence and estimates justify them.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly82%
  • passAddresses the actual decision98%
  • passRespects explicit constraints38%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims72%
  • passProduces the required deliverable30%
  • passEvery commitment serves the goals97%
  • passHandles the Sales promise with evidence100%
  • passOutcomes, with certainty that falls with distance84%
  • passFits the stated capacity85%
  • passSequences around dependencies73%
Run
Run
#1
API response time
40 s
Submitted
30 Sept 2026

Needs more than a quick edit5 of 11 checks passed

The output contains an unsupported benchmark claim, lacks slack in the capacity plan, and commits deposits without tying it to the goals, requiring rework of evidence and reasoning before use.

Output

To: Dana Okafor (CPO), Squad Leads CC: Sales, Design From: Product Manager, Booking & Scheduling Subject: Product Roadmap: Q4 2026 – Q1 2027 (and Beyond)

---

Executive Summary

Our mandate is clear: cut no-shows and halt multi-site clinic churn.

Multi-site clinics represent 38% of our ARR ($1.48M) but churn at more than double the rate of single-site clinics (2.9% vs. 1.3%/month). Concurrently, no-shows cost our clinics 11% of their billable volume, despite evidence that pre-visit SMS reminders cut this to 6%.

This roadmap directly tackles these two threats while paying down technical debt to unlock 15% more product development capacity starting in Q1 2027.

---

Capacity Model

  • Standard Capacity: 2 squads × 10 roadmap weeks/quarter = 20 squad-weeks/quarter (3 weeks/squad reserved for support/maintenance).
  • Q4 2026: 20 squad-weeks available.
  • Q1 2027: 23 squad-weeks available (the Q4 calendar sync fix recovers 3 squad-weeks of support overhead per quarter).

---

Q4 2026: Multi-Site Retention & Core Platform

Squad 1: Clinic Operations (10 squad-weeks) * Multi-site Dashboard (10 weeks): Builds unified multi-location visibility into bookings and practitioner utilisation. * Rationale: 74% of churned multi-site clinics (23 of 31) cited the lack of a cross-location view as their reason for leaving. It is also the sole “Must Have” listed by FitPhysio Group ($64.8k ARR).

Squad 2: Platform & Messaging (10 squad-weeks) * Calendar Sync Fix (5 weeks): Resolves root causes of sync failures. * Rationale: Drives 38% of all support tickets. Resolving this unlocks 3 squad-weeks of engineering capacity every quarter starting Q1 2027. * Messaging Service Foundation (5 of 8 weeks): SMS gateway provider integration, consent management, and compliance architecture. * Rationale: Mandatory architectural foundation for automated patient communications in Q1.

---

Q1 2027: Slashing No-Shows & Revenue Recovery

Total Capacity: 23 squad-weeks (includes +3 weeks unlocked from Calendar Sync).

Squad 1: Patient Engagement (11 squad-weeks) * Complete Messaging Service (3 weeks): Finalise templates and delivery logging. * SMS Reminders with Confirm/Cancel (6 weeks): Automated 24-hour reminders with bidirectional text response. * Rationale: Targets our #1 clinic promise. The 40 clinics currently texting reminders manually see no-shows drop from 11% to 6%. * Deposit Discovery & Technical Spec (2 weeks): Lay groundwork for payment gateway integration once legal clears contracts in mid-February.

Squad 2: Clinic Automation (12 squad-weeks) * Waitlist Auto-Fill (7 weeks): Automatically texts waitlisted patients when slots open. * Rationale: SMS confirm/cancel in Squad 1 will generate advance cancellations. Auto-fill instantly monetises those open slots. * Deposits at Booking – Phase 1 (5 of 14 weeks): Begin implementation mid-February once payment contracts execute.

---

What Comes After (Q2 2027+)

  • Deposits at Booking – Phase 2 (9 weeks remaining): Complete deposit workflows to reduce late cancellations and guarantee revenue for booked slots.
  • Integration Ecosystem (Physitrack / EHRs): Standardised webhooks and partner app marketplace based on clinic demand.

---

What We Are Not Doing (And Why)

#### 1. Physitrack Integration (8 weeks) * The Context: Sales requested this to close FitPhysio Group ($64.8k ARR) before Christmas. * Why We Are Deferring: FitPhysio’s operations team explicitly evaluated Physitrack as a “Nice to have,” whereas the Multi-site Dashboard is their sole “Must have.” Building Physitrack in Q4 would force us to drop the Multi-site Dashboard, risking both the FitPhysio close and the broader 38% of our ARR currently vulnerable to churn. We will re-evaluate partner integrations in Q2 2027.

#### 2. Booking Page Redesign (9 weeks) * The Context: Proposed by Design to refresh the booking interface. * Why We Are Deferring: Our current page converts at 64%—well above healthcare industry benchmarks. In 18 clinic interviews this past quarter, not one customer raised booking page friction. Spending 9 weeks here does not move our core metrics of reducing no-shows or stopping churn.

#### 3. Deposits at Booking in Q4 2026 * Why We Are Waiting: Legal confirmed the payments partner contract cannot be signed until mid-February 2027. We will not commit engineering resources ahead of executed legal agreements.

#### 4. Waitlist Auto-Fill in Q4 2026 * Why We Are Sequencing in Q1: Waitlist auto-fill requires cancellations with advance notice to work. Today, patients simply no-show. We must launch SMS confirm/cancel first to create the advance openings that waitlist auto-fill recovers.

Claims the judge couldn’t find in the brief

  • Booking page conversion is well above healthcare industry benchmarks

Grades and run details

Decision model 82 · LLM judge 6 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly17%
  • passAddresses the actual decision70%
  • passRespects explicit constraints11%
  • partialIdentifies material uncertainty32%
  • partialAvoids unsupported claims52%
  • failProduces the required deliverable14%
  • passEvery commitment serves the goals56%
  • passHandles the Sales promise with evidence82%
  • passOutcomes, with certainty that falls with distance36%
  • passFits the stated capacity23%
  • passSequences around dependencies70%
Run
Run
#1
API response time
38 s
Submitted
1 Oct 2026

Needs more than a quick edit9 of 11 checks passed

A PM could use most of it, but the Q1 calendar sync placement undermines the capacity logic and would need resequencing before it is fully decision-ready.

Output

Tidewell Booking Roadmap | Q4 2026–Q1 2027

For: Dana Okafor, CPO; squad leads; Sales and Design North star: Reduce clinic no-shows and multi-site churn. We have 20 roadmap squad-weeks per quarter across two squads; support and bugs are planned separately.

Q4 2026 — Build the multi-site view and messaging foundation

  • Multi-site dashboard — 10 squad-weeks. Show bookings and utilisation across every location in one view. This directly addresses the most-cited multi-site exit reason (23 of 31 surveyed departures) and FitPhysio’s stated must-have.
  • Messaging service — 8 squad-weeks. Build the SMS provider, patient consent records and templates needed for reminders. This is foundational work, not a patient-facing launch; reminders follow in Q1.

Capacity: One squad takes the dashboard; the other builds messaging, with two squad-weeks uncommitted.

Sales: Jordan’s Christmas commitment for Physitrack is not in this roadmap. FitPhysio called it “nice to have,” while the multi-site view is a must-have. Please correct the commitment with FitPhysio now; do not sell a Christmas delivery date.

Q1 2027 — Launch reminders and fix calendar sync

  • SMS reminders with confirm/cancel — 6 squad-weeks. Launch using the Q4 messaging foundation, with rollout and measurement. Forty clinics that send their own day-before texts average 6% no-shows versus 11% overall. That’s a promising signal, not proof of causation; measure results against a comparison group and track cancellations as well as no-shows.
  • Calendar sync fix — 5 squad-weeks. Address a cause of 38% of support tickets. Engineering expects it to free about three squad-weeks of support time per quarter, starting Q2.

The remaining Q1 roadmap capacity is reserved for rollout learning and iteration rather than adding another large commitment.

Success measures: No-show rate for reminder-enabled appointments versus comparison; reminder delivery, confirmations and cancellations; dashboard adoption across multi-site locations; and multi-site support and churn trends.

After Q1 — Reassess before committing

  • Waitlist auto-fill: Revisit once reminders give patients a way to cancel and we know whether cancellations happen with enough notice. Today, patients mostly don’t cancel, so auto-fill has few openings to act on.
  • Deposits: Reassess after the payments contract can be signed (earliest mid-February) and we have evidence on reminder impact and clinic appetite. The 14-week estimate and partner dependency make it unsuitable for this committed plan.
  • Physitrack: Reassess against customer evidence and capacity; it is not committed for Q4 or Q1.
  • Booking page redesign: Revisit if research identifies a meaningful problem. Booking-page conversion is 64%, and nobody in 18 recent interviews raised it; it is not the best use of capacity against current goals.

Explicitly not doing now

We are not building the booking redesign, waitlist auto-fill, deposits or Physitrack in these two quarters. This is a focus decision: prioritize the evidenced churn driver and test the clearest no-show intervention before investing in lower-evidence or dependency-heavy work.

Grades and run details

Decision model 91 · LLM judge 10 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly27%
  • passAddresses the actual decision81%
  • failRespects explicit constraints11%
  • passIdentifies material uncertainty80%
  • passAvoids unsupported claims53%
  • passProduces the required deliverable15%
  • passEvery commitment serves the goals91%
  • passHandles the Sales promise with evidence100%
  • passOutcomes, with certainty that falls with distance85%
  • passFits the stated capacity32%
  • passSequences around dependencies81%
Run
Run
#1
API response time
31 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 7

Uses the supplied evidence correctlyRightMixedRight
GPT-6.1 Sol · API

The output’s current-situation facts, figures, customer statements, estimates, dates, and capacity numbers all come from the supplied context or follow directly from it.

Gemini 3.8 Flash · API

The output claims the booking page conversion is 'well above healthcare industry benchmarks,' which is not in the supplied context and is an unsupported statement about the current situation.

GPT-6 Luna · API

The output uses the supplied numbers, dates, customer statements, and capacity facts correctly, and its few inferences follow from the context.

Respects explicit constraintsRightRightMixed
GPT-6.1 Sol · API

It is addressed to the named readers, stays under 800 words, covers Q4/Q1 and after, says what is not being done and why, and its allocations fit the stated capacity.

Gemini 3.8 Flash · API

The output is under 800 words, in memo form, addressed to the named readers, and includes what is not being done and why.

GPT-6 Luna · API

It is under 800 words, addressed to the named readers, includes what is not being done and why, and respects the stated capacity and dependencies.

Identifies material uncertaintyRightWrongRight
GPT-6.1 Sol · API

It names material unknowns such as reminder causation, recovered support time, dashboard adoption, timely cancellations, and Physitrack demand, and says how they would be resolved.

Gemini 3.8 Flash · API

The output does not name specific unknowns that could change the decision, such as delays in the messaging service or the calendar sync fix not freeing the expected capacity, nor does it say how they would be resolved.

GPT-6 Luna · API

It names material unknowns such as reminder causation, cancellation notice, contract timing, and booking-page evidence, and says how they would be resolved.

Avoids unsupported claimsRightWrongRight
GPT-6.1 Sol · API

Interpretations and forecasts are labelled as evidence, assumptions, or verification items rather than presented as established facts.

Gemini 3.8 Flash · API

The output presents the booking page conversion as 'well above healthcare industry benchmarks' as an established fact, but no benchmark data is provided in the context.

GPT-6 Luna · API

It labels the 6% versus 11% signal as not proof of causation and frames later items as reassessments rather than committed facts.

Produces the required deliverableRightMixedRight
GPT-6.1 Sol · API

It is a usable roadmap for the CPO, squad leads, Sales, and Design, with near-term commitments, success measures, deferrals, and post-Q1 direction.

Gemini 3.8 Flash · API

The roadmap memo is complete, under 800 words, addressed to the required readers, and could be acted on with light edits.

GPT-6 Luna · API

It is a usable roadmap for the CPO, squad leads, Sales, and Design, with committed work, deferred work, success measures, and a sales message.

Every commitment serves the goalsRightMixedMixed
GPT-6.1 Sol · API

Every committed item serves no-show reduction, multi-site churn reduction, or capacity freed for those goals, and non-serving items are deferred with reasons.

Gemini 3.8 Flash · API

Deposits at Booking Phase 1 is committed in Q1 but is not tied to either cutting no-shows or reducing multi-site churn, nor is it presented as freeing capacity.

GPT-6 Luna · API

The calendar sync fix is committed in Q1 but only frees support capacity starting Q2, so it does not serve the two-quarter goals or free capacity within the half as required.

Fits the stated capacityRightMixedRight
GPT-6.1 Sol · API

The committed work totals 29 squad-weeks against 40 available, with checkable quarter sums and explicit slack/reserved capacity.

Gemini 3.8 Flash · API

The committed work sums to exactly the available capacity (20 weeks in Q4, 23 in Q1) with no slack, while the brief expects some slack in the plan.

GPT-6 Luna · API

The committed work totals 29 squad-weeks against 40 available, with explicit slack and deferrals to make it fit.

All got right 4

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

It commits to a clear two-quarter roadmap and named deferrals, and states the evidence or conditions that would change later choices.

Gemini 3.8 Flash · API

The output commits to a clear two-quarter roadmap with specific items, says what is not being done and why, and is framed for the CPO and squad leads.

GPT-6 Luna · API

It commits to a clear two-quarter roadmap for the named readers and states conditions that would change later decisions, though it could be stronger on what would change the committed plan itself.

Handles the Sales promise with evidenceRightRightRight
GPT-6.1 Sol · API

It uses FitPhysio’s own must-have/nice-to-have notes to tell Sales to correct the Physitrack promise and lead with the dashboard.

Gemini 3.8 Flash · API

The roadmap uses FitPhysio's own notes to show the multi-site dashboard is their must-have, defers Physitrack, and gives Sales a clear line to lead with the dashboard.

GPT-6 Luna · API

It uses FitPhysio's own must-have/nice-to-have notes to tell Sales to correct the Christmas Physitrack promise and lead with the committed dashboard.

Outcomes, with certainty that falls with distanceRightRightRight
GPT-6.1 Sol · API

Items are framed by outcomes, near-term work is specific, and later work is deliberately looser pending evidence.

Gemini 3.8 Flash · API

Every item names its outcome or problem, near-term items are specific with weeks and scope, and later items like 'Integration Ecosystem' are deliberately looser.

GPT-6 Luna · API

Each item names the problem or outcome it serves, and later items are deliberately looser than near-term commitments.

Sequences around dependenciesRightRightRight
GPT-6.1 Sol · API

Messaging precedes reminders, deposits are deferred until after the contract, and the dashboard is committed before Physitrack is reconsidered.

Gemini 3.8 Flash · API

All dependencies are respected: messaging service before SMS reminders and waitlist, SMS reminders before waitlist auto-fill, deposits after contract signing, and calendar sync fix before capacity increase.

GPT-6 Luna · API

It places reminders after the messaging service, defers waitlist until reminders create cancellations, and defers deposits until the contract can be signed.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI89.6100.02None
2GPT-6 AstrawithChatGPT91.787.52None
3GPT-6.1 SolwithAPI91.784.62None
4GPT-6 LunawithAPI85.076.32None
5Opus 5.5withClaude80.565.12None
6Gemini 3.8 FlashwithAPI70.151.92None
7Gemini 3.5 Flash-LitewithGemini28.48.322 capped

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review