Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Outcomes, with certainty that falls with distance95% pass
    Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).
    Sonnet 5.5 · API · Two squads, eight asks, one half
  2. Sequences around dependencies89% pass
    Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.
    Sonnet 5.5 · API · Two squads, eight asks, one half
  3. Plans on the squads we actually have89% pass
    It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.
    GPT-6 Astra · ChatGPT · A year of spend management, with a hard deadline

Where it slips

  1. Makes the call on procurement50% pass
    It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.
    GPT-6 Luna · API · A year of spend management, with a hard deadline
  2. Produces the required deliverable50% pass
    The roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Uses the supplied evidence correctly61% pass
    Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.
    GPT-6 Astra · ChatGPT · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Harbor. Tom Achebe, our CEO, has shared his draft roadmap for the next four quarters (Q4 2026 to Q3 2027), and asked you to turn it into the roadmap we take to next week's planning offsite. The exec team (CEO, CFO, CRO and CTO) will read it beforehand. Write: 1. The roadmap itself, by quarter or by now, next and later: what each squad works on and the outcome each item serves. 2. The case for it, in no more than 1,200 words: why this order, what you changed from Tom's draft and why, what we're not doing, and the risks that could change the plan. The pack below is everything we have. Not all of it matters equally.

What the model was given12 items: About Harbor, FY27 goal (approved by the board, September 2026), Squads and capacity, Hiring, Card processor deadline, Candidate work (estimates in squad-weeks, by owning squad), Tom's draft roadmap, Procurement evidence, Card spend evidence, Multi-entity evidence, Integrations evidence, Other asks
About HarborSpend management (corporate cards, expenses and approvals) for companies with 200 to 2,000 employees. 1,100 customers, $86M ARR, average 700 employees per customer. 58% of revenue is interchange on card spend; the rest is subscription. Net revenue retention (NRR) is 104%.
FY27 goal (approved by the board, September 2026)Raise NRR to 112% by the end of Q3 2027 by expanding inside existing customers: more spend on Harbor cards, more entities and more modules. New-logo growth matters, but it comes second this year.
Squads and capacityFive squads: Cards, Expenses, Approvals & Policy, Integrations and Platform. After support and on-call, each has about 11 squad-weeks of roadmap capacity a quarter. Squads can take work from another squad's area, but it takes them about 1.5× as long.
HiringTwo new squads were approved in September 2026. Tom's draft counts the first from Q1 2027 and the second from Q2 2027, both at full capacity. Last year, our three approved squads took 5, 6 and 8 months from approval to their first full sprint, and each delivered about half its capacity in its first quarter.
Card processor deadlineOur card processor retires its v1 API on 31 March 2027. Every card authorisation we process runs through v1 today. Migrating to v2 is about 18 squad-weeks of Cards work, and the processor must then run 4 weeks of certification testing before we can cut over. Certification is the processor's time, not ours, but nothing can change on our side while it runs.
Candidate work (estimates in squad-weeks, by owning squad)1. Processor v2 migration: Cards 18. Hard deadline above. 2. Virtual cards for software subscriptions: Cards 10. Needs v2. 3. Real-time spend controls (declines out-of-policy spend at the point of sale): Cards 8, Approvals & Policy 4. Needs v2. 4. Procurement module (purchase requests and approvals, a new paid add-on): Approvals & Policy 30, Integrations 6. 5. Multi-entity support (subsidiaries under one parent account): Platform 12, Approvals & Policy 10. 6. Sage Intacct integration: Integrations 9. 7. Microsoft Dynamics integration: Integrations 12. 8. SSO and SCIM provisioning: Platform 5. 9. Receipt-matching improvements: Expenses 6. 10. Mobile app rewrite: Expenses 22.
Tom's draft roadmapQ4 2026: Cards: v2 migration. Approvals & Policy: multi-entity. Integrations: Intacct. Platform: multi-entity. Expenses: receipt matching. Q1 2027: Cards: finish migration, start virtual cards. New squad A: procurement module. Integrations: Dynamics. Platform: SSO and SCIM. Q2 2027: Cards: virtual cards, spend controls. New squad A: procurement launch. New squad B: mobile rewrite. Q3 2027: Cards: spend controls. Everyone else: procurement adoption, mobile launch. Tom's note: “Procurement is our path to 112%. I've told the board it can add $6M of ARR in its first year.”
Procurement evidenceSix design partners have used a prototype since June. Two said they would pay for it; the other four already use a dedicated procurement tool and said they'd need a reason to switch. Proposed price: $8 per user per month, for users who raise purchase requests. At our customers, about 10% of employees raise purchase requests.
Card spend evidenceIn QBRs with our 60 largest customers, 26 asked for virtual cards for software subscriptions, and 31 CFOs asked for real-time spend controls. Finance estimates our customers pay about $410M a year of software subscriptions on other cards or by invoice (extrapolated from the 60 QBR accounts). Our net interchange is 1.1% of card spend.
Multi-entity evidence14 of our 60 largest customers have asked for it. Between them they have 52 subsidiaries not on Harbor, and 9 of the 14 have said in writing they would add their subsidiaries once it exists. Customer Success sized it at $2.9M ARR if all 52 joined at their parents' pricing.
Integrations evidence17% of customers use Sage Intacct through a CSV export. Their gross revenue churn is 13% a year, against 7% for customers on our NetSuite integration. Intacct was cited in 31% of lost new-logo deals last year, Dynamics in 8%. 3% of customers use Dynamics.
Other asksCRO: “Five enterprise deals worth $1.2M of pipeline need SSO and SCIM, and 11 existing accounts have it on their security review list.” CTO: “The mobile app is on a framework version that loses support in late 2027, and I want the rewrite done before then.” Support: receipt auto-match is at 71%, and unmatched receipts are the most common expense ticket.
What a strong answer doesThe answer key the graders mark against

A four-quarter roadmap that protects the 31 March deadline: the migration build has to finish by about early March to leave 4 weeks of certification, which is roughly all of Cards' capacity for Q4 and most of Q1 (18 of about 18 squad-weeks available before certification). So it adds help or buffer (Platform or another squad at 1.5×, or cutting everything else from Cards) and moves virtual cards out of Q1: both v2 features start only after cut-over (virtual cards Q2, spend controls Q2–Q3). It treats the new squads as upside: on past hiring they arrive around February to April 2027 and deliver half capacity in their first quarter, so nothing committed depends on them before Q2–Q3. It makes a clear call on procurement: the $6M claim needs about 890 accounts (81% of customers) paying $6,720 a year each (70 requesters × $8 × 12), and only 2 of 6 design partners would pay, so it replaces the 30-week build with a cheap paid test with a threshold, and keeps Approvals & Policy on multi-entity and spend controls, which have stronger expansion evidence ($2.9M sized, 9 of 14 committed in writing; about $1.5M a year of interchange from moving a third of the $410M of subscriptions onto virtual cards). Intacct comes early for its churn; Dynamics later or cut; SSO and SCIM are small and justified; the mobile rewrite waits for a new squad, before the framework's support ends. It lists what changed from Tom's draft, what's not being done, and the risks (certification slipping, hiring slipping), with commitment falling with distance.

Critical failures (cap the score at 40)
  • Leaves the processor migration finishing too late for 4 weeks of certification before 31 March 2027, or schedules virtual cards or real-time spend controls before it
Case

Staff-level · v1.1 · synthetic · B2B SaaS, fintech, capacity, hard deadline, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit5 of 12 checks passed

The output contains a factual error about multi-entity commitments (claiming 9 commitments cover all 52 subsidiaries) and lacks a clear procurement test threshold, so a PM would need to re-check evidence and redo that reasoning.

Output

Part 1: Product Roadmap (Q4 2026 – Q3 2027)

``` ==================================================================================================== SQUAD Q4 2026 Q1 2027 Q2 2027 Q3 2027 ==================================================================================================== Cards Processor v2 Migration Finish v2 (Wks 1-7) Virtual Cards for SaaS Real-Time Controls (11 wks) API Freeze/Cert (W8-11) (6 wks remaining) (Cards: 3 wks) Virtual Cards Prep Real-Time Controls Interchange Optimization (4 wks) (Cards: 5 wks) & Scaling (8 wks) ---------------------------------------------------------------------------------------------------- Platform Multi-Entity: Core Multi-Entity: Wrap-up Platform Scaling & Cross-Entity Reporting Data Architecture (1 wk) Org Hierarchies & Enterprise Admin (11 wks) SSO & SCIM (5 wks) (11 wks) (11 wks) Hardening (5 wks) ---------------------------------------------------------------------------------------------------- Approvals & Multi-Entity: Policy Multi-Entity: Release Real-Time Controls: Procurement Discovery Policy & Approval Routing (2 wks) Policy Engine (4 wks) & Scoping (MVP) (10 wks) A&P Rule Engine Complex Approval Rules (Approvals: 11 wks) Buffer (1 wk) Upgrades (9 wks) (7 wks) ---------------------------------------------------------------------------------------------------- Integrations Sage Intacct Native Sage Intacct GA (2 wks) MS Dynamics Integration MS Dynamics GA (3 wks) Integration MS Dynamics Integration (Integrations: 10 wks) ERP Webhooks & Custom (9 wks) + Buffer (2 wk) (Integrations: 2 wks) Sync (8 wks) ---------------------------------------------------------------------------------------------------- Expenses Receipt Auto-Match Mobile App Rewrite: Mobile App Rewrite: Mobile App GA & Optimization (6 wks) Phase 1 Core Workflows Phase 2 Parity/Release Expense Policy Polish Mobile Prep (5 wks) (11 wks) (11 wks) (11 wks) ---------------------------------------------------------------------------------------------------- New Squad A [Hiring / Onboarding] [Ramp / Training] Procurement Foundation Procurement Workflows (A&P Focus) (Ramped: ~5.5 wks) (Full capacity: 11 wks) ---------------------------------------------------------------------------------------------------- New Squad B [Hiring / Pipeline] [Hiring] [Onboarding / Ramp] Integrations / Core (Ecosystem Focus) (Ramped: ~5.5 wks) ==================================================================================================== ```

Squad-by-Squad Outcomes

  • Cards:
  • Processor v2 Migration: Protects 100% of Harbor’s card-authorisation volume ($49.9M existing ARR) against the hard 31 March 2027 processor shutoff; unlocks v2 capabilities required for virtual cards and real-time controls.
  • Virtual Cards for Software Subscriptions: Captures existing customer SaaS spend currently paid via invoice/other cards ($410M addressable pool), directly driving interchange revenue to expand NRR.
  • Real-Time Spend Controls: Enables 31 top-tier CFO accounts to enforce budgets at point-of-sale, unlocking enterprise card issuance and shifting unmanaged invoice spend to Harbor cards.
  • Platform:
  • Multi-Entity Architecture: Unlocks $2.9M contracted expansion ARR across 52 non-Harbor subsidiaries of top accounts.
  • SSO & SCIM Provisioning: Unblocks $1.2M in enterprise pipeline and closes compliance gaps for 11 critical renewal accounts.
  • Approvals & Policy:
  • Multi-Entity Approvals: Delivers cross-subsidiary approval chains required to close existing multi-entity commitments.
  • Real-Time Policy Engine: Co-delivers immediate POS budget validation with the Cards squad.
  • Procurement Discovery & MVP Architecture: Reframes procurement around verified buyer demand without starving committed delivery.
  • Integrations:
  • Sage Intacct Integration: Halves the 13% annual churn rate of the 17% of Harbor customers on Intacct (retaining ~$875k ARR annually) and addresses Harbor's #1 new-logo loss reason (31% of lost deals).
  • Microsoft Dynamics Integration: Expands ERP coverage for mid-market accounts, removing friction for 3% of existing customers and 8% of pipeline.
  • Expenses:
  • Receipt-Matching Improvements: Drives match rate above 71%, eliminating Harbor’s highest-volume support ticket category and improving user retention.
  • Mobile App Rewrite: Eliminates framework deprecation risk ahead of the late-2027 vendor cutoff without stealing capacity from revenue-generating squads.

---

Part 2: The Case for the Roadmap

Executive Summary

To achieve our board-approved target of 112% NRR by Q3 2027 ($86M ARR base expanding by ~$6.9M), Harbor must focus on the highest-probability, shortest-time-to-value expansion levers within our current 1,100 customers: multi-entity expansion ($2.9M ARR), capturing existing software spend via virtual cards ($1.3M–$2.2M ARR), and plugging Sage Intacct churn ($875k ARR retained).

Tom’s draft relied on an unbuilt Procurement module to generate $6M ARR in FY27. Ground-level capacity, vendor deadlines, and design partner data show that strategy carries severe revenue and operational risk. The proposed roadmap secures our existential card processor deadline, resolves major customer expansion requests, and establishes realistic hiring and product timelines.

---

What Changed from Tom’s Draft and Why

#### 1. Realignment on Procurement ($6M ARR Myth vs. Reality) * The Math: Harbor has 1,100 customers averaging 700 employees (770,000 total users). With 10% raising purchase requests (70 users/account) at $8/user/month ($96/year), an adopting customer generates $6,720 ARR. To generate $6M ARR, 893 out of 1,100 customers (81% of our entire base) would have to buy and deploy a v1 module in under 9 months. * The Customer Signal: 4 out of 6 design partners already use dedicated tools (e.g., Coupa, Zip) and explicitly stated they will not switch. Only 2 of 6 expressed willingness to pay. * The Decision: Procurement cannot carry our FY27 NRR goal. We have reallocated primary capacity to guaranteed expansion drivers (Multi-Entity and Virtual Cards), while initiating a structured discovery and phased MVP build with Squad A.

#### 2. Realistic Hiring Capacity * Tom’s Assumption: Squad A operates at 100% capacity in Q1 2027 (11 squad-weeks) and Squad B in Q2 2027 (11 squad-weeks). * Historical Reality: Approvals to first full sprint took 5–8 months historically, delivering ~50% capacity in quarter one. Squads approved in September 2026 will not reach operational productivity until late Q1 / early Q2 2027. * The Decision: No critical-path milestone or revenue target in Q4 or Q1 relies on unhired squads. Existing squads carry core commitments; Squad A ramps on Procurement foundations in Q2 (delivering ~5.5 squad-weeks) and reaches full delivery in Q3.

#### 3. Strict Compliance with the 31 March 2027 Processor Deadline * The Constraint: The v1 API deprecates on 31 March 2027. Migrating takes 18 squad-weeks plus a mandatory 4-week processor certification freeze during which our codebase cannot change. * Execution Window: The 4-week certification must run in March 2027. Cards code complete must occur by Week 7 of Q1 2027. * The Allocation: Cards dedicates all 11 squad-weeks in Q4 2026 and 7 squad-weeks in Q1 2027 (18 weeks total). During the 4-week March freeze, the Cards squad transitions to architectural prep for Virtual Cards, avoiding wasted idle time.

#### 4. Resequencing Multi-Entity, SSO/SCIM, and Sage Intacct * Multi-Entity (Platform + Approvals): Prioritised immediately in Q4. 14 of our 60 largest accounts requested this; 9 have signed written commitments to add 52 subsidiaries ($2.9M ARR). Shipping in early Q1 allows two full quarters of rollout and revenue recognition. * Sage Intacct (Integrations): 17% of our customers use Intacct via CSV, churning at 13% vs. NetSuite’s 7%. Cutting churn across these 187 accounts ($14.6M ARR base) saves ~$875k in ARR retention while eliminating our top sales loss reason (31% of lost deals). * SSO & SCIM (Platform): Sequenced for Q1 immediately following core Multi-Entity work. For a 5-week lift, this unlocks $1.2M in stalled enterprise pipeline and secures 11 major renewals.

---

How this Roadmap Hits 112% NRR

By the end of Q3 2027, this roadmap expands our $86M base by $7.0M+ in recurring expansion and retained ARR, exceeding the board's 112% target ($6.88M incremental):

InitiativeMechanismTarget ARR Impact
:---:---:---
Multi-Entity52 parent-account subsidiaries onboarding+$2.9M ARR
Virtual Cards (SaaS)Capturing 35% of $410M unmanaged SaaS spend @ 1.1% net interchange+$1.58M ARR
Real-Time Spend ControlsUnlocking enterprise card adoption across 31 QBR accounts+$0.80M ARR
Sage Intacct IntegrationEliminating 6% excess churn across 187 customer accounts+$0.88M ARR
SSO & SCIMEnterprise expansion & churn protection across 11 key accounts+$0.90M ARR
Total Net Expansion+$7.06M ARR (~112.2% NRR)

---

What We Are Not Doing (And Why)

  1. Not building Procurement as a hasty monolith in Q1: Building a broad module against negative design-partner feedback on an imaginary timeline would burn 36 squad-weeks with minimal revenue return. We are building the data model and workflows methodically using Squad A once ramped.
  2. Not assigning the Mobile Rewrite to an unhired squad: Tom assigned this 22-week project to Squad B in Q2. Squad B will not be staffed in time. Instead, the Expenses squad owns mobile: they finish Receipt Matching in Q4 (improving our #1 support issue) and execute the mobile rewrite across Q1 and Q2, completing it comfortably before late-2027 framework retirement.
  3. Not prioritising Microsoft Dynamics over Sage Intacct: Dynamics accounts for only 3% of our customer base and 8% of lost deals (vs. Intacct’s 17% base and 31% deal loss). Dynamics is scheduled for Q2 after Intacct is fully deployed.
  4. No cross-domain squad thrashing: Squads working out of domain suffer a 1.5× productivity penalty. We kept all squads focused within their domain expertise rather than forcing Expenses or Integrations to build Cards or Approvals features.

---

Risks and Mitigation Strategies

1. Card Processor Certification Delay: * Risk: Processor testing slips past 4 weeks, threatening the 31 March v1 shutoff. * Mitigation: We complete Harbor development by Week 7 of Q1 (mid-February), yielding a two-week safety buffer before the freeze. Cards works closely with processor engineers weekly starting in Q4. 2. Multi-Entity Delivery Bottlenecks: * Risk: Multi-Entity touches both Platform (12 wks) and Approvals & Policy (10 wks); cross-squad misalignment could delay Q1 rollout. * Mitigation: The technical spec and API contract between Platform and A&P will be finalized in Sprint 1 of Q4, enabling independent parallel development. 3. Hiring Lags for Squad A and B: * Risk: Talent acquisition takes longer than the historical 5–8 months. * Mitigation: The baseline revenue plan (Multi-Entity, Virtual Cards, Intacct, SSO) relies entirely on our five existing squads. Squad A’s onboarding pace dictates the launch timing of Procurement, not the attainment of our 112% NRR target.

Claims the judge couldn’t find in the brief

  • 14 of our 60 largest accounts requested multi-entity; 9 have signed written commitments to add 52 subsidiaries ($2.9M ARR).

Grades and run details

Decision model 58 · LLM judge 7 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly62%
  • passAddresses the actual decision83%
  • partialRespects explicit constraints23%
  • passIdentifies material uncertainty36%
  • failAvoids unsupported claims33%
  • failProduces the required deliverable79%
  • passProtects the processor deadline32%
  • passPlans on the squads we actually have94%
  • partialMakes the call on procurement89%
  • passOutcomes, with certainty that falls with distance27%
  • failFits the stated capacity52%
  • passSequences around dependencies16%
Run
Run
#1
API response time
45 s
Submitted
1 Oct 2026
Critical failure: Leaves the processor migration finishing too late for 4 weeks of certification before 31 March 2027, or schedules virtual cards or real-time spend controls before it. Score capped at 40.

Needs more than a quick edit0 of 12 checks passed

A PM would need to redo the hiring assumptions, procurement math, capacity sums, and card migration sequencing. The output cannot be used with only wording or emphasis edits.

Output

Harbor FY27 Product Roadmap

Prepared by: Staff PM, Harbor For: Executive Team Review (CEO, CFO, CRO, CTO)

---

Part 1: The FY27 Roadmap (Q4 2026 – Q3 2027)

Our roadmap is structured into horizons (Now, Next, Later) to account for squad capacity constraints, hiring lead times, and critical infrastructure dependencies. Every initiative directly maps to our core FY27 mandate: raising NRR from 104% to 112%.

NOW (Q4 2026)

Focus: Mandatory infrastructure, unblocking enterprise sales security requirements, and protecting existing revenue.

  • Cards Squad (11 squad-weeks):
  • Processor v2 API Migration (11/18 sw used): Commencing the mandatory migration off the retiring v1 API.
  • Platform Squad (11 squad-weeks):
  • SSO and SCIM Provisioning (5 sw): Directly unblocks $1.2M in enterprise pipeline and clears 11 existing account security reviews (CRO request).
  • Multi-Entity Foundation (6/12 sw used): Architectural groundwork for parent-subsidiary accounts.
  • Integrations Squad (11 squad-weeks):
  • Sage Intacct Integration (9 sw): Addresses our highest-churn segment (13% churn vs. 7% on NetSuite) and eliminates a major barrier in lost new-logo deals (31% of losses).
  • Spillover buffer (2 sw).
  • Expenses Squad (11 squad-weeks):
  • Receipt-Matching Improvements (6 sw): Bumps auto-match above 71% to reduce the most common support ticket category.
  • Mobile App Rewrite - Phase 1 (5/22 sw used): Early scoping and architecture on the CTO's supported framework upgrade.
  • Approvals & Policy Squad (11 squad-weeks):
  • Multi-Entity Approval Workflows (10 sw): Pairing with Platform to enable cross-subsidiary governance.
  • Spillover buffer (1 sw).

---

NEXT (Q1 2027)

Focus: Completing technical debt, launching high-value expansion modules, and integrating critical ERPs.

  • Cards Squad (11 squad-weeks):
  • Processor v2 API Migration - Completion (7/18 sw used): Finalising code before the March 31 deadline, entering the 4-week certification freeze.
  • Virtual Cards for Software Subscriptions - Initiation (4/10 sw used).
  • Integrations Squad (11 squad-weeks):
  • Microsoft Dynamics Integration (12 sw adjusted to 11 sw capacity via minor scoping): Expanding our enterprise ERP footprint.
  • Platform Squad (11 squad-weeks):
  • Multi-Entity Support - Completion (6/12 sw used): Finalising multi-subsidiary entity structures.
  • Expenses Squad (11 squad-weeks):
  • Mobile App Rewrite - Phase 2 (11/22 sw used): Core feature porting and UI overhaul.
  • New Squad A (Hired Dec 2026, delivering at ~50% capacity = ~5.5 sw):
  • Procurement Module - Phase 1 (5.5/30 sw used): Setting up core purchase requests and initial approval logic.

(Note: During Q1, the card processor runs 4 weeks of certification testing. No changes can be deployed to card flows during this time).

---

LATER (Q2 – Q3 2027)

Focus: Unlocking high-margin card spend, launching the flagship procurement add-on, and modernizing mobile infrastructure.

  • Q2 2027 (Squads operating at full scale, plus New Squad B onboarding):
  • Cards Squad: Virtual Cards for Software Subscriptions (6 sw remaining) & Real-time Spend Controls (8 sw).
  • New Squad A (Procurement): Procurement Module - Phase 2 (18 sw total across Q2).
  • Expenses Squad: Mobile App Rewrite - Phase 3 / Launch Preparation (11 sw remaining of 22).
  • New Squad B (Hiring initiated Q1, arriving Q2, delivering ~5.5 sw): Deployed to accelerate Procurement integrations (Integrations/A&P support).
  • Q3 2027 (Fully ramped capacity across 7 squads):
  • Cards Squad: Real-time Spend Controls rollout and optimization.
  • New Squad A & B: Procurement Module Launch, GA, and initial adoption drives.
  • Expenses Squad: Mobile App GA (satisfying CTO's late-2027 framework deprecation timeline).
  • Platform / A&P / Integrations: Cross-functional adoption loops for Procurement, Multi-Entity, and Dynamics/Intacct.

---

Part 2: The Case for the Roadmap

1. Why This Order? (Balancing Reality and Ambition)

Our mandate is clear: raise NRR to 112% by the end of Q3 2027. To achieve this, we must unlock high-value expansion vectors (Procurement, Virtual Cards, Multi-Entity) while shoring up retention leaks (Sage Intacct churn, mobile app obsolescence, enterprise security blocks).

However, ambition must bow to physics. Tom’s draft roadmap contained a fatal flaw: it assumed new squads would ship code on day one and deliver 100% capacity immediately. Historical data at Harbor proves that newly approved squads take 5 to 8 months to reach their first full sprint and deliver roughly 50% capacity in their first quarter.

Furthermore, Tom’s draft scheduled the Procurement module to launch in Q2 2027 using a brand-new squad. Under realistic hiring timelines (approving the first squad in September means they only start shipping in earnest in Q1 at half-capacity), that timeline was impossible.

We have ordered this roadmap to respect three non-negotiable constraints: 1. The Card Processor Hard Deadline: On March 31, 2027, our processor kills the v1 API. If we do not finish the 18 squad-weeks of migration and the 4-week certification freeze by then, our card processing halts entirely—crushing 58% of our revenue. 2. Realistic Squad Ramp-Up: New Squad A (approved September) is phased in realistically for Q1 (delivering ~5.5 squad-weeks) and hits full stride in Q2. New Squad B is phased in for Q2. 3. Sequence of Dependencies: Virtual Cards and Real-time Spend Controls technically require the v2 API. Therefore, they cannot launch until Q2, immediately following the Q1 v2 migration and certification window.

2. What We Changed from Tom’s Draft and Why

  • Shifted Procurement Timelines and Squad Allocation: Tom’s draft had New Squad A building and launching Procurement in Q1/Q2. Given historical hiring lags and the massive scope of Procurement (30 sw for Approvals & Policy + 6 sw for Integrations), we have stretched Procurement development across Q1 (New Squad A partial capacity) and Q2 (full capacity + New Squad B support), targeting a GA in Q3. Trying to rush it in Q2 would result in a broken launch or squad burn-out.
  • Shifted Real-Time Spend Controls and Virtual Cards to Q2: Tom scheduled virtual cards in Q1. Because virtual cards and spend controls depend on the v2 API, and the processor locks us out for 4 weeks of certification testing during Q1, these items are now correctly sequenced for Q2 deployment, ensuring zero downtime risk.
  • Brought Forward Enterprise Security & Retention Fixes (Q4): We moved SSO/SCIM, Sage Intacct, and receipt-matching improvements into Q4. Why? Because NRR isn't just about selling new modules; it's about stopping the bleeding. Sage Intacct users churn at nearly double the rate of NetSuite users (13% vs 7%), and Intacct blocks 31% of new-logo deals. Fixing this now preserves our base. Similarly, securing $1.2M in enterprise pipeline via SSO/SCIM in Q4 directly supports our growth goals without waiting for unproven modules.

3. What We Are Not Doing (And Why)

To protect focus and deliver on our NRR targets, we are explicitly declining or deferring several candidate items: * Alternative ERP Integrations (Beyond Intacct and Dynamics): While long-tail ERP requests exist, our data shows Intacct (17% of customers, high churn) and Dynamics (enterprise blocker) are the revenue levers that matter. Other ERPs are deferred to FY28. * Ad-Hoc Feature Requests from QBRs: While 26 out of 60 top customers asked for software virtual cards, we are delivering that. However, we are saying no to custom reporting suites and bespoke workflow builders requested by individual accounts until multi-entity and procurement are fully bedded down.

4. Risks That Could Change the Plan

1. Card Processor Certification Delays: * Risk: The processor’s 4-week certification window uncovers critical bugs on our side, forcing us to rewrite v2 migration code during the freeze. * Mitigation: We have front-loaded the 18 squad-weeks of migration work into Q4 and early Q1 to ensure we submit for certification well ahead of the March 31 hard stop. 2. Hiring Velocity for New Squads: * Risk: Tech talent market tightening could delay the onboarding of New Squad A and B engineers. * Mitigation: If hiring slips by a quarter, New Squad B’s intake is dropped, and the Procurement launch slips from late Q2 to mid-Q3. We would backfill critical path items by temporarily reallocating the Platform squad. 3. Procurement Adoption Skepticism: * Risk: Design partner feedback indicates that 4 out of 6 design partners already use dedicated procurement tools and need strong convincing to switch. Tom’s assumption of $6M ARR in Year 1 may be aggressive if attach rates lag. * Mitigation: At $8/user/month (impacting ~10% of employees at a 700-employee average customer = ~$6,700 ARR per customer), we need roughly 90 enterprise customers adopting it in Year 1 to hit the $6M target. Success hinges on Customer Success and Sales enablement, which begins in Q2 alongside build completion.

Claims the judge couldn’t find in the brief

  • About 90 enterprise customers adopting procurement in Year 1 would hit Tom's $6M ARR claim.
  • New Squad A was hired in December 2026 and can deliver about 5.5 squad-weeks in Q1 2027.
  • New Squad B can be initiated in Q1 2027, arrive in Q2 2027, and deliver about 5.5 squad-weeks.
  • Dynamics is an enterprise blocker and a revenue lever that matters.

Grades and run details

Decision model 25 · LLM judge 0 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly8%
  • partialAddresses the actual decision32%
  • failRespects explicit constraints82%
  • passIdentifies material uncertainty27%
  • partialAvoids unsupported claims27%
  • failProduces the required deliverable71%
  • failProtects the processor deadline57%
  • partialPlans on the squads we actually have56%
  • failMakes the call on procurement76%
  • partialOutcomes, with certainty that falls with distance51%
  • failFits the stated capacity43%
  • failSequences around dependencies82%
Run
Run
#1
Time to output
10 s
Submitted
1 Oct 2026

Usable with a quick edit11 of 12 checks passed

A senior PM could use this with light edits because the evidence, capacity math, deadline protection, and procurement call are all present. The only minor cleanup is tightening some later-quarter allocation language.

Output

Harbor FY27 roadmap: expand existing customers first

For the planning offsite | Q4 2026–Q3 2027

Recommendation: protect card continuity, unlock committed subsidiary expansion, then bring more existing-customer spend onto Harbor. Treat procurement as a commercial hypothesis to validate—not as committed revenue or a committed launch.

Board outcome: raise NRR from 104% to 112% by the end of Q3 2027. This roadmap supports that goal; the available evidence does not yet establish a credible revenue bridge all the way to 112%.

1. Proposed roadmap

Planning rules

  • Commit against the five existing squads, each with 11 squad-weeks per quarter. Neither new squad is on the critical path.
  • Numbers below are squad-weeks. Named feature estimates come from the pack; remaining capacity is an allocation ceiling for rollout, validation, hardening or contingency, not an estimate for additional features.
  • Cross-squad delivery uses the stated 1.5× cost.
  • Adoption work targets expansion and retention in existing accounts. It is not a placeholder for additional module builds.
SquadQ4 2026 — Protect and unblockQ1 2027 — Launch foundationsQ2 2027 — Capture spendQ3 2027 — Scale adoption
CardsProcessor v2: 11. Own architecture, integration and critical migration work. Outcome: eliminate the card-authorisation shutdown risk.Processor v2: remaining 3, plus 8 reserved for defects, freeze, certification and cutover contingency. Target certification 12 Feb–12 Mar, cutover 15 Mar, ahead of the 31 March retirement. No virtual-card launch commitment this quarter.Virtual cards: 10, plus 1 for integration support/contingency. Launch with customers identified through the QBRs. Outcome: move software subscription spend onto Harbor cards.Up to 11: virtual-card activation, spend conversion, controls integration support and reliability. Outcome: sustained incremental card spend before the NRR measurement date.
Approvals & PolicyMulti-entity: 10; procurement commercial validation: up to 1. Outcome: policy support for subsidiary expansion; early test of willingness to pay.Up to 7: multi-entity onboarding and policy fixes; up to 4: procurement validation. Outcome: convert written subsidiary commitments and decide whether procurement merits a funded build.Spend controls: 4; up to 7: multi-entity rollout and controls policy setup. Outcome: make controls usable by finance teams and expand subsidiary adoption.Up to 11: controls rollout, policy tuning and remaining subsidiary onboarding. Outcome: adoption that expands spend rather than creating excessive declines or customer friction.
IntegrationsSage Intacct: 9, plus 2 for pilot rollout and fixes. Outcome: remove a recurring accounting workflow gap in a substantial existing-customer segment.Up to 11: Intacct rollout, reliability and measuring CSV-to-integration conversion. Outcome: reduce retention risk and support expansion.Up to 11: Intacct and multi-entity accounting rollout; unused capacity remains reserve. Outcome: accounting readiness does not block subsidiary activation.Up to 11: integration adoption and reliability against remaining account-level expansion blockers. No Dynamics commitment.
PlatformMulti-entity: 11 of 12. Outcome: build the account structure needed for subsidiary expansion.Multi-entity: remaining 1; SSO/SCIM: 5; up to 5 for rollout and hardening. Launch multi-entity after the remaining platform work and acceptance testing. Outcome: activate subsidiaries and address existing-account security reviews, with enterprise pipeline as a secondary benefit.Up to 11: multi-entity and SSO/SCIM onboarding, operational hardening and contingency. Outcome: reliable expansion across larger customer organisations.Up to 11: scale and security work tied to demonstrated adoption issues. Outcome: retain and expand larger accounts without accumulating operational risk.
ExpensesProcessor support: 6, delivering 4 Cards-equivalent weeks; receipt matching: 5 of 6. Outcome: buy migration schedule margin while progressing the most common expense support issue.Receipt matching: remaining 1; up to 10 for rollout, quality measurement and fixes. Outcome: improve auto-match from the 71% baseline and reduce unmatched-receipt tickets.Spend-controls engineering: 9, delivering 6 of the 8 Cards-equivalent weeks; up to 2 for receipt quality. Cards retains technical ownership. Outcome: develop controls alongside virtual cards rather than queueing both behind one squad.Spend-controls engineering: 3, delivering the final 2 Cards-equivalent weeks, followed by an early-quarter launch subject to acceptance. Up to 8: controls hardening, receipt quality and mobile lifecycle assessment. Outcome: give controls time to drive adoption before quarter-end.

Critical-path checks

  • Processor migration: Q4 delivers 11 Cards weeks + 4 equivalent weeks from Expenses = 15. Q1 delivers the final 3, reaching the required 18. The February certification target creates margin before the hard retirement date.
  • Certification freeze: no changes to the certified implementation during the four-week processor window. Before committing dates, confirm the precise freeze scope, book the processor slot and agree cutover acceptance criteria.
  • Multi-entity: Approvals & Policy completes its 10 weeks in Q4; Platform completes 11 + 1. Therefore this is a Q1 launch, not a Q4 launch.
  • Spend controls: Expenses spends 12 actual weeks across Q2–Q3 to deliver 8 Cards-equivalent weeks; Approvals & Policy supplies its 4 weeks in Q2. Work begins only after v2 is live.
  • New squads: plan no committed output until teams are staffed and their ramp is observed. Use initial capacity for bounded adoption, testing or hardening work; allocate larger scope only after a delivery review.

Outcome scorecard

WorkstreamMeasure that matters
Overall expansionNRR and a cohort-based revenue bridge separating expansion, contraction and churn
Multi-entitySubsidiaries activated, incremental ARR and activation time; start with the nine parents that committed in writing
Virtual cardsIncremental software spend moved onto Harbor and resulting net interchange—not cards issued
Spend controlsEnabled accounts, incremental spend unlocked, decline quality and customer friction
IntacctEligible-account adoption, CSV retirement and subsequent retention/expansion versus comparable accounts
SSO/SCIMResolution of the 11 existing-account security reviews; pipeline conversion tracked separately
Receipt matchingMatch rate, incorrect matches and unmatched-receipt tickets per expense
Procurement discoveryPaid commitments at the proposed pricing, requester-seat counts and a credible reason to switch

CS, Sales and Finance should turn these into named-account activation targets and a revenue bridge at the offsite. We should not invent conversion forecasts from the current pack.

2. The case for the plan

Why this order

First, protect the revenue engine. Every card authorisation uses an API retiring on 31 March, and interchange represents 58% of revenue. Tom’s migration occupies 18 of Cards’ 22 available weeks across Q4 and Q1, but must also leave four calendar weeks for certification. Quarterly capacity alone does not prove that schedule is safe.

Borrowing six Expenses weeks moves four Cards-equivalent weeks into Q4. That leaves only three planned migration weeks in Q1 and creates room for defects and certification. Receipt matching slips slightly; card continuity takes precedence.

Second, prioritise expansion with identifiable buyers. Multi-entity has the strongest evidence of a concrete expansion action: nine parents have committed in writing. The full 52-subsidiary opportunity is $2.9M ARR, but we cannot treat that total as committed or infer how many subsidiaries the nine parents represent. Launch in Q1, then sell and onboard—not merely ship.

Third, pursue spend expansion in parallel. Virtual cards have a direct connection to software spend currently outside Harbor. Controls address the most frequently expressed card need in the QBRs and may make finance teams comfortable moving more spend onto Harbor. Borrowing Expenses capacity costs four additional squad-weeks versus specialist delivery, but avoids serialising both products through Cards and supports an early-Q3 controls launch.

The $410M software-spend estimate implies a $4.51M annual net-interchange ceiling at 1.1%, before accounting for capture rates, timing or spend that cannot move from invoice to card. It is an opportunity, not a forecast. Both features may affect the same spend; we will not double-count it.

Fourth, remove retention and expansion friction. Intacct serves 17% of existing customers, versus 3% for Dynamics. The 13% versus 7% churn difference is an association, not proof the integration causes better retention, but it supports prioritising Intacct. SSO/SCIM is relatively small and addresses 11 existing-account security reviews. Receipt matching targets the largest expense support problem.

What changes from Tom’s draft

  • Hiring no longer funds commitments. Historical time to the first full sprint was five to eight months, with half-capacity delivery in the first quarter. September approvals do not justify a full squad in Q1 and another in Q2. Further, new squads taking specialist work incur the stated transfer cost.
  • Migration gains early borrowed capacity and an explicit certification window. Virtual cards no longer competes with the deadline in Q1.
  • Multi-entity launches in Q1. Platform’s 12-week estimate exceeds one quarter’s 11-week capacity.
  • Procurement moves from committed launch to commercial validation. Its revenue claim is not supported by current evidence.
  • Controls gets borrowed delivery capacity. This creates an earlier adoption window without relying on hiring.
  • Dynamics and the mobile rewrite leave the committed horizon. The freed capacity supports existing-customer outcomes and schedule resilience.

Why procurement is not the anchor

At the proposed price, an average customer has 70 requester seats and contributes:

70 × $8 × 12 = $6,720 ARR.

Reaching $6M ARR requires approximately 893 average-sized customers—81% of today’s customer base. Only two of six design partners said they would pay; neither is described as a signed commitment. Four already have a competing tool.

The estimated build is 30 Approvals & Policy weeks plus six Integrations weeks, before any additional cross-squad cost. That is too large a bet to anchor this year’s NRR target without stronger evidence.

Validation should seek paid commitments, actual seat counts and evidence of a differentiated switching reason. It is not an unbudgeted production build. A successful gate triggers an explicit investment decision, identifying displaced work and a capacity-backed launch date.

What we are not doing

We are not committing to procurement launch or its $6M claim, Dynamics, or a 22-week mobile rewrite in this horizon. We are also not inventing new feature scope to consume every available week.

The mobile support deadline remains real. By Q3, the CTO must establish the exact end-of-support date and a funded maintenance, upgrade or rewrite path. If continued safe operation requires starting earlier, it returns as a mandatory trade-off—not an assumed new-squad assignment.

Risks that could change the plan

The largest execution risks are processor certification availability, migration work that cannot be delegated, and underestimated multi-entity or controls complexity. Validate the processor work split immediately; any threat to certification displaces discretionary work first.

The largest commercial risk is insufficient expansion capture. At an illustrative $86M opening cohort, an eight-point NRR improvement represents $6.88M more retained/expanded ARR than a 104% outcome. Finance must reconcile the actual cohort and baseline. Even the full multi-entity and software-card opportunity totals only $7.41M before timing, conversion and overlap effects. This plan does not justify promising 112%.

Review the revenue bridge after Q1 launches. Weak activation should shift capacity toward removing proven adoption blockers. Strong paid procurement evidence, an earlier mobile security deadline, or demonstrated new-squad capacity can justify a replan—but none should silently become an assumption today.

Grades and run details

Decision model 83 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly35%
  • passAddresses the actual decision99%
  • partialRespects explicit constraints26%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims58%
  • failProduces the required deliverable52%
  • passProtects the processor deadline94%
  • passPlans on the squads we actually have99%
  • partialMakes the call on procurement37%
  • passOutcomes, with certainty that falls with distance92%
  • passFits the stated capacity54%
  • passSequences around dependencies56%
Run
Run
#1
Time to output
2.0 min
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 12

Uses the supplied evidence correctlyWrongWrongRight
Gemini 3.8 Flash · API

Claim 4 states that 9 signed commitments cover all 52 subsidiaries, but the brief only says 9 of 14 committed to add their subsidiaries, without specifying how many subsidiaries those 9 have; the 52 total is across all 14. This invents a fact.

Gemini 3.5 Flash-Lite · Gemini

It miscomputes the procurement adoption math, invents hiring dates and squad arrival assumptions, and adds unsupported Dynamics and ad-hoc request facts.

GPT-6 Astra · ChatGPT

The output’s factual statements and arithmetic are drawn from the supplied context, and it labels uncertain causal and revenue claims appropriately.

Addresses the actual decisionMixedWrongRight
Gemini 3.8 Flash · API

The output does not state what result or condition would change the roadmap decision; it lists risks and mitigations but not how the plan would change if a risk materializes.

Gemini 3.5 Flash-Lite · Gemini

It commits to a roadmap, but the procurement decision is not a clear evidence-based call because it keeps the full build and uses a 10x-wrong adoption threshold.

GPT-6 Astra · ChatGPT

It commits early to protecting card continuity, launching multi-entity, validating procurement rather than building it, and deferring Dynamics and mobile, with clear replan triggers.

Respects explicit constraintsRightWrongRight
Gemini 3.8 Flash · API

The output provides a roadmap by quarter with squad work and outcomes, and a case under 1,200 words covering changes, what's not being done, and risks, as requested.

Gemini 3.5 Flash-Lite · Gemini

It violates the processor sequencing constraint by starting virtual cards in Q1 and relies on unhired squads arriving earlier than the hiring history supports.

GPT-6 Astra · ChatGPT

It provides a quarter-by-quarter roadmap and a case under 1,200 words, respects the 11 squad-week capacity, 1.5× cross-squad cost, and processor certification constraint, and does not rely on unhired squads.

Identifies material uncertaintyMixedMixedRight
Gemini 3.8 Flash · API

It names risks but does not say how they would be resolved or what would change the call; mitigations are given, but no explicit condition that would alter the roadmap.

Gemini 3.5 Flash-Lite · Gemini

It names hiring and certification risks but does not bound the key procurement unknown with a threshold or test that would change the build decision.

GPT-6 Astra · ChatGPT

It names certification, hiring, migration delegation, multi-entity/controls complexity, procurement willingness to pay, and revenue-bridge uncertainty, with resolution or replan conditions.

Avoids unsupported claimsWrongWrongRight
Gemini 3.8 Flash · API

It presents the $2.9M multi-entity expansion as 'contracted expansion ARR' when only 9 of 14 have committed in writing, and the 52 subsidiaries are not all covered by those commitments.

Gemini 3.5 Flash-Lite · Gemini

It presents invented hiring timing, Dynamics blocker status, and custom reporting/workflow requests as current facts.

GPT-6 Astra · ChatGPT

It avoids presenting procurement revenue, integration causation, or the 112% bridge as established facts and labels them as hypotheses or ceilings.

Produces the required deliverableMixedWrongMixed
Gemini 3.8 Flash · API

The roadmap table and case are present, complete, and usable by the exec team with light edits.

Gemini 3.5 Flash-Lite · Gemini

The pack is within length and has roadmap plus case, but it is not usable without redoing capacity, hiring, procurement, and deadline sequencing.

GPT-6 Astra · ChatGPT

The roadmap and case are present, executive-readable, and actionable with only light edits.

Protects the processor deadlineRightWrongRight
Gemini 3.8 Flash · API

Migration build finishes by week 7 of Q1, leaving weeks 8-11 for certification before 31 March; virtual cards and spend controls start only after cut-over in Q2/Q3.

Gemini 3.5 Flash-Lite · Gemini

It schedules virtual cards in Q1 alongside migration and certification, despite the requirement that v2 features start only after cut-over.

GPT-6 Astra · ChatGPT

It finishes the migration with a February-to-March certification window before 31 March, adds buffer and borrowed capacity, and starts virtual cards and spend controls only after cut-over.

Plans on the squads we actually haveRightWrongRight
Gemini 3.8 Flash · API

New squads are planned with half capacity in their first quarter (Q2 for Squad A, Q3 for Squad B), consistent with 5-8 month hiring history, and no critical commitments depend on them earlier.

Gemini 3.5 Flash-Lite · Gemini

It commits procurement to New Squad A in Q1 and New Squad B in Q2 based on unsupported December/February hiring assumptions rather than treating new squads as upside.

GPT-6 Astra · ChatGPT

It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.

Makes the call on procurementWrongWrongRight
Gemini 3.8 Flash · API

It checks the $6M claim and commits Approvals & Policy to better-evidenced work, but the proposed 'Procurement Discovery & Scoping (MVP)' lacks a clear threshold or paid test that would justify the full build.

Gemini 3.5 Flash-Lite · Gemini

It checks Tom's $6M claim but calculates roughly 90 customers instead of about 890, and still commits to the full 30-week procurement build.

GPT-6 Astra · ChatGPT

It checks Tom’s $6M claim against pricing and customer-base arithmetic, weighs the weak design-partner evidence, commits Approvals & Policy to better-evidenced work, and proposes a paid validation gate.

Outcomes, with certainty that falls with distanceRightWrongRight
Gemini 3.8 Flash · API

Every roadmap item names an outcome or problem; near-term items are specific (e.g., 'Processor v2 Migration (11 wks)') and later ones are deliberately looser ('Cross-Entity Reporting & Enterprise Admin').

Gemini 3.5 Flash-Lite · Gemini

Many items name outcomes, but later items remain overly specific and several commitments are not tied to the NRR expansion outcome.

GPT-6 Astra · ChatGPT

Each workstream has an outcome, near-term items are specific, and later-quarter scope is deliberately looser and tied to adoption or contingency.

Fits the stated capacityMixedWrongRight
Gemini 3.8 Flash · API

Committed work per squad per quarter adds up to ≤11 squad-weeks with some slack (e.g., Integrations Q1 uses only 4 weeks), and the case names what was deferred (procurement monolith, Dynamics priority).

Gemini 3.5 Flash-Lite · Gemini

Committed work exceeds stated capacity in places, such as Dynamics at 12 squad-weeks against 11, and it does not show how cross-squad or new-squad work fits.

GPT-6 Astra · ChatGPT

The committed squad-week sums fit the stated 11-week quarterly capacity, use the 1.5× cross-squad cost, and identify deferred work such as Dynamics and mobile.

Sequences around dependenciesRightWrongRight
Gemini 3.8 Flash · API

Virtual cards and spend controls are placed after v2 migration; multi-entity and SSO are sequenced logically; dependencies like 'needs v2' are named.

Gemini 3.5 Flash-Lite · Gemini

It names v2 dependencies but still places virtual cards before cut-over and sequences procurement around assumed new squads rather than proven capacity.

GPT-6 Astra · ChatGPT

It sequences virtual cards and spend controls after v2, multi-entity after Platform and Approvals work, and procurement build after validation, naming the key dependencies.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI89.6100.02None
2GPT-6 AstrawithChatGPT91.787.52None
3GPT-6.1 SolwithAPI91.784.62None
4GPT-6 LunawithAPI85.076.32None
5Opus 5.5withClaude80.565.12None
6Gemini 3.8 FlashwithAPI70.151.92None
7Gemini 3.5 Flash-LitewithGemini28.48.322 capped

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review