Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Outcomes, with certainty that falls with distance95% pass
    Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).
    Sonnet 5.5 · API · Two squads, eight asks, one half
  2. Sequences around dependencies89% pass
    Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.
    Sonnet 5.5 · API · Two squads, eight asks, one half
  3. Plans on the squads we actually have89% pass
    It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.
    GPT-6 Astra · ChatGPT · A year of spend management, with a hard deadline

Where it slips

  1. Makes the call on procurement50% pass
    It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.
    GPT-6 Luna · API · A year of spend management, with a hard deadline
  2. Produces the required deliverable50% pass
    The roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Uses the supplied evidence correctly61% pass
    Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.
    GPT-6 Astra · ChatGPT · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Harbor. Tom Achebe, our CEO, has shared his draft roadmap for the next four quarters (Q4 2026 to Q3 2027), and asked you to turn it into the roadmap we take to next week's planning offsite. The exec team (CEO, CFO, CRO and CTO) will read it beforehand. Write: 1. The roadmap itself, by quarter or by now, next and later: what each squad works on and the outcome each item serves. 2. The case for it, in no more than 1,200 words: why this order, what you changed from Tom's draft and why, what we're not doing, and the risks that could change the plan. The pack below is everything we have. Not all of it matters equally.

What the model was given12 items: About Harbor, FY27 goal (approved by the board, September 2026), Squads and capacity, Hiring, Card processor deadline, Candidate work (estimates in squad-weeks, by owning squad), Tom's draft roadmap, Procurement evidence, Card spend evidence, Multi-entity evidence, Integrations evidence, Other asks
About HarborSpend management (corporate cards, expenses and approvals) for companies with 200 to 2,000 employees. 1,100 customers, $86M ARR, average 700 employees per customer. 58% of revenue is interchange on card spend; the rest is subscription. Net revenue retention (NRR) is 104%.
FY27 goal (approved by the board, September 2026)Raise NRR to 112% by the end of Q3 2027 by expanding inside existing customers: more spend on Harbor cards, more entities and more modules. New-logo growth matters, but it comes second this year.
Squads and capacityFive squads: Cards, Expenses, Approvals & Policy, Integrations and Platform. After support and on-call, each has about 11 squad-weeks of roadmap capacity a quarter. Squads can take work from another squad's area, but it takes them about 1.5× as long.
HiringTwo new squads were approved in September 2026. Tom's draft counts the first from Q1 2027 and the second from Q2 2027, both at full capacity. Last year, our three approved squads took 5, 6 and 8 months from approval to their first full sprint, and each delivered about half its capacity in its first quarter.
Card processor deadlineOur card processor retires its v1 API on 31 March 2027. Every card authorisation we process runs through v1 today. Migrating to v2 is about 18 squad-weeks of Cards work, and the processor must then run 4 weeks of certification testing before we can cut over. Certification is the processor's time, not ours, but nothing can change on our side while it runs.
Candidate work (estimates in squad-weeks, by owning squad)1. Processor v2 migration: Cards 18. Hard deadline above. 2. Virtual cards for software subscriptions: Cards 10. Needs v2. 3. Real-time spend controls (declines out-of-policy spend at the point of sale): Cards 8, Approvals & Policy 4. Needs v2. 4. Procurement module (purchase requests and approvals, a new paid add-on): Approvals & Policy 30, Integrations 6. 5. Multi-entity support (subsidiaries under one parent account): Platform 12, Approvals & Policy 10. 6. Sage Intacct integration: Integrations 9. 7. Microsoft Dynamics integration: Integrations 12. 8. SSO and SCIM provisioning: Platform 5. 9. Receipt-matching improvements: Expenses 6. 10. Mobile app rewrite: Expenses 22.
Tom's draft roadmapQ4 2026: Cards: v2 migration. Approvals & Policy: multi-entity. Integrations: Intacct. Platform: multi-entity. Expenses: receipt matching. Q1 2027: Cards: finish migration, start virtual cards. New squad A: procurement module. Integrations: Dynamics. Platform: SSO and SCIM. Q2 2027: Cards: virtual cards, spend controls. New squad A: procurement launch. New squad B: mobile rewrite. Q3 2027: Cards: spend controls. Everyone else: procurement adoption, mobile launch. Tom's note: “Procurement is our path to 112%. I've told the board it can add $6M of ARR in its first year.”
Procurement evidenceSix design partners have used a prototype since June. Two said they would pay for it; the other four already use a dedicated procurement tool and said they'd need a reason to switch. Proposed price: $8 per user per month, for users who raise purchase requests. At our customers, about 10% of employees raise purchase requests.
Card spend evidenceIn QBRs with our 60 largest customers, 26 asked for virtual cards for software subscriptions, and 31 CFOs asked for real-time spend controls. Finance estimates our customers pay about $410M a year of software subscriptions on other cards or by invoice (extrapolated from the 60 QBR accounts). Our net interchange is 1.1% of card spend.
Multi-entity evidence14 of our 60 largest customers have asked for it. Between them they have 52 subsidiaries not on Harbor, and 9 of the 14 have said in writing they would add their subsidiaries once it exists. Customer Success sized it at $2.9M ARR if all 52 joined at their parents' pricing.
Integrations evidence17% of customers use Sage Intacct through a CSV export. Their gross revenue churn is 13% a year, against 7% for customers on our NetSuite integration. Intacct was cited in 31% of lost new-logo deals last year, Dynamics in 8%. 3% of customers use Dynamics.
Other asksCRO: “Five enterprise deals worth $1.2M of pipeline need SSO and SCIM, and 11 existing accounts have it on their security review list.” CTO: “The mobile app is on a framework version that loses support in late 2027, and I want the rewrite done before then.” Support: receipt auto-match is at 71%, and unmatched receipts are the most common expense ticket.
What a strong answer doesThe answer key the graders mark against

A four-quarter roadmap that protects the 31 March deadline: the migration build has to finish by about early March to leave 4 weeks of certification, which is roughly all of Cards' capacity for Q4 and most of Q1 (18 of about 18 squad-weeks available before certification). So it adds help or buffer (Platform or another squad at 1.5×, or cutting everything else from Cards) and moves virtual cards out of Q1: both v2 features start only after cut-over (virtual cards Q2, spend controls Q2–Q3). It treats the new squads as upside: on past hiring they arrive around February to April 2027 and deliver half capacity in their first quarter, so nothing committed depends on them before Q2–Q3. It makes a clear call on procurement: the $6M claim needs about 890 accounts (81% of customers) paying $6,720 a year each (70 requesters × $8 × 12), and only 2 of 6 design partners would pay, so it replaces the 30-week build with a cheap paid test with a threshold, and keeps Approvals & Policy on multi-entity and spend controls, which have stronger expansion evidence ($2.9M sized, 9 of 14 committed in writing; about $1.5M a year of interchange from moving a third of the $410M of subscriptions onto virtual cards). Intacct comes early for its churn; Dynamics later or cut; SSO and SCIM are small and justified; the mobile rewrite waits for a new squad, before the framework's support ends. It lists what changed from Tom's draft, what's not being done, and the risks (certification slipping, hiring slipping), with commitment falling with distance.

Critical failures (cap the score at 40)
  • Leaves the processor migration finishing too late for 4 weeks of certification before 31 March 2027, or schedules virtual cards or real-time spend controls before it
Case

Staff-level · v1.1 · synthetic · B2B SaaS, fintech, capacity, hard deadline, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit8 of 12 checks passed

A PM would need to add real buffer or help for the v2 migration and define a procurement validation threshold before taking this to the exec team.

Output

Harbor FY27 roadmap — Q4 2026 to Q3 2027

Planning basis: Each existing squad has about 11 roadmap squad-weeks per quarter. Figures below are the candidate-work estimates; capacity not assigned to a feature is deliberately held for launch, customer activation, fixes and contingency. New squads are not counted at full capacity until their ramp is proven.

QuarterSquadWork and outcome
Q4 2026CardsProcessor v2 migration — 11 of 18 weeks. Start the mandatory migration and establish the path to retire v1 before the processor deadline.
Approvals & PolicyMulti-entity — 10 of 10 weeks. Build parent/subsidiary approval and policy support so existing customers can bring more entities onto Harbor.
PlatformMulti-entity — 11 of 12 weeks. Build the platform foundation for subsidiaries under a parent account.
IntegrationsSage Intacct — 9 weeks. Replace CSV workarounds to reduce a retention risk among existing customers.
ExpensesReceipt matching — 6 weeks. Improve the most common expense-ticket driver; use remaining capacity for fixes and customer feedback.
Q1 2027CardsFinish v2 — 7 weeks. Target implementation complete by the end of February, leaving March for the processor’s four-week certification. No card-side changes during certification.
Approvals & PolicyMulti-entity launch and customer activation. Support the first subsidiary rollouts. Use remaining capacity for a time-bounded procurement willingness-to-pay and switching validation; do not build the module.
PlatformFinish multi-entity — 1 week; SSO/SCIM — 5 weeks. Complete the entity foundation, then address security blockers affecting existing accounts and enterprise pipeline.
IntegrationsMicrosoft Dynamics — 11 of 12 weeks. Address a cited reason for lost new-logo deals; complete the final week in Q2.
ExpensesMobile rewrite — 11 of 22 weeks. Start early to reduce the risk of running out of time before the current framework loses support in late 2027.
Q2 2027CardsVirtual cards for software subscriptions — 10 weeks. After v2 certification, give customers a way to move subscription spend onto Harbor cards and capture more interchange.
Approvals & PolicyReal-time spend controls — 4 weeks. Build the policy logic for point-of-sale declines, in step with Cards’ Q3 work.
PlatformSSO/SCIM and multi-entity activation. Support adoption and resolve launch issues; no new major platform feature is precommitted.
IntegrationsFinish Dynamics — 1 week; customer rollout. Complete the integration and support implementation.
ExpensesFinish mobile rewrite — 11 weeks. Complete the rewrite by Q2, ahead of the framework support deadline.
Q3 2027CardsReal-time spend controls — 8 weeks. Complete and release with Approvals & Policy; help finance teams enforce policy at the point of sale.
Approvals & PolicyRelease and adoption for real-time controls; multi-entity activation. Focus on turning shipped capabilities into customer usage and expansion.
PlatformActivation and reliability for multi-entity and SSO/SCIM. Prioritize customer rollout and fixes over an unvalidated new feature.
IntegrationsIntacct and Dynamics adoption. Help customers get live and measure retention and deal impact before choosing further integration work.
ExpensesMobile rollout and stabilization; measure receipt-match performance. Address launch issues and confirm the support-ticket outcome.

New Squads A and B: Plan for onboarding and ramp, not full feature capacity. When productive, use them for bounded rollout, quality or activation work. No committed launch depends on their availability.

Case for the plan

Why this order

The board’s priority is to move NRR from 104% to 112% by the end of Q3, primarily through expansion in existing customers. The plan therefore favors credible opportunities to retain customers or deepen their use of Harbor, while keeping some room for new-logo needs.

First, v2 is a hard dependency and a hard deadline. All authorisations still run through v1, which retires on 31 March. The 18 Cards squad-weeks consume 11 weeks in Q4 and seven in Q1. That leaves time to finish implementation before the processor’s four-week certification and to keep the required freeze during testing. Virtual cards and real-time controls must wait until certification is complete.

Next, multi-entity has concrete expansion evidence: 14 of the 60 largest customers asked, and nine have said in writing that they would add subsidiaries. The $2.9M ARR estimate assumes all 52 subsidiaries join, so it is an upside case, not a commitment. Completing the Platform and Approvals & Policy work across Q4 and Q1 creates a path to test that demand with real rollouts.

Virtual cards are the first post-migration Cards feature because 26 large customers asked for them and Finance estimates $410M of annual subscription spend is currently elsewhere. At Harbor’s 1.1% net interchange, full capture would represent about $4.5M in annual interchange revenue, but that is a ceiling, not a forecast: the spend estimate is extrapolated, and adoption and capture are unknown. Real-time controls follow, with coordinated work across Cards and Approvals & Policy. The 31 CFO requests are strong demand evidence, but the plan avoids trying to fit both features into Cards’ 11-week quarterly capacity.

Intacct comes before Dynamics. Intacct is used by 17% of customers through CSV, and those customers have higher observed gross revenue churn than customers on NetSuite. That comparison is correlational, but it supports addressing a retention risk. Dynamics has lower existing use (3%) and a more limited new-logo case, but the 8% of lost deals citing it justifies completing the 12-week integration after Intacct.

SSO/SCIM is a relatively small, five-week Platform investment tied to five enterprise deals representing $1.2M of pipeline and 11 existing security reviews. It supports commercial outcomes without displacing the larger expansion bets.

What changed from Tom’s draft

  • Removed procurement from committed build and launch. Six design partners are not enough evidence for the board’s $6M first-year ARR claim: only two said they would pay, and four already use another tool. At the proposed price, an average 700-employee customer would generate about $6,720 ARR if 10% of employees were buyers. Reaching $6M would require roughly 893 such customer equivalents. We should validate willingness to switch and pricing before committing 30 Approvals & Policy and six Integrations squad-weeks.
  • Corrected the hiring assumption. Prior squads took five to eight months from approval to a first full sprint and delivered about half capacity in their first quarter. The plan does not count either new squad at full capacity in Q1 or Q2, or make a feature dependent on them.
  • Made the v2 critical path explicit. The migration spans Q4 and Q1, with implementation targeted ahead of March certification. Tom’s draft did not protect that freeze clearly enough.
  • Sequenced Cards work to fit capacity. Virtual cards and real-time controls require 18 Cards squad-weeks together, more than one quarter’s 11-week capacity. They are sequenced across Q2 and Q3.
  • Started the mobile rewrite with the existing Expenses squad. The work is 22 weeks, so starting in Q1 targets completion in Q2 without relying on an unproven new squad.

What we are not doing

We are not committing to build procurement, starting a mobile rewrite with a new squad at assumed full capacity, or taking on another major integration or platform feature in Q3 without evidence from the launches above. New-logo work remains in the plan, but expansion and retention come first.

Risks that could change the plan

  • Processor timing: Any migration slip puts card authorisations at risk. Track delivery and certification weekly; protect the Q1 card schedule and do not pull dependent features forward.
  • Expansion may not materialize: Multi-entity interest and card requests do not guarantee adoption or spend migration. Track subsidiary activation, card spend moved, and controls usage—not just releases.
  • Commercial estimates are uncertain: The $2.9M multi-entity estimate and $4.5M card-spend ceiling are not forecasts. Procurement’s $6M claim is especially unsupported by current evidence.
  • Mobile has little slack: The 22-week rewrite fills two quarters of Expenses capacity. Scope growth, delivery delays or an earlier-than-expected framework deadline would require replanning.
  • Hiring, integration and cross-squad delivery may slip: Ramp timing is uncertain; the Intacct churn difference may not be caused by the integration; and spend controls depend on coordinated Cards and Approvals & Policy work.

This roadmap gives Harbor credible bets aligned to the NRR goal, but it does not establish that the 112% target is guaranteed. We should review expansion and retention indicators at each quarterly planning checkpoint and change the plan if customer evidence or delivery risk materially shifts.

Grades and run details

Decision model 79 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly56%
  • passAddresses the actual decision96%
  • partialRespects explicit constraints16%
  • passIdentifies material uncertainty89%
  • passAvoids unsupported claims64%
  • failProduces the required deliverable9%
  • partialProtects the processor deadline35%
  • passPlans on the squads we actually have98%
  • partialMakes the call on procurement84%
  • passOutcomes, with certainty that falls with distance72%
  • passFits the stated capacity62%
  • passSequences around dependencies54%
Run
Run
#1
API response time
1.9 min
Submitted
1 Oct 2026

Needs more than a quick edit5 of 12 checks passed

The output contains a factual error about multi-entity commitments (claiming 9 commitments cover all 52 subsidiaries) and lacks a clear procurement test threshold, so a PM would need to re-check evidence and redo that reasoning.

Output

Part 1: Product Roadmap (Q4 2026 – Q3 2027)

``` ==================================================================================================== SQUAD Q4 2026 Q1 2027 Q2 2027 Q3 2027 ==================================================================================================== Cards Processor v2 Migration Finish v2 (Wks 1-7) Virtual Cards for SaaS Real-Time Controls (11 wks) API Freeze/Cert (W8-11) (6 wks remaining) (Cards: 3 wks) Virtual Cards Prep Real-Time Controls Interchange Optimization (4 wks) (Cards: 5 wks) & Scaling (8 wks) ---------------------------------------------------------------------------------------------------- Platform Multi-Entity: Core Multi-Entity: Wrap-up Platform Scaling & Cross-Entity Reporting Data Architecture (1 wk) Org Hierarchies & Enterprise Admin (11 wks) SSO & SCIM (5 wks) (11 wks) (11 wks) Hardening (5 wks) ---------------------------------------------------------------------------------------------------- Approvals & Multi-Entity: Policy Multi-Entity: Release Real-Time Controls: Procurement Discovery Policy & Approval Routing (2 wks) Policy Engine (4 wks) & Scoping (MVP) (10 wks) A&P Rule Engine Complex Approval Rules (Approvals: 11 wks) Buffer (1 wk) Upgrades (9 wks) (7 wks) ---------------------------------------------------------------------------------------------------- Integrations Sage Intacct Native Sage Intacct GA (2 wks) MS Dynamics Integration MS Dynamics GA (3 wks) Integration MS Dynamics Integration (Integrations: 10 wks) ERP Webhooks & Custom (9 wks) + Buffer (2 wk) (Integrations: 2 wks) Sync (8 wks) ---------------------------------------------------------------------------------------------------- Expenses Receipt Auto-Match Mobile App Rewrite: Mobile App Rewrite: Mobile App GA & Optimization (6 wks) Phase 1 Core Workflows Phase 2 Parity/Release Expense Policy Polish Mobile Prep (5 wks) (11 wks) (11 wks) (11 wks) ---------------------------------------------------------------------------------------------------- New Squad A [Hiring / Onboarding] [Ramp / Training] Procurement Foundation Procurement Workflows (A&P Focus) (Ramped: ~5.5 wks) (Full capacity: 11 wks) ---------------------------------------------------------------------------------------------------- New Squad B [Hiring / Pipeline] [Hiring] [Onboarding / Ramp] Integrations / Core (Ecosystem Focus) (Ramped: ~5.5 wks) ==================================================================================================== ```

Squad-by-Squad Outcomes

  • Cards:
  • Processor v2 Migration: Protects 100% of Harbor’s card-authorisation volume ($49.9M existing ARR) against the hard 31 March 2027 processor shutoff; unlocks v2 capabilities required for virtual cards and real-time controls.
  • Virtual Cards for Software Subscriptions: Captures existing customer SaaS spend currently paid via invoice/other cards ($410M addressable pool), directly driving interchange revenue to expand NRR.
  • Real-Time Spend Controls: Enables 31 top-tier CFO accounts to enforce budgets at point-of-sale, unlocking enterprise card issuance and shifting unmanaged invoice spend to Harbor cards.
  • Platform:
  • Multi-Entity Architecture: Unlocks $2.9M contracted expansion ARR across 52 non-Harbor subsidiaries of top accounts.
  • SSO & SCIM Provisioning: Unblocks $1.2M in enterprise pipeline and closes compliance gaps for 11 critical renewal accounts.
  • Approvals & Policy:
  • Multi-Entity Approvals: Delivers cross-subsidiary approval chains required to close existing multi-entity commitments.
  • Real-Time Policy Engine: Co-delivers immediate POS budget validation with the Cards squad.
  • Procurement Discovery & MVP Architecture: Reframes procurement around verified buyer demand without starving committed delivery.
  • Integrations:
  • Sage Intacct Integration: Halves the 13% annual churn rate of the 17% of Harbor customers on Intacct (retaining ~$875k ARR annually) and addresses Harbor's #1 new-logo loss reason (31% of lost deals).
  • Microsoft Dynamics Integration: Expands ERP coverage for mid-market accounts, removing friction for 3% of existing customers and 8% of pipeline.
  • Expenses:
  • Receipt-Matching Improvements: Drives match rate above 71%, eliminating Harbor’s highest-volume support ticket category and improving user retention.
  • Mobile App Rewrite: Eliminates framework deprecation risk ahead of the late-2027 vendor cutoff without stealing capacity from revenue-generating squads.

---

Part 2: The Case for the Roadmap

Executive Summary

To achieve our board-approved target of 112% NRR by Q3 2027 ($86M ARR base expanding by ~$6.9M), Harbor must focus on the highest-probability, shortest-time-to-value expansion levers within our current 1,100 customers: multi-entity expansion ($2.9M ARR), capturing existing software spend via virtual cards ($1.3M–$2.2M ARR), and plugging Sage Intacct churn ($875k ARR retained).

Tom’s draft relied on an unbuilt Procurement module to generate $6M ARR in FY27. Ground-level capacity, vendor deadlines, and design partner data show that strategy carries severe revenue and operational risk. The proposed roadmap secures our existential card processor deadline, resolves major customer expansion requests, and establishes realistic hiring and product timelines.

---

What Changed from Tom’s Draft and Why

#### 1. Realignment on Procurement ($6M ARR Myth vs. Reality) * The Math: Harbor has 1,100 customers averaging 700 employees (770,000 total users). With 10% raising purchase requests (70 users/account) at $8/user/month ($96/year), an adopting customer generates $6,720 ARR. To generate $6M ARR, 893 out of 1,100 customers (81% of our entire base) would have to buy and deploy a v1 module in under 9 months. * The Customer Signal: 4 out of 6 design partners already use dedicated tools (e.g., Coupa, Zip) and explicitly stated they will not switch. Only 2 of 6 expressed willingness to pay. * The Decision: Procurement cannot carry our FY27 NRR goal. We have reallocated primary capacity to guaranteed expansion drivers (Multi-Entity and Virtual Cards), while initiating a structured discovery and phased MVP build with Squad A.

#### 2. Realistic Hiring Capacity * Tom’s Assumption: Squad A operates at 100% capacity in Q1 2027 (11 squad-weeks) and Squad B in Q2 2027 (11 squad-weeks). * Historical Reality: Approvals to first full sprint took 5–8 months historically, delivering ~50% capacity in quarter one. Squads approved in September 2026 will not reach operational productivity until late Q1 / early Q2 2027. * The Decision: No critical-path milestone or revenue target in Q4 or Q1 relies on unhired squads. Existing squads carry core commitments; Squad A ramps on Procurement foundations in Q2 (delivering ~5.5 squad-weeks) and reaches full delivery in Q3.

#### 3. Strict Compliance with the 31 March 2027 Processor Deadline * The Constraint: The v1 API deprecates on 31 March 2027. Migrating takes 18 squad-weeks plus a mandatory 4-week processor certification freeze during which our codebase cannot change. * Execution Window: The 4-week certification must run in March 2027. Cards code complete must occur by Week 7 of Q1 2027. * The Allocation: Cards dedicates all 11 squad-weeks in Q4 2026 and 7 squad-weeks in Q1 2027 (18 weeks total). During the 4-week March freeze, the Cards squad transitions to architectural prep for Virtual Cards, avoiding wasted idle time.

#### 4. Resequencing Multi-Entity, SSO/SCIM, and Sage Intacct * Multi-Entity (Platform + Approvals): Prioritised immediately in Q4. 14 of our 60 largest accounts requested this; 9 have signed written commitments to add 52 subsidiaries ($2.9M ARR). Shipping in early Q1 allows two full quarters of rollout and revenue recognition. * Sage Intacct (Integrations): 17% of our customers use Intacct via CSV, churning at 13% vs. NetSuite’s 7%. Cutting churn across these 187 accounts ($14.6M ARR base) saves ~$875k in ARR retention while eliminating our top sales loss reason (31% of lost deals). * SSO & SCIM (Platform): Sequenced for Q1 immediately following core Multi-Entity work. For a 5-week lift, this unlocks $1.2M in stalled enterprise pipeline and secures 11 major renewals.

---

How this Roadmap Hits 112% NRR

By the end of Q3 2027, this roadmap expands our $86M base by $7.0M+ in recurring expansion and retained ARR, exceeding the board's 112% target ($6.88M incremental):

InitiativeMechanismTarget ARR Impact
:---:---:---
Multi-Entity52 parent-account subsidiaries onboarding+$2.9M ARR
Virtual Cards (SaaS)Capturing 35% of $410M unmanaged SaaS spend @ 1.1% net interchange+$1.58M ARR
Real-Time Spend ControlsUnlocking enterprise card adoption across 31 QBR accounts+$0.80M ARR
Sage Intacct IntegrationEliminating 6% excess churn across 187 customer accounts+$0.88M ARR
SSO & SCIMEnterprise expansion & churn protection across 11 key accounts+$0.90M ARR
Total Net Expansion+$7.06M ARR (~112.2% NRR)

---

What We Are Not Doing (And Why)

  1. Not building Procurement as a hasty monolith in Q1: Building a broad module against negative design-partner feedback on an imaginary timeline would burn 36 squad-weeks with minimal revenue return. We are building the data model and workflows methodically using Squad A once ramped.
  2. Not assigning the Mobile Rewrite to an unhired squad: Tom assigned this 22-week project to Squad B in Q2. Squad B will not be staffed in time. Instead, the Expenses squad owns mobile: they finish Receipt Matching in Q4 (improving our #1 support issue) and execute the mobile rewrite across Q1 and Q2, completing it comfortably before late-2027 framework retirement.
  3. Not prioritising Microsoft Dynamics over Sage Intacct: Dynamics accounts for only 3% of our customer base and 8% of lost deals (vs. Intacct’s 17% base and 31% deal loss). Dynamics is scheduled for Q2 after Intacct is fully deployed.
  4. No cross-domain squad thrashing: Squads working out of domain suffer a 1.5× productivity penalty. We kept all squads focused within their domain expertise rather than forcing Expenses or Integrations to build Cards or Approvals features.

---

Risks and Mitigation Strategies

1. Card Processor Certification Delay: * Risk: Processor testing slips past 4 weeks, threatening the 31 March v1 shutoff. * Mitigation: We complete Harbor development by Week 7 of Q1 (mid-February), yielding a two-week safety buffer before the freeze. Cards works closely with processor engineers weekly starting in Q4. 2. Multi-Entity Delivery Bottlenecks: * Risk: Multi-Entity touches both Platform (12 wks) and Approvals & Policy (10 wks); cross-squad misalignment could delay Q1 rollout. * Mitigation: The technical spec and API contract between Platform and A&P will be finalized in Sprint 1 of Q4, enabling independent parallel development. 3. Hiring Lags for Squad A and B: * Risk: Talent acquisition takes longer than the historical 5–8 months. * Mitigation: The baseline revenue plan (Multi-Entity, Virtual Cards, Intacct, SSO) relies entirely on our five existing squads. Squad A’s onboarding pace dictates the launch timing of Procurement, not the attainment of our 112% NRR target.

Claims the judge couldn’t find in the brief

  • 14 of our 60 largest accounts requested multi-entity; 9 have signed written commitments to add 52 subsidiaries ($2.9M ARR).

Grades and run details

Decision model 58 · LLM judge 7 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly62%
  • passAddresses the actual decision83%
  • partialRespects explicit constraints23%
  • passIdentifies material uncertainty36%
  • failAvoids unsupported claims33%
  • failProduces the required deliverable79%
  • passProtects the processor deadline32%
  • passPlans on the squads we actually have94%
  • partialMakes the call on procurement89%
  • passOutcomes, with certainty that falls with distance27%
  • failFits the stated capacity52%
  • passSequences around dependencies16%
Run
Run
#1
API response time
45 s
Submitted
1 Oct 2026

Needs more than a quick edit5 of 12 checks passed

A PM would need to shorten the case, correct the false draft claim, and rework procurement and virtual-card dependencies before using it. The core sequencing is strong but not ready without substantive edits.

Output

Harbor Roadmap, Q4 2026 to Q3 2027

Pre-read for the planning offsite. Prepared from Tom's draft.

Part 1: The roadmap

Planning assumptions - Each existing squad has 11 squad-weeks of roadmap capacity a quarter. Work done in another squad's area costs 1.5×. - The two new squads are planned at the dates our history supports, not the approved dates. Last year's squads took 5 to 8 months from approval to their first full sprint. - Squad A: 0 weeks in Q1, about 5 in Q2 (half capacity), about 11 in Q3. - Squad B: 0 weeks in Q2, about 5 in Q3. - Nothing committed depends on either new squad. - The fiscal year runs Q4 2026 (Oct to Dec) through Q3 2027 (Jul to Sep).

Squad plan by quarter (squad-weeks in brackets)

SquadQ4 2026Q1 2027Q2 2027Q3 2027
Cardsv2 migration (11)v2 migration: code complete by ~31 Jan (4). Certification runs through Feb with no card-stack changes. Cutover ~1 Mar. Cutover contingency (3). Virtual cards design and sandbox build, no production changes (4)Virtual cards build and launch ~May (6). Real-time spend controls (5)Spend controls launch ~Jul (3). Adoption work on virtual cards and controls with CS (8)
Platformv2 migration support (5, which buys ~3.3 Cards-weeks). Multi-entity (6)Multi-entity, ships ~end Feb (6). SSO and SCIM, ships by end Mar (5)Subsidiary onboarding tooling and SSO rollout to the 11 accounts (5). Procurement billing and permissions, if the gate passes (6)New-squad onboarding and spend-controls latency support (5). Reserve (6)
Approvals & PolicyMulti-entity (10). Procurement pricing test with design partners (1)Procurement build, phase 1 (11). Gate at end of MarchSpend controls policy engine (4). Procurement (7)Procurement finish (8). Beta, then GA ~Sep (3)
IntegrationsSage Intacct, ships ~early Dec (9). CSV-to-integration migration tooling (2)Move ~187 Intacct CSV customers onto the integration, with CS (5). Dynamics (6)Dynamics, ships ~May (6). Procurement integrations (5)Procurement integrations (1). Hardening and reserve (10)
ExpensesReceipt matching (6). Mobile rewrite (5)Mobile rewrite (11)Mobile rewrite, feature-complete ~May (6). Beta and fixes (5)Staged rollout. Old app retired by end Sep, a quarter or more before framework support ends
New squad A (upside)——Onboarding (~5)~11: accelerate procurement, then own it. Otherwise, adoption work on cards
New squad B (upside)———Onboarding (~5). Scope set at Q2 planning

Items, outcomes and key dates

ItemOutcome it servesKey dateOwner
Processor v2 migrationKeeps every card authorisation running. That protects the 58% of revenue that is interchangeCode complete ~31 Jan, cutover ~1 Mar, about 4 weeks of buffer before 31 MarCards + Platform
Multi-entityMore entities: 9 customers have committed in writing to add subsidiariesShips ~end Feb. Subsidiaries onboard in Q2 and Q3Platform + A&P
Virtual cardsMore card spend: brings software subscriptions onto Harbor cardsLaunches ~MayCards
Real-time spend controlsMore card spend: gives CFOs the confidence to issue more cards (31 CFOs asked for it)Launches ~JulCards + A&P
Intacct integration and CSV migrationRetention: Intacct-on-CSV customers churn at 13%, against 7% on NetSuite. Also addresses 31% of lost dealsIntegration ships Dec. Customers move in Q1Integrations
SSO and SCIMRetention of the 11 accounts with it on their security review list. Also $1.2M of new-logo pipelineShips MarPlatform
ProcurementMore modules: a new paid add-onGate end Mar. GA ~SepA&P
DynamicsNew-logo win rate (8% of lost deals)Ships ~MayIntegrations
Receipt matchingLower support load: the most common expense ticketShips DecExpenses
Mobile rewriteRemoves platform risk before the framework loses supportOld app retired by SepExpenses

NRR bridge: what this plan plausibly delivers by end of Q3 2027

Moving NRR from 104% to 112% needs about $6.9M of additional net expansion (8 points of $86M).

LeverBasisEstimate in window
Multi-entity9 of the 14 requesting customers committed in writing, roughly 33 of 52 subsidiaries. The $2.9M figure assumes all 52 join$1.5M to $2.9M
Virtual cards$410M × 1.1% = $4.5M ceiling. Assumes 15 to 35% captured within ~5 months of launch$0.7M to $1.6M
Intacct churn reduction~$14.6M of ARR on CSV. Closing the churn gap is worth ~$0.9M a year, and only part of that lands in window$0.3M to $0.6M
ProcurementGA in ~SepUnder $0.3M
Spend controls, SSO and SCIMNot sizedUpside or protection
Total$2.8M to $5.4M, or roughly 107% to 110% NRR

---

Part 2: The case for it

The short version

Tom's draft has the right ingredients. It does, however, have three problems:

  1. No buffer on the processor deadline. The migration has no room to slip against a date that could stop every card authorisation.
  2. Key work sits on squads we won't have. Procurement and the mobile rewrite depend on new squads that history says won't exist at full capacity until Q3 at the earliest.
  3. Procurement can't carry the goal. It is expected to deliver the NRR target, but the evidence doesn't support $6M and the timing puts almost none of it inside the measurement window.

This plan protects the deadline first. It then sequences the levers with the strongest evidence so they land early enough to count. Even so, product alone likely gets us to 107% to 110%, not 112%. The exec team should know that now, not in Q3.

Why this order

1. The processor migration comes first, with buffer.

Every authorisation runs through v1, so a missed cutover puts most of our revenue at risk. On Tom's draft: - Cards alone finishes the 18 weeks around late February. - The 4-week certification then ends in the last days of March. - That leaves under a week of slack, with December holidays not counted and no room for a failed certification.

Lending 5 Platform weeks in Q4 changes this: - Code complete moves to about 31 January. - Cutover lands around 1 March. - We have roughly four weeks of buffer.

This also means virtual cards can't "start in Q1" as the draft says. During certification nothing on our card stack can change. Cards will do design and sandbox work in Q1 and start production work after cutover.

2. Multi-entity and virtual cards next, because they have the best evidence and they land in time.

NRR is measured at the end of Q3, so anything launching after about June barely counts. Our two strongest levers are: - Multi-entity: 9 customers have committed in writing to add subsidiaries, which is about $1.5M to $2.9M. - Virtual cards: a $4.5M ceiling. 26 of our 60 largest customers asked for them.

Multi-entity ships in February. Virtual cards ship in May. Real-time spend controls follow in July because they share the Cards squad. They matter to 31 CFOs, but we haven't sized them.

3. Retention work runs in parallel, because it is cheap.

  • Intacct: 9 weeks of work. It addresses a segment churning at nearly twice our NetSuite rate and appears in 31% of lost deals. The integration only reduces churn if customers actually move off CSV, so I added a Q1 migration push with CS.
  • SSO and SCIM: 5 weeks of work. It protects 11 accounts and unblocks $1.2M of pipeline.

4. Procurement is real but gated, and it is not the FY27 lever.

What I changed from Tom's draft, and why

  1. Migration buffer. Platform lends 5 weeks in Q4. The cost is that multi-entity moves from December to February. I think that trade is clearly right, because the downside of a missed cutover is existential.

2. New squads planned realistically. - Squads approved in September reach their first full sprint between February and May, based on last year's 5 to 8 months. - Each then runs at half capacity for a quarter. - Squad A realistically gives about 5 weeks in Q2. Squad B gives nothing before Q3. - I've kept every committed item on the existing five squads. The new squads are upside.

3. Procurement moved to Approvals & Policy and gated. - A&P owns the domain, has capacity after multi-entity, and was left without work in the draft. - Build starts in Q1. At the end of March we need signed paid-pilot commitments from about 10 customers beyond the two design partners, or we redirect A&P. - GA lands around September. If Squad A arrives on time, it accelerates the work.

4. The $6M procurement figure needs resetting with the board. - At $8 per user per month and 10% of our ~770,000 employees raising requests, the ceiling is about $7.4M a year if every customer adopts. $6M therefore implies about 81% adoption. - Only 2 of 6 design partners said they'd pay. The other four already use a dedicated tool. - At a one-in-three adoption rate, procurement is worth about $2.5M at full run-rate, reached over years, and less than $0.3M inside FY27. - I recommend Tom and the CFO correct this at the next board touchpoint, before the board starts tracking it.

5. Mobile rewrite moved to Expenses, starting in Q4. - Expenses owns the app and the draft left it idle after Q4. - The rewrite finishes by May, and the old app is retired by September. - This gives the CTO's deadline real margin. Leaving it on Squad B would risk missing it.

  1. Spend controls now include the A&P work (4 weeks). The draft omitted this.
  1. Dynamics moves from Q1 to May. The Intacct customer migration goes first, because Intacct matters roughly five times more by customer count and lost deals.

What we're not doing

  • Not claiming procurement drives FY27 NRR. We are also not building past the March gate without paid commitments.
  • Not starting virtual cards or anything else on the card stack before cutover.
  • Not putting any deadline-critical or NRR-critical work on unhired squads.
  • Not adding new-logo-only features beyond Intacct, Dynamics and SSO, which also serve retention. New logos come second this year, as the board set.
  • Not filling the Q3 reserve on Platform and Integrations. That slack absorbs migration fallout, integration hardening and new-squad onboarding. If the migration goes cleanly, we can release it to card-spend adoption work at Q2 planning.

Risks that could change the plan

1. Certification fails or needs a second round. A re-run would consume the entire buffer. - Book the certification slot now. - Run a pre-certification dry run in January. - Ask the processor in writing for a contingency extension and an emergency-fix protocol for the freeze window. - If code complete slips past mid-February, pull more Platform capacity immediately.

2. The $410M software-spend figure is extrapolated from our 60 largest accounts. Smaller customers may put proportionally less software spend on cards. - In Q4, validate it with AP data from about 20 mid-sized customers. - If it halves, virtual cards is worth about $0.4M to $0.8M in window and spend controls rises in priority.

  1. Multi-entity conversion comes in below the written commitments. CS should agree subsidiary onboarding plans with the 9 customers in Q1, before the feature ships.
  1. Real-time spend controls sit on the authorisation path. A latency problem could cause false declines. Platform supports this work in Q3, and we should launch in stages.
  1. Hiring slips further. This doesn't threaten committed dates, but procurement would stay in Q3 and our adoption capacity shrinks.
  1. The procurement gate fails. A&P redirects to deeper spend controls and multi-entity policy features, both of which feed card spend.
  1. The mobile framework end-of-support date is only "late 2027". The CTO should confirm the exact date.

Decisions we need at the offsite

  1. Approve Platform's Q4 loan to Cards, accepting that multi-entity moves to February.
  2. Agree the board message on procurement, and the gate criteria behind it.
  3. Close the NRR gap. This plan lands about 107% to 110%. Closing the remaining $1.5M to $4M needs commercial levers the CRO and CFO own, such as CS-led card activation, packaging and pricing, and a multi-entity sales motion beyond the 9. Otherwise we tell the board now that 112% is a stretch.
  4. Authorise the processor extension request.

Claims the judge couldn’t find in the brief

  • Tom's draft left Approvals & Policy without work.

Grades and run details

Decision model 79 · LLM judge 5 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly16%
  • passAddresses the actual decision98%
  • partialRespects explicit constraints33%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims20%
  • failProduces the required deliverable66%
  • passProtects the processor deadline98%
  • passPlans on the squads we actually have96%
  • passMakes the call on procurement59%
  • passOutcomes, with certainty that falls with distance88%
  • passFits the stated capacity39%
  • passSequences around dependencies62%
Run
Run
#1
Time to output
3.5 min
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 10

Uses the supplied evidence correctlyRightWrongWrong
GPT-6 Luna · API

The output's factual statements about Harbor, evidence, estimates, and arithmetic are supported by the supplied context.

Gemini 3.8 Flash · API

Claim 4 states that 9 signed commitments cover all 52 subsidiaries, but the brief only says 9 of 14 committed to add their subsidiaries, without specifying how many subsidiaries those 9 have; the 52 total is across all 14. This invents a fact.

Opus 5.5 · Claude

Most quantitative evidence is used correctly, but the output makes at least one false current-situation claim about Tom's draft leaving Approvals & Policy without work.

Addresses the actual decisionRightMixedRight
GPT-6 Luna · API

It commits to a clear four-quarter roadmap and procurement call, and says quarterly evidence checkpoints and risks could change the plan.

Gemini 3.8 Flash · API

The output does not state what result or condition would change the roadmap decision; it lists risks and mitigations but not how the plan would change if a risk materializes.

Opus 5.5 · Claude

It commits to a clear four-quarter order and decision set for the exec team, including the migration buffer, procurement gate, and NRR gap.

Respects explicit constraintsWrongRightWrong
GPT-6 Luna · API

It does not enforce the processor-deadline constraint with real buffer or help, and the procurement validation lacks a threshold that would justify the full build.

Gemini 3.8 Flash · API

The output provides a roadmap by quarter with squad work and outcomes, and a case under 1,200 words covering changes, what's not being done, and risks, as requested.

Opus 5.5 · Claude

It exceeds the 1,200-word limit for the case and proposes virtual-card sandbox work before cut-over despite the stated v2 dependency.

Identifies material uncertaintyRightMixedRight
GPT-6 Luna · API

It names key unknowns such as certification slip, hiring ramp, adoption, and correlational churn evidence, and says they would be reviewed at quarterly checkpoints.

Gemini 3.8 Flash · API

It names risks but does not say how they would be resolved or what would change the call; mitigations are given, but no explicit condition that would alter the roadmap.

Opus 5.5 · Claude

It names material unknowns such as certification failure, spend extrapolation, multi-entity conversion, latency, hiring, and procurement gate failure, with mitigations or decision triggers.

Avoids unsupported claimsRightWrongMixed
GPT-6 Luna · API

It labels the $4.5M interchange figure as a ceiling, the churn comparison as correlational, and the $2.9M multi-entity figure as upside rather than forecast.

Gemini 3.8 Flash · API

It presents the $2.9M multi-entity expansion as 'contracted expansion ARR' when only 9 of 14 have committed in writing, and the 52 subsidiaries are not all covered by those commitments.

Opus 5.5 · Claude

It presents false or under-supported current-situation claims, including that Approvals & Policy was idle in Tom's draft and that Intacct matters roughly five times more by lost deals.

Produces the required deliverableMixedMixedWrong
GPT-6 Luna · API

It provides both a quarter-by-quarter roadmap with outcomes and a case under 1,200 words for the exec team.

Gemini 3.8 Flash · API

The roadmap table and case are present, complete, and usable by the exec team with light edits.

Opus 5.5 · Claude

The roadmap and case are present and usable, but the case is too long for the requested pre-read format.

Protects the processor deadlineWrongRightRight
GPT-6 Luna · API

It finishes the migration with no Cards slack or added help, relying on an exact end-February target rather than protecting the four-week certification window.

Gemini 3.8 Flash · API

Migration build finishes by week 7 of Q1, leaving weeks 8-11 for certification before 31 March; virtual cards and spend controls start only after cut-over in Q2/Q3.

Opus 5.5 · Claude

It finishes migration by about 31 January, books February certification, cuts over around 1 March with buffer, and keeps production card-stack changes out of the freeze.

Makes the call on procurementWrongWrongMixed
GPT-6 Luna · API

It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.

Gemini 3.8 Flash · API

It checks the $6M claim and commits Approvals & Policy to better-evidenced work, but the proposed 'Procurement Discovery & Scoping (MVP)' lacks a clear threshold or paid test that would justify the full build.

Opus 5.5 · Claude

It checks the $6M claim but still commits a large procurement build before a paid-pilot threshold, rather than replacing the 30-week build with a cheap test.

Fits the stated capacityRightMixedMixed
GPT-6 Luna · API

The listed squad-weeks fit the stated 11-week quarterly capacity, and the output names deferred or non-committed work.

Gemini 3.8 Flash · API

Committed work per squad per quarter adds up to ≤11 squad-weeks with some slack (e.g., Integrations Q1 uses only 4 weeks), and the case names what was deferred (procurement monolith, Dynamics priority).

Opus 5.5 · Claude

Several committed items exceed candidate estimates or add unestimated work, such as virtual cards, procurement, and mobile beta, so the sums do not cleanly fit the stated capacity.

Sequences around dependenciesRightRightMixed
GPT-6 Luna · API

It places virtual cards and spend controls after v2 certification and names the key deadline, hiring, and cross-squad dependencies.

Gemini 3.8 Flash · API

Virtual cards and spend controls are placed after v2 migration; multi-entity and SSO are sequenced logically; dependencies like 'needs v2' are named.

Opus 5.5 · Claude

It respects the processor freeze but violates the virtual-cards dependency on v2 by scheduling sandbox build work before cut-over.

All got right 2

Plans on the squads we actually haveRightRightRight
GPT-6 Luna · API

It uses the hiring history to avoid counting new squads at full capacity and commits no important launch to them.

Gemini 3.8 Flash · API

New squads are planned with half capacity in their first quarter (Q2 for Squad A, Q3 for Squad B), consistent with 5-8 month hiring history, and no critical commitments depend on them earlier.

Opus 5.5 · Claude

It plans new squads as upside, with no committed deadline-critical or NRR-critical work depending on them before Q2-Q3.

Outcomes, with certainty that falls with distanceRightRightRight
GPT-6 Luna · API

Each item names an outcome or problem, and later-quarter items are deliberately looser than near-term commitments.

Gemini 3.8 Flash · API

Every roadmap item names an outcome or problem; near-term items are specific (e.g., 'Processor v2 Migration (11 wks)') and later ones are deliberately looser ('Cross-Entity Reporting & Enterprise Admin').

Opus 5.5 · Claude

Each item names an outcome, and later-quarter scope is looser than near-quarter commitments.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI89.6100.02None
2GPT-6 AstrawithChatGPT91.787.52None
3GPT-6.1 SolwithAPI91.784.62None
4GPT-6 LunawithAPI85.076.32None
5Opus 5.5withClaude80.565.12None
6Gemini 3.8 FlashwithAPI70.151.92None
7Gemini 3.5 Flash-LitewithGemini28.48.322 capped

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review