Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Outcomes, with certainty that falls with distance95% pass
    Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).
    Sonnet 5.5 · API · Two squads, eight asks, one half
  2. Sequences around dependencies89% pass
    Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.
    Sonnet 5.5 · API · Two squads, eight asks, one half
  3. Plans on the squads we actually have89% pass
    It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.
    GPT-6 Astra · ChatGPT · A year of spend management, with a hard deadline

Where it slips

  1. Makes the call on procurement50% pass
    It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.
    GPT-6 Luna · API · A year of spend management, with a hard deadline
  2. Produces the required deliverable50% pass
    The roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Uses the supplied evidence correctly61% pass
    Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.
    GPT-6 Astra · ChatGPT · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Harbor. Tom Achebe, our CEO, has shared his draft roadmap for the next four quarters (Q4 2026 to Q3 2027), and asked you to turn it into the roadmap we take to next week's planning offsite. The exec team (CEO, CFO, CRO and CTO) will read it beforehand. Write: 1. The roadmap itself, by quarter or by now, next and later: what each squad works on and the outcome each item serves. 2. The case for it, in no more than 1,200 words: why this order, what you changed from Tom's draft and why, what we're not doing, and the risks that could change the plan. The pack below is everything we have. Not all of it matters equally.

What the model was given12 items: About Harbor, FY27 goal (approved by the board, September 2026), Squads and capacity, Hiring, Card processor deadline, Candidate work (estimates in squad-weeks, by owning squad), Tom's draft roadmap, Procurement evidence, Card spend evidence, Multi-entity evidence, Integrations evidence, Other asks
About HarborSpend management (corporate cards, expenses and approvals) for companies with 200 to 2,000 employees. 1,100 customers, $86M ARR, average 700 employees per customer. 58% of revenue is interchange on card spend; the rest is subscription. Net revenue retention (NRR) is 104%.
FY27 goal (approved by the board, September 2026)Raise NRR to 112% by the end of Q3 2027 by expanding inside existing customers: more spend on Harbor cards, more entities and more modules. New-logo growth matters, but it comes second this year.
Squads and capacityFive squads: Cards, Expenses, Approvals & Policy, Integrations and Platform. After support and on-call, each has about 11 squad-weeks of roadmap capacity a quarter. Squads can take work from another squad's area, but it takes them about 1.5× as long.
HiringTwo new squads were approved in September 2026. Tom's draft counts the first from Q1 2027 and the second from Q2 2027, both at full capacity. Last year, our three approved squads took 5, 6 and 8 months from approval to their first full sprint, and each delivered about half its capacity in its first quarter.
Card processor deadlineOur card processor retires its v1 API on 31 March 2027. Every card authorisation we process runs through v1 today. Migrating to v2 is about 18 squad-weeks of Cards work, and the processor must then run 4 weeks of certification testing before we can cut over. Certification is the processor's time, not ours, but nothing can change on our side while it runs.
Candidate work (estimates in squad-weeks, by owning squad)1. Processor v2 migration: Cards 18. Hard deadline above. 2. Virtual cards for software subscriptions: Cards 10. Needs v2. 3. Real-time spend controls (declines out-of-policy spend at the point of sale): Cards 8, Approvals & Policy 4. Needs v2. 4. Procurement module (purchase requests and approvals, a new paid add-on): Approvals & Policy 30, Integrations 6. 5. Multi-entity support (subsidiaries under one parent account): Platform 12, Approvals & Policy 10. 6. Sage Intacct integration: Integrations 9. 7. Microsoft Dynamics integration: Integrations 12. 8. SSO and SCIM provisioning: Platform 5. 9. Receipt-matching improvements: Expenses 6. 10. Mobile app rewrite: Expenses 22.
Tom's draft roadmapQ4 2026: Cards: v2 migration. Approvals & Policy: multi-entity. Integrations: Intacct. Platform: multi-entity. Expenses: receipt matching. Q1 2027: Cards: finish migration, start virtual cards. New squad A: procurement module. Integrations: Dynamics. Platform: SSO and SCIM. Q2 2027: Cards: virtual cards, spend controls. New squad A: procurement launch. New squad B: mobile rewrite. Q3 2027: Cards: spend controls. Everyone else: procurement adoption, mobile launch. Tom's note: “Procurement is our path to 112%. I've told the board it can add $6M of ARR in its first year.”
Procurement evidenceSix design partners have used a prototype since June. Two said they would pay for it; the other four already use a dedicated procurement tool and said they'd need a reason to switch. Proposed price: $8 per user per month, for users who raise purchase requests. At our customers, about 10% of employees raise purchase requests.
Card spend evidenceIn QBRs with our 60 largest customers, 26 asked for virtual cards for software subscriptions, and 31 CFOs asked for real-time spend controls. Finance estimates our customers pay about $410M a year of software subscriptions on other cards or by invoice (extrapolated from the 60 QBR accounts). Our net interchange is 1.1% of card spend.
Multi-entity evidence14 of our 60 largest customers have asked for it. Between them they have 52 subsidiaries not on Harbor, and 9 of the 14 have said in writing they would add their subsidiaries once it exists. Customer Success sized it at $2.9M ARR if all 52 joined at their parents' pricing.
Integrations evidence17% of customers use Sage Intacct through a CSV export. Their gross revenue churn is 13% a year, against 7% for customers on our NetSuite integration. Intacct was cited in 31% of lost new-logo deals last year, Dynamics in 8%. 3% of customers use Dynamics.
Other asksCRO: “Five enterprise deals worth $1.2M of pipeline need SSO and SCIM, and 11 existing accounts have it on their security review list.” CTO: “The mobile app is on a framework version that loses support in late 2027, and I want the rewrite done before then.” Support: receipt auto-match is at 71%, and unmatched receipts are the most common expense ticket.
What a strong answer doesThe answer key the graders mark against

A four-quarter roadmap that protects the 31 March deadline: the migration build has to finish by about early March to leave 4 weeks of certification, which is roughly all of Cards' capacity for Q4 and most of Q1 (18 of about 18 squad-weeks available before certification). So it adds help or buffer (Platform or another squad at 1.5×, or cutting everything else from Cards) and moves virtual cards out of Q1: both v2 features start only after cut-over (virtual cards Q2, spend controls Q2–Q3). It treats the new squads as upside: on past hiring they arrive around February to April 2027 and deliver half capacity in their first quarter, so nothing committed depends on them before Q2–Q3. It makes a clear call on procurement: the $6M claim needs about 890 accounts (81% of customers) paying $6,720 a year each (70 requesters × $8 × 12), and only 2 of 6 design partners would pay, so it replaces the 30-week build with a cheap paid test with a threshold, and keeps Approvals & Policy on multi-entity and spend controls, which have stronger expansion evidence ($2.9M sized, 9 of 14 committed in writing; about $1.5M a year of interchange from moving a third of the $410M of subscriptions onto virtual cards). Intacct comes early for its churn; Dynamics later or cut; SSO and SCIM are small and justified; the mobile rewrite waits for a new squad, before the framework's support ends. It lists what changed from Tom's draft, what's not being done, and the risks (certification slipping, hiring slipping), with commitment falling with distance.

Critical failures (cap the score at 40)
  • Leaves the processor migration finishing too late for 4 weeks of certification before 31 March 2027, or schedules virtual cards or real-time spend controls before it
Case

Staff-level · v1.1 · synthetic · B2B SaaS, fintech, capacity, hard deadline, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 12 checks passed

The output is well-structured, evidence-based, and addresses all required elements; a PM could use it with minor wording edits without re-checking evidence or reasoning.

Output

Part 1: Roadmap, Q4 2026 to Q3 2027

Planning basis: Five existing squads at 11 squad-weeks a quarter (55 a quarter). The two new squads are counted as zero in the commitments (see the case). Lending between squads is costed at 1.5×. Figures in brackets are squad-weeks of the named squad.

Q4 2026: Protect the card rail, start entities

SquadWorkOutcome served
Cardsv2 migration (11). Ring-fenced, nothing else.Protect continuity of card processing (~$50M interchange)
IntegrationsLend to v2 migration (6, ≈4 Cards-weeks). Intacct (5 of 9).Protect the deadline; cut churn in Intacct customers
PlatformMulti-entity (11 of 12)More entities
Approvals & PolicyMulti-entity (10). Procurement pricing test and paid-pilot outreach (1).More entities; more modules
ExpensesReceipt matching (6). Mobile options and architecture, including an in-place upgrade check (5).Fewer support tickets; keep the mobile app supportable

Q1 2027: Cut over, ship entities, start controls

SquadWorkOutcome served
CardsFinish v2 (3), code complete by late January. Certification and change freeze in February. Cutover by end of February. Real-time spend controls build (8), released after cutover.Protect processing; more spend on Harbor cards
Approvals & PolicyReal-time spend controls policy side (4). Procurement build (7). Gate: procurement go/no-go at end of Q1.Card spend; more modules
PlatformFinish multi-entity (1), GA in January. SSO and SCIM (5). Reserve (5).More entities; retention and new-logo pipeline
IntegrationsFinish Intacct (4), GA. Procurement ERP sync (3). Reserve (4).Lower churn; more modules
ExpensesLend to procurement (5). Mobile rewrite (6).More modules; app supportability

Q2 2027: Virtual cards, procurement beta

SquadWorkOutcome served
CardsVirtual cards for software subscriptions (10). Reserve (1).More spend on Harbor cards
Approvals & PolicyProcurement (11), beta with design partners by end of Q2More modules
PlatformLend to procurement (9, ≈6 A&P-weeks). Reserve (2).More modules
IntegrationsProcurement ERP sync (3). Dynamics starts only if capacity remains (stretch).More modules
ExpensesLend to procurement (4). Mobile rewrite (7).More modules; app supportability
New squad A (if in seat)Bounded procurement items, ≈5 weeks at half capacity. Upside only.

Q3 2027: Launch and land

SquadWorkOutcome served
Approvals & PolicyProcurement GA and hardening (11), target JulyMore modules
ExpensesFinish and ship mobile (4). Reserve (7).App supported before the framework deadline
Cards, Platform, IntegrationsUncommitted. Post-launch iteration on controls and virtual cards, multi-entity rollout follow-ups, Dynamics. Allocated at the Q2 review.
New squadsSquad A: procurement. Squad B: Dynamics or the first FY28 area. Charters set at the Q2 review.

About 60 of 220 squad-weeks are deliberately uncommitted, mostly in Q3.

---

Part 2: The case (≈1,150 words)

The short version

The order is set by one thing we can't negotiate and by the evidence for everything else. The processor retires v1 on 31 March, every authorisation runs through it, and interchange is 58% of revenue. Cards does nothing else until cutover. After that, work is ordered by evidence and cost per squad-week.

The plan gets our exit run-rate to roughly 109–111% NRR, not reliably 112%. We should tell the board that now rather than in Q3.

Why this order

1. v2 first, with real slack. Tom's draft has Cards doing 18 weeks at 11 a quarter, so build finishes around late February. Four weeks of certification then ends around 25 March, about a week before a deadline that would stop all card authorisation. My plan ring-fences Cards in Q4 and has Integrations, which does API work daily, lend 6 weeks (≈4 Cards-weeks). Code is complete in late January, certification runs in February, and we keep about five weeks of slack.

2. Multi-entity next. It has the best evidence we hold. Nine of 14 requesting customers said in writing they'd add subsidiaries, which means ~33 of the 52 subsidiaries, or ~$1.9M ARR. Customer Success's $2.9M assumes all 52 join. The cost is 22 weeks, and it is a direct "more entities" lever.

3. Intacct, then SSO. 17% of customers (~187) use Intacct via CSV and churn at 13% against 7% on NetSuite. If the integration closes that gap, it retains ~$0.9M a year for 9 weeks of work. That assumes average ARR of $78k and that the gap is causal, which is unproven. Intacct also appears in 31% of lost deals. SSO and SCIM costs 5 weeks and covers $1.2M of pipeline plus 11 existing accounts' security reviews. Dynamics is 12 weeks for 3% of customers and 8% of lost deals, and it serves new-logo growth, so it is stretch.

4. Controls before virtual cards. Both need v2 live, so both follow cutover. 31 CFOs asked for real-time controls against 26 for virtual cards. The 4 Approvals & Policy weeks for controls fit in Q1, before procurement takes that squad's capacity. Virtual cards is the only card item we can size: $410M × 10–25% capture × 1.1% ≈ $0.5–1.1M. The $410M is extrapolated from our 60 largest accounts, so it is probably high. This is a judgment call. Swapping the order is cheap, but the Approvals & Policy slot would collide with the procurement build.

5. Procurement: build it, resize the promise. At $8 × 10% of a 700-person company's employees, a customer pays $6,720 a year. Even 100% attach across 1,100 customers is $7.4M, so $6M needs ~890 customers (81%). Two of six design partners would pay (33%), and four already own a procurement tool. At 10–33% attach, year-one ARR is $0.7–2.5M, and with GA in Q3 only a small part lands before Q3 ends. It is still worth building: it is our only new paid module, it recurs, and it matches the board's "more modules" lever. But it is not a path to 112% on its own.

The arithmetic: 8 points of NRR on $86M is ~$6.9M. Sized items total roughly $4–6M of exit run-rate: entities $1.9M, Intacct $0.9M, virtual cards $0.5–1.1M, and procurement $0.7–2.5M (mostly after Q3). Controls and SSO retention are unsized upside. Recognised NRR at Q3 will lag run-rate because late launches haven't ramped.

What I changed from Tom's draft

  • New squads count as zero. Last year's three squads took 5, 6 and 8 months from approval to first full sprint, then delivered half capacity in their first quarter. From a September approval, squad A is likely in seat between February and May and squad B later. A realistic ~5 weeks in Q2 and ~11 in Q3 for A, plus ~5 in Q3 for B, is about 22 weeks, against 55 in the draft. Even at Tom's capacity, procurement (36 weeks) couldn't launch in Q2, because squad A would have 22 weeks by the end of Q2.
  • Procurement moves to Approvals & Policy. They own the domain and, in the draft, have no work after Q4. Expenses and Platform lend 9 weeks each (≈6 A&P-weeks each) to get a beta by the end of Q2 and GA in Q3. New squads join as upside.
  • Mobile moves to Expenses. The draft puts it on squad B, which may not exist until Q3. Expenses has ~38 uncommitted weeks, and the rewrite ships in Q3, ahead of the late-2027 framework cutoff.
  • Controls before virtual cards (point 4 above).
  • Intacct GA moves from Q4 to Q1 to fund the v2 help. The cost is roughly one quarter of a ~$0.9M-a-year benefit, about $0.2M.
  • Q3's "everyone else: adoption" is replaced by a held reserve, because adoption isn't squad work and Q3 launches can't move FY27 NRR much.
  • A procurement checkpoint is added at the end of Q1 (see risks).

What we're not doing

  • Dynamics, beyond a stretch start in Q2–Q3.
  • A Q2 procurement launch, or putting the $6M figure in front of the board.
  • Any new Cards scope before cutover.
  • A mobile rewrite staffed by a squad that doesn't exist yet.
  • Planning on capacity from squads we haven't hired.

Risks that could change the plan

  1. Processor timeline. If v2 isn't code-complete by 31 January, Platform joins Cards and controls and virtual cards slip a quarter. Two asks for this month: a written fallback or extension from the processor, and confirmation that feature work behind flags is allowed during the certification freeze (my assumption). Book the certification slot now.
  2. Procurement demand. At the end-of-Q1 gate I propose requiring at least 10 signed paid commitments at $8 from ~30 accounts approached (≈$67k ARR). If we miss it, we cut scope and redirect Approvals & Policy and the lenders to multi-entity rollout and card-spend work. The threshold is my proposal and is open to debate.
  3. The NRR target itself. Exec decision: hold 112% as a Q3 recognised-NRR target, or restate it as an exit run-rate target, which I think is achievable only with upside from controls and procurement.
  4. Evidence quality. The $410M is extrapolated from our largest accounts. The Intacct churn gap is a correlation. The multi-entity sizing rests on 14 of 60 customers. Any could be off by a factor of two.
  5. Lending efficiency. The 1.5× penalty is an average. Expenses and Platform have less procurement context, and if they're slower than assumed, GA slips into Q4.
  6. Hiring. If squad A lands in February, procurement GA can move earlier. If squad B lands late, nothing in the plan changes.
  7. Mobile date. The CTO should confirm the exact end-of-support date and whether an in-place upgrade avoids a 22-week rewrite. Q4's scoping work answers this.

Grades and run details

Decision model 79 · LLM judge 13 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly14%
  • passAddresses the actual decision93%
  • passRespects explicit constraints30%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims25%
  • failProduces the required deliverable8%
  • passProtects the processor deadline75%
  • passPlans on the squads we actually have99%
  • passMakes the call on procurement71%
  • passOutcomes, with certainty that falls with distance93%
  • passFits the stated capacity44%
  • passSequences around dependencies51%
Run
Run
#1
API response time
3.4 min
Submitted
30 Sept 2026
Critical failure: Leaves the processor migration finishing too late for 4 weeks of certification before 31 March 2027, or schedules virtual cards or real-time spend controls before it. Score capped at 40.

Needs more than a quick edit0 of 12 checks passed

A PM would need to redo the hiring assumptions, procurement math, capacity sums, and card migration sequencing. The output cannot be used with only wording or emphasis edits.

Output

Harbor FY27 Product Roadmap

Prepared by: Staff PM, Harbor For: Executive Team Review (CEO, CFO, CRO, CTO)

---

Part 1: The FY27 Roadmap (Q4 2026 – Q3 2027)

Our roadmap is structured into horizons (Now, Next, Later) to account for squad capacity constraints, hiring lead times, and critical infrastructure dependencies. Every initiative directly maps to our core FY27 mandate: raising NRR from 104% to 112%.

NOW (Q4 2026)

Focus: Mandatory infrastructure, unblocking enterprise sales security requirements, and protecting existing revenue.

  • Cards Squad (11 squad-weeks):
  • Processor v2 API Migration (11/18 sw used): Commencing the mandatory migration off the retiring v1 API.
  • Platform Squad (11 squad-weeks):
  • SSO and SCIM Provisioning (5 sw): Directly unblocks $1.2M in enterprise pipeline and clears 11 existing account security reviews (CRO request).
  • Multi-Entity Foundation (6/12 sw used): Architectural groundwork for parent-subsidiary accounts.
  • Integrations Squad (11 squad-weeks):
  • Sage Intacct Integration (9 sw): Addresses our highest-churn segment (13% churn vs. 7% on NetSuite) and eliminates a major barrier in lost new-logo deals (31% of losses).
  • Spillover buffer (2 sw).
  • Expenses Squad (11 squad-weeks):
  • Receipt-Matching Improvements (6 sw): Bumps auto-match above 71% to reduce the most common support ticket category.
  • Mobile App Rewrite - Phase 1 (5/22 sw used): Early scoping and architecture on the CTO's supported framework upgrade.
  • Approvals & Policy Squad (11 squad-weeks):
  • Multi-Entity Approval Workflows (10 sw): Pairing with Platform to enable cross-subsidiary governance.
  • Spillover buffer (1 sw).

---

NEXT (Q1 2027)

Focus: Completing technical debt, launching high-value expansion modules, and integrating critical ERPs.

  • Cards Squad (11 squad-weeks):
  • Processor v2 API Migration - Completion (7/18 sw used): Finalising code before the March 31 deadline, entering the 4-week certification freeze.
  • Virtual Cards for Software Subscriptions - Initiation (4/10 sw used).
  • Integrations Squad (11 squad-weeks):
  • Microsoft Dynamics Integration (12 sw adjusted to 11 sw capacity via minor scoping): Expanding our enterprise ERP footprint.
  • Platform Squad (11 squad-weeks):
  • Multi-Entity Support - Completion (6/12 sw used): Finalising multi-subsidiary entity structures.
  • Expenses Squad (11 squad-weeks):
  • Mobile App Rewrite - Phase 2 (11/22 sw used): Core feature porting and UI overhaul.
  • New Squad A (Hired Dec 2026, delivering at ~50% capacity = ~5.5 sw):
  • Procurement Module - Phase 1 (5.5/30 sw used): Setting up core purchase requests and initial approval logic.

(Note: During Q1, the card processor runs 4 weeks of certification testing. No changes can be deployed to card flows during this time).

---

LATER (Q2 – Q3 2027)

Focus: Unlocking high-margin card spend, launching the flagship procurement add-on, and modernizing mobile infrastructure.

  • Q2 2027 (Squads operating at full scale, plus New Squad B onboarding):
  • Cards Squad: Virtual Cards for Software Subscriptions (6 sw remaining) & Real-time Spend Controls (8 sw).
  • New Squad A (Procurement): Procurement Module - Phase 2 (18 sw total across Q2).
  • Expenses Squad: Mobile App Rewrite - Phase 3 / Launch Preparation (11 sw remaining of 22).
  • New Squad B (Hiring initiated Q1, arriving Q2, delivering ~5.5 sw): Deployed to accelerate Procurement integrations (Integrations/A&P support).
  • Q3 2027 (Fully ramped capacity across 7 squads):
  • Cards Squad: Real-time Spend Controls rollout and optimization.
  • New Squad A & B: Procurement Module Launch, GA, and initial adoption drives.
  • Expenses Squad: Mobile App GA (satisfying CTO's late-2027 framework deprecation timeline).
  • Platform / A&P / Integrations: Cross-functional adoption loops for Procurement, Multi-Entity, and Dynamics/Intacct.

---

Part 2: The Case for the Roadmap

1. Why This Order? (Balancing Reality and Ambition)

Our mandate is clear: raise NRR to 112% by the end of Q3 2027. To achieve this, we must unlock high-value expansion vectors (Procurement, Virtual Cards, Multi-Entity) while shoring up retention leaks (Sage Intacct churn, mobile app obsolescence, enterprise security blocks).

However, ambition must bow to physics. Tom’s draft roadmap contained a fatal flaw: it assumed new squads would ship code on day one and deliver 100% capacity immediately. Historical data at Harbor proves that newly approved squads take 5 to 8 months to reach their first full sprint and deliver roughly 50% capacity in their first quarter.

Furthermore, Tom’s draft scheduled the Procurement module to launch in Q2 2027 using a brand-new squad. Under realistic hiring timelines (approving the first squad in September means they only start shipping in earnest in Q1 at half-capacity), that timeline was impossible.

We have ordered this roadmap to respect three non-negotiable constraints: 1. The Card Processor Hard Deadline: On March 31, 2027, our processor kills the v1 API. If we do not finish the 18 squad-weeks of migration and the 4-week certification freeze by then, our card processing halts entirely—crushing 58% of our revenue. 2. Realistic Squad Ramp-Up: New Squad A (approved September) is phased in realistically for Q1 (delivering ~5.5 squad-weeks) and hits full stride in Q2. New Squad B is phased in for Q2. 3. Sequence of Dependencies: Virtual Cards and Real-time Spend Controls technically require the v2 API. Therefore, they cannot launch until Q2, immediately following the Q1 v2 migration and certification window.

2. What We Changed from Tom’s Draft and Why

  • Shifted Procurement Timelines and Squad Allocation: Tom’s draft had New Squad A building and launching Procurement in Q1/Q2. Given historical hiring lags and the massive scope of Procurement (30 sw for Approvals & Policy + 6 sw for Integrations), we have stretched Procurement development across Q1 (New Squad A partial capacity) and Q2 (full capacity + New Squad B support), targeting a GA in Q3. Trying to rush it in Q2 would result in a broken launch or squad burn-out.
  • Shifted Real-Time Spend Controls and Virtual Cards to Q2: Tom scheduled virtual cards in Q1. Because virtual cards and spend controls depend on the v2 API, and the processor locks us out for 4 weeks of certification testing during Q1, these items are now correctly sequenced for Q2 deployment, ensuring zero downtime risk.
  • Brought Forward Enterprise Security & Retention Fixes (Q4): We moved SSO/SCIM, Sage Intacct, and receipt-matching improvements into Q4. Why? Because NRR isn't just about selling new modules; it's about stopping the bleeding. Sage Intacct users churn at nearly double the rate of NetSuite users (13% vs 7%), and Intacct blocks 31% of new-logo deals. Fixing this now preserves our base. Similarly, securing $1.2M in enterprise pipeline via SSO/SCIM in Q4 directly supports our growth goals without waiting for unproven modules.

3. What We Are Not Doing (And Why)

To protect focus and deliver on our NRR targets, we are explicitly declining or deferring several candidate items: * Alternative ERP Integrations (Beyond Intacct and Dynamics): While long-tail ERP requests exist, our data shows Intacct (17% of customers, high churn) and Dynamics (enterprise blocker) are the revenue levers that matter. Other ERPs are deferred to FY28. * Ad-Hoc Feature Requests from QBRs: While 26 out of 60 top customers asked for software virtual cards, we are delivering that. However, we are saying no to custom reporting suites and bespoke workflow builders requested by individual accounts until multi-entity and procurement are fully bedded down.

4. Risks That Could Change the Plan

1. Card Processor Certification Delays: * Risk: The processor’s 4-week certification window uncovers critical bugs on our side, forcing us to rewrite v2 migration code during the freeze. * Mitigation: We have front-loaded the 18 squad-weeks of migration work into Q4 and early Q1 to ensure we submit for certification well ahead of the March 31 hard stop. 2. Hiring Velocity for New Squads: * Risk: Tech talent market tightening could delay the onboarding of New Squad A and B engineers. * Mitigation: If hiring slips by a quarter, New Squad B’s intake is dropped, and the Procurement launch slips from late Q2 to mid-Q3. We would backfill critical path items by temporarily reallocating the Platform squad. 3. Procurement Adoption Skepticism: * Risk: Design partner feedback indicates that 4 out of 6 design partners already use dedicated procurement tools and need strong convincing to switch. Tom’s assumption of $6M ARR in Year 1 may be aggressive if attach rates lag. * Mitigation: At $8/user/month (impacting ~10% of employees at a 700-employee average customer = ~$6,700 ARR per customer), we need roughly 90 enterprise customers adopting it in Year 1 to hit the $6M target. Success hinges on Customer Success and Sales enablement, which begins in Q2 alongside build completion.

Claims the judge couldn’t find in the brief

  • About 90 enterprise customers adopting procurement in Year 1 would hit Tom's $6M ARR claim.
  • New Squad A was hired in December 2026 and can deliver about 5.5 squad-weeks in Q1 2027.
  • New Squad B can be initiated in Q1 2027, arrive in Q2 2027, and deliver about 5.5 squad-weeks.
  • Dynamics is an enterprise blocker and a revenue lever that matters.

Grades and run details

Decision model 25 · LLM judge 0 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly8%
  • partialAddresses the actual decision32%
  • failRespects explicit constraints82%
  • passIdentifies material uncertainty27%
  • partialAvoids unsupported claims27%
  • failProduces the required deliverable71%
  • failProtects the processor deadline57%
  • partialPlans on the squads we actually have56%
  • failMakes the call on procurement76%
  • partialOutcomes, with certainty that falls with distance51%
  • failFits the stated capacity43%
  • failSequences around dependencies82%
Run
Run
#1
Time to output
10 s
Submitted
1 Oct 2026

Needs more than a quick edit5 of 12 checks passed

A PM would need to shorten the case, correct the false draft claim, and rework procurement and virtual-card dependencies before using it. The core sequencing is strong but not ready without substantive edits.

Output

Harbor Roadmap, Q4 2026 to Q3 2027

Pre-read for the planning offsite. Prepared from Tom's draft.

Part 1: The roadmap

Planning assumptions - Each existing squad has 11 squad-weeks of roadmap capacity a quarter. Work done in another squad's area costs 1.5×. - The two new squads are planned at the dates our history supports, not the approved dates. Last year's squads took 5 to 8 months from approval to their first full sprint. - Squad A: 0 weeks in Q1, about 5 in Q2 (half capacity), about 11 in Q3. - Squad B: 0 weeks in Q2, about 5 in Q3. - Nothing committed depends on either new squad. - The fiscal year runs Q4 2026 (Oct to Dec) through Q3 2027 (Jul to Sep).

Squad plan by quarter (squad-weeks in brackets)

SquadQ4 2026Q1 2027Q2 2027Q3 2027
Cardsv2 migration (11)v2 migration: code complete by ~31 Jan (4). Certification runs through Feb with no card-stack changes. Cutover ~1 Mar. Cutover contingency (3). Virtual cards design and sandbox build, no production changes (4)Virtual cards build and launch ~May (6). Real-time spend controls (5)Spend controls launch ~Jul (3). Adoption work on virtual cards and controls with CS (8)
Platformv2 migration support (5, which buys ~3.3 Cards-weeks). Multi-entity (6)Multi-entity, ships ~end Feb (6). SSO and SCIM, ships by end Mar (5)Subsidiary onboarding tooling and SSO rollout to the 11 accounts (5). Procurement billing and permissions, if the gate passes (6)New-squad onboarding and spend-controls latency support (5). Reserve (6)
Approvals & PolicyMulti-entity (10). Procurement pricing test with design partners (1)Procurement build, phase 1 (11). Gate at end of MarchSpend controls policy engine (4). Procurement (7)Procurement finish (8). Beta, then GA ~Sep (3)
IntegrationsSage Intacct, ships ~early Dec (9). CSV-to-integration migration tooling (2)Move ~187 Intacct CSV customers onto the integration, with CS (5). Dynamics (6)Dynamics, ships ~May (6). Procurement integrations (5)Procurement integrations (1). Hardening and reserve (10)
ExpensesReceipt matching (6). Mobile rewrite (5)Mobile rewrite (11)Mobile rewrite, feature-complete ~May (6). Beta and fixes (5)Staged rollout. Old app retired by end Sep, a quarter or more before framework support ends
New squad A (upside)——Onboarding (~5)~11: accelerate procurement, then own it. Otherwise, adoption work on cards
New squad B (upside)———Onboarding (~5). Scope set at Q2 planning

Items, outcomes and key dates

ItemOutcome it servesKey dateOwner
Processor v2 migrationKeeps every card authorisation running. That protects the 58% of revenue that is interchangeCode complete ~31 Jan, cutover ~1 Mar, about 4 weeks of buffer before 31 MarCards + Platform
Multi-entityMore entities: 9 customers have committed in writing to add subsidiariesShips ~end Feb. Subsidiaries onboard in Q2 and Q3Platform + A&P
Virtual cardsMore card spend: brings software subscriptions onto Harbor cardsLaunches ~MayCards
Real-time spend controlsMore card spend: gives CFOs the confidence to issue more cards (31 CFOs asked for it)Launches ~JulCards + A&P
Intacct integration and CSV migrationRetention: Intacct-on-CSV customers churn at 13%, against 7% on NetSuite. Also addresses 31% of lost dealsIntegration ships Dec. Customers move in Q1Integrations
SSO and SCIMRetention of the 11 accounts with it on their security review list. Also $1.2M of new-logo pipelineShips MarPlatform
ProcurementMore modules: a new paid add-onGate end Mar. GA ~SepA&P
DynamicsNew-logo win rate (8% of lost deals)Ships ~MayIntegrations
Receipt matchingLower support load: the most common expense ticketShips DecExpenses
Mobile rewriteRemoves platform risk before the framework loses supportOld app retired by SepExpenses

NRR bridge: what this plan plausibly delivers by end of Q3 2027

Moving NRR from 104% to 112% needs about $6.9M of additional net expansion (8 points of $86M).

LeverBasisEstimate in window
Multi-entity9 of the 14 requesting customers committed in writing, roughly 33 of 52 subsidiaries. The $2.9M figure assumes all 52 join$1.5M to $2.9M
Virtual cards$410M × 1.1% = $4.5M ceiling. Assumes 15 to 35% captured within ~5 months of launch$0.7M to $1.6M
Intacct churn reduction~$14.6M of ARR on CSV. Closing the churn gap is worth ~$0.9M a year, and only part of that lands in window$0.3M to $0.6M
ProcurementGA in ~SepUnder $0.3M
Spend controls, SSO and SCIMNot sizedUpside or protection
Total$2.8M to $5.4M, or roughly 107% to 110% NRR

---

Part 2: The case for it

The short version

Tom's draft has the right ingredients. It does, however, have three problems:

  1. No buffer on the processor deadline. The migration has no room to slip against a date that could stop every card authorisation.
  2. Key work sits on squads we won't have. Procurement and the mobile rewrite depend on new squads that history says won't exist at full capacity until Q3 at the earliest.
  3. Procurement can't carry the goal. It is expected to deliver the NRR target, but the evidence doesn't support $6M and the timing puts almost none of it inside the measurement window.

This plan protects the deadline first. It then sequences the levers with the strongest evidence so they land early enough to count. Even so, product alone likely gets us to 107% to 110%, not 112%. The exec team should know that now, not in Q3.

Why this order

1. The processor migration comes first, with buffer.

Every authorisation runs through v1, so a missed cutover puts most of our revenue at risk. On Tom's draft: - Cards alone finishes the 18 weeks around late February. - The 4-week certification then ends in the last days of March. - That leaves under a week of slack, with December holidays not counted and no room for a failed certification.

Lending 5 Platform weeks in Q4 changes this: - Code complete moves to about 31 January. - Cutover lands around 1 March. - We have roughly four weeks of buffer.

This also means virtual cards can't "start in Q1" as the draft says. During certification nothing on our card stack can change. Cards will do design and sandbox work in Q1 and start production work after cutover.

2. Multi-entity and virtual cards next, because they have the best evidence and they land in time.

NRR is measured at the end of Q3, so anything launching after about June barely counts. Our two strongest levers are: - Multi-entity: 9 customers have committed in writing to add subsidiaries, which is about $1.5M to $2.9M. - Virtual cards: a $4.5M ceiling. 26 of our 60 largest customers asked for them.

Multi-entity ships in February. Virtual cards ship in May. Real-time spend controls follow in July because they share the Cards squad. They matter to 31 CFOs, but we haven't sized them.

3. Retention work runs in parallel, because it is cheap.

  • Intacct: 9 weeks of work. It addresses a segment churning at nearly twice our NetSuite rate and appears in 31% of lost deals. The integration only reduces churn if customers actually move off CSV, so I added a Q1 migration push with CS.
  • SSO and SCIM: 5 weeks of work. It protects 11 accounts and unblocks $1.2M of pipeline.

4. Procurement is real but gated, and it is not the FY27 lever.

What I changed from Tom's draft, and why

  1. Migration buffer. Platform lends 5 weeks in Q4. The cost is that multi-entity moves from December to February. I think that trade is clearly right, because the downside of a missed cutover is existential.

2. New squads planned realistically. - Squads approved in September reach their first full sprint between February and May, based on last year's 5 to 8 months. - Each then runs at half capacity for a quarter. - Squad A realistically gives about 5 weeks in Q2. Squad B gives nothing before Q3. - I've kept every committed item on the existing five squads. The new squads are upside.

3. Procurement moved to Approvals & Policy and gated. - A&P owns the domain, has capacity after multi-entity, and was left without work in the draft. - Build starts in Q1. At the end of March we need signed paid-pilot commitments from about 10 customers beyond the two design partners, or we redirect A&P. - GA lands around September. If Squad A arrives on time, it accelerates the work.

4. The $6M procurement figure needs resetting with the board. - At $8 per user per month and 10% of our ~770,000 employees raising requests, the ceiling is about $7.4M a year if every customer adopts. $6M therefore implies about 81% adoption. - Only 2 of 6 design partners said they'd pay. The other four already use a dedicated tool. - At a one-in-three adoption rate, procurement is worth about $2.5M at full run-rate, reached over years, and less than $0.3M inside FY27. - I recommend Tom and the CFO correct this at the next board touchpoint, before the board starts tracking it.

5. Mobile rewrite moved to Expenses, starting in Q4. - Expenses owns the app and the draft left it idle after Q4. - The rewrite finishes by May, and the old app is retired by September. - This gives the CTO's deadline real margin. Leaving it on Squad B would risk missing it.

  1. Spend controls now include the A&P work (4 weeks). The draft omitted this.
  1. Dynamics moves from Q1 to May. The Intacct customer migration goes first, because Intacct matters roughly five times more by customer count and lost deals.

What we're not doing

  • Not claiming procurement drives FY27 NRR. We are also not building past the March gate without paid commitments.
  • Not starting virtual cards or anything else on the card stack before cutover.
  • Not putting any deadline-critical or NRR-critical work on unhired squads.
  • Not adding new-logo-only features beyond Intacct, Dynamics and SSO, which also serve retention. New logos come second this year, as the board set.
  • Not filling the Q3 reserve on Platform and Integrations. That slack absorbs migration fallout, integration hardening and new-squad onboarding. If the migration goes cleanly, we can release it to card-spend adoption work at Q2 planning.

Risks that could change the plan

1. Certification fails or needs a second round. A re-run would consume the entire buffer. - Book the certification slot now. - Run a pre-certification dry run in January. - Ask the processor in writing for a contingency extension and an emergency-fix protocol for the freeze window. - If code complete slips past mid-February, pull more Platform capacity immediately.

2. The $410M software-spend figure is extrapolated from our 60 largest accounts. Smaller customers may put proportionally less software spend on cards. - In Q4, validate it with AP data from about 20 mid-sized customers. - If it halves, virtual cards is worth about $0.4M to $0.8M in window and spend controls rises in priority.

  1. Multi-entity conversion comes in below the written commitments. CS should agree subsidiary onboarding plans with the 9 customers in Q1, before the feature ships.
  1. Real-time spend controls sit on the authorisation path. A latency problem could cause false declines. Platform supports this work in Q3, and we should launch in stages.
  1. Hiring slips further. This doesn't threaten committed dates, but procurement would stay in Q3 and our adoption capacity shrinks.
  1. The procurement gate fails. A&P redirects to deeper spend controls and multi-entity policy features, both of which feed card spend.
  1. The mobile framework end-of-support date is only "late 2027". The CTO should confirm the exact date.

Decisions we need at the offsite

  1. Approve Platform's Q4 loan to Cards, accepting that multi-entity moves to February.
  2. Agree the board message on procurement, and the gate criteria behind it.
  3. Close the NRR gap. This plan lands about 107% to 110%. Closing the remaining $1.5M to $4M needs commercial levers the CRO and CFO own, such as CS-led card activation, packaging and pricing, and a multi-entity sales motion beyond the 9. Otherwise we tell the board now that 112% is a stretch.
  4. Authorise the processor extension request.

Claims the judge couldn’t find in the brief

  • Tom's draft left Approvals & Policy without work.

Grades and run details

Decision model 79 · LLM judge 5 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly16%
  • passAddresses the actual decision98%
  • partialRespects explicit constraints33%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims20%
  • failProduces the required deliverable66%
  • passProtects the processor deadline98%
  • passPlans on the squads we actually have96%
  • passMakes the call on procurement59%
  • passOutcomes, with certainty that falls with distance88%
  • passFits the stated capacity39%
  • passSequences around dependencies62%
Run
Run
#1
Time to output
3.5 min
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 12

Uses the supplied evidence correctlyMixedWrongWrong
Sonnet 5.5 · API

Every factual claim about the current situation is taken directly from the brief or supplied context, with no invented facts.

Gemini 3.5 Flash-Lite · Gemini

It miscomputes the procurement adoption math, invents hiring dates and squad arrival assumptions, and adds unsupported Dynamics and ad-hoc request facts.

Opus 5.5 · Claude

Most quantitative evidence is used correctly, but the output makes at least one false current-situation claim about Tom's draft leaving Approvals & Policy without work.

Addresses the actual decisionRightWrongRight
Sonnet 5.5 · API

The output commits to a clear roadmap and case, makes an explicit call on procurement with a gate, and states what would change the plan.

Gemini 3.5 Flash-Lite · Gemini

It commits to a roadmap, but the procurement decision is not a clear evidence-based call because it keeps the full build and uses a 10x-wrong adoption threshold.

Opus 5.5 · Claude

It commits to a clear four-quarter order and decision set for the exec team, including the migration buffer, procurement gate, and NRR gap.

Respects explicit constraintsRightWrongWrong
Sonnet 5.5 · API

The roadmap is by quarter with squads and outcomes, the case is under 1,200 words, and all requested elements (order, changes, not doing, risks) are included.

Gemini 3.5 Flash-Lite · Gemini

It violates the processor sequencing constraint by starting virtual cards in Q1 and relies on unhired squads arriving earlier than the hiring history supports.

Opus 5.5 · Claude

It exceeds the 1,200-word limit for the case and proposes virtual-card sandbox work before cut-over despite the stated v2 dependency.

Identifies material uncertaintyRightMixedRight
Sonnet 5.5 · API

The output names specific unknowns (processor timeline, procurement demand, NRR target, evidence quality, lending efficiency, hiring, mobile date) and says how they would be resolved or change the plan.

Gemini 3.5 Flash-Lite · Gemini

It names hiring and certification risks but does not bound the key procurement unknown with a threshold or test that would change the build decision.

Opus 5.5 · Claude

It names material unknowns such as certification failure, spend extrapolation, multi-entity conversion, latency, hiring, and procurement gate failure, with mitigations or decision triggers.

Avoids unsupported claimsRightWrongMixed
Sonnet 5.5 · API

Interpretations and forecasts are clearly labelled as such (e.g., 'probably high', 'correlation', 'could be off by a factor of two'), and confident claims are supported by the evidence.

Gemini 3.5 Flash-Lite · Gemini

It presents invented hiring timing, Dynamics blocker status, and custom reporting/workflow requests as current facts.

Opus 5.5 · Claude

It presents false or under-supported current-situation claims, including that Approvals & Policy was idle in Tom's draft and that Intacct matters roughly five times more by lost deals.

Produces the required deliverableMixedWrongWrong
Sonnet 5.5 · API

The roadmap and case are complete, in the requested form for the exec team, within the word limit, and usable with light edits.

Gemini 3.5 Flash-Lite · Gemini

The pack is within length and has roadmap plus case, but it is not usable without redoing capacity, hiring, procurement, and deadline sequencing.

Opus 5.5 · Claude

The roadmap and case are present and usable, but the case is too long for the requested pre-read format.

Protects the processor deadlineRightWrongRight
Sonnet 5.5 · API

The migration finishes code-complete by late January, certification runs in February, cutover by end of February, with slack before 31 March; virtual cards and spend controls start only after cutover.

Gemini 3.5 Flash-Lite · Gemini

It schedules virtual cards in Q1 alongside migration and certification, despite the requirement that v2 features start only after cut-over.

Opus 5.5 · Claude

It finishes migration by about 31 January, books February certification, cuts over around 1 March with buffer, and keeps production card-stack changes out of the freeze.

Plans on the squads we actually haveRightWrongRight
Sonnet 5.5 · API

The new squads are treated as zero in commitments, with realistic arrival and half capacity based on past hiring, and nothing important depends on them before Q2-Q3.

Gemini 3.5 Flash-Lite · Gemini

It commits procurement to New Squad A in Q1 and New Squad B in Q2 based on unsupported December/February hiring assumptions rather than treating new squads as upside.

Opus 5.5 · Claude

It plans new squads as upside, with no committed deadline-critical or NRR-critical work depending on them before Q2-Q3.

Makes the call on procurementRightWrongMixed
Sonnet 5.5 · API

The output checks the $6M claim against pricing and adoption evidence, shows it implies ~81% attach, weighs it against 2 of 6 design partners, commits Approvals & Policy to better-evidenced work, and proposes a paid test with a threshold.

Gemini 3.5 Flash-Lite · Gemini

It checks Tom's $6M claim but calculates roughly 90 customers instead of about 890, and still commits to the full 30-week procurement build.

Opus 5.5 · Claude

It checks the $6M claim but still commits a large procurement build before a paid-pilot threshold, rather than replacing the 30-week build with a cheap test.

Outcomes, with certainty that falls with distanceRightWrongRight
Sonnet 5.5 · API

Every roadmap item names its outcome or problem, near-term items are specific with dates, and later items are deliberately looser (e.g., Q3 uncommitted, allocated at Q2 review).

Gemini 3.5 Flash-Lite · Gemini

Many items name outcomes, but later items remain overly specific and several commitments are not tied to the NRR expansion outcome.

Opus 5.5 · Claude

Each item names an outcome, and later-quarter scope is looser than near-quarter commitments.

Fits the stated capacityRightWrongMixed
Sonnet 5.5 · API

The committed work sums to 55 squad-weeks per quarter with slack, the sums are checkable, and it names what was cut or deferred (Dynamics, Q2 procurement launch, etc.).

Gemini 3.5 Flash-Lite · Gemini

Committed work exceeds stated capacity in places, such as Dynamics at 12 squad-weeks against 11, and it does not show how cross-squad or new-squad work fits.

Opus 5.5 · Claude

Several committed items exceed candidate estimates or add unestimated work, such as virtual cards, procurement, and mobile beta, so the sums do not cleanly fit the stated capacity.

Sequences around dependenciesRightWrongMixed
Sonnet 5.5 · API

All dependencies are respected (v2 before virtual cards and controls, hiring before new squad work, procurement gate), and the ones driving the order are named.

Gemini 3.5 Flash-Lite · Gemini

It names v2 dependencies but still places virtual cards before cut-over and sequences procurement around assumed new squads rather than proven capacity.

Opus 5.5 · Claude

It respects the processor freeze but violates the virtual-cards dependency on v2 by scheduling sandbox build work before cut-over.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI89.6100.02None
2GPT-6 AstrawithChatGPT91.787.52None
3GPT-6.1 SolwithAPI91.784.62None
4GPT-6 LunawithAPI85.076.32None
5Opus 5.5withClaude80.565.12None
6Gemini 3.8 FlashwithAPI70.151.92None
7Gemini 3.5 Flash-LitewithGemini28.48.322 capped

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review