Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Outcomes, with certainty that falls with distance95% pass
    Every item names the outcome or problem it serves; near-term items have specific targets (mid-December, January) while later items are looser (trigger-based, contract-dependent).
    Sonnet 5.5 · API · Two squads, eight asks, one half
  2. Sequences around dependencies89% pass
    Messaging service is built before SMS reminders and waitlist auto-fill, reminders before waitlist, and deposits are placed after the contract can be signed; the key dependencies are named.
    Sonnet 5.5 · API · Two squads, eight asks, one half
  3. Plans on the squads we actually have89% pass
    It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.
    GPT-6 Astra · ChatGPT · A year of spend management, with a hard deadline

Where it slips

  1. Makes the call on procurement50% pass
    It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.
    GPT-6 Luna · API · A year of spend management, with a hard deadline
  2. Produces the required deliverable50% pass
    The roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Uses the supplied evidence correctly61% pass
    Most numbers and quotes are correct, but the output presents 'undermines booking reliability' as a current fact when the supplied context only says calendar sync failures cause 38% of support tickets.
    GPT-6 Astra · ChatGPT · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Harbor. Tom Achebe, our CEO, has shared his draft roadmap for the next four quarters (Q4 2026 to Q3 2027), and asked you to turn it into the roadmap we take to next week's planning offsite. The exec team (CEO, CFO, CRO and CTO) will read it beforehand. Write: 1. The roadmap itself, by quarter or by now, next and later: what each squad works on and the outcome each item serves. 2. The case for it, in no more than 1,200 words: why this order, what you changed from Tom's draft and why, what we're not doing, and the risks that could change the plan. The pack below is everything we have. Not all of it matters equally.

What the model was given12 items: About Harbor, FY27 goal (approved by the board, September 2026), Squads and capacity, Hiring, Card processor deadline, Candidate work (estimates in squad-weeks, by owning squad), Tom's draft roadmap, Procurement evidence, Card spend evidence, Multi-entity evidence, Integrations evidence, Other asks
About HarborSpend management (corporate cards, expenses and approvals) for companies with 200 to 2,000 employees. 1,100 customers, $86M ARR, average 700 employees per customer. 58% of revenue is interchange on card spend; the rest is subscription. Net revenue retention (NRR) is 104%.
FY27 goal (approved by the board, September 2026)Raise NRR to 112% by the end of Q3 2027 by expanding inside existing customers: more spend on Harbor cards, more entities and more modules. New-logo growth matters, but it comes second this year.
Squads and capacityFive squads: Cards, Expenses, Approvals & Policy, Integrations and Platform. After support and on-call, each has about 11 squad-weeks of roadmap capacity a quarter. Squads can take work from another squad's area, but it takes them about 1.5× as long.
HiringTwo new squads were approved in September 2026. Tom's draft counts the first from Q1 2027 and the second from Q2 2027, both at full capacity. Last year, our three approved squads took 5, 6 and 8 months from approval to their first full sprint, and each delivered about half its capacity in its first quarter.
Card processor deadlineOur card processor retires its v1 API on 31 March 2027. Every card authorisation we process runs through v1 today. Migrating to v2 is about 18 squad-weeks of Cards work, and the processor must then run 4 weeks of certification testing before we can cut over. Certification is the processor's time, not ours, but nothing can change on our side while it runs.
Candidate work (estimates in squad-weeks, by owning squad)1. Processor v2 migration: Cards 18. Hard deadline above. 2. Virtual cards for software subscriptions: Cards 10. Needs v2. 3. Real-time spend controls (declines out-of-policy spend at the point of sale): Cards 8, Approvals & Policy 4. Needs v2. 4. Procurement module (purchase requests and approvals, a new paid add-on): Approvals & Policy 30, Integrations 6. 5. Multi-entity support (subsidiaries under one parent account): Platform 12, Approvals & Policy 10. 6. Sage Intacct integration: Integrations 9. 7. Microsoft Dynamics integration: Integrations 12. 8. SSO and SCIM provisioning: Platform 5. 9. Receipt-matching improvements: Expenses 6. 10. Mobile app rewrite: Expenses 22.
Tom's draft roadmapQ4 2026: Cards: v2 migration. Approvals & Policy: multi-entity. Integrations: Intacct. Platform: multi-entity. Expenses: receipt matching. Q1 2027: Cards: finish migration, start virtual cards. New squad A: procurement module. Integrations: Dynamics. Platform: SSO and SCIM. Q2 2027: Cards: virtual cards, spend controls. New squad A: procurement launch. New squad B: mobile rewrite. Q3 2027: Cards: spend controls. Everyone else: procurement adoption, mobile launch. Tom's note: “Procurement is our path to 112%. I've told the board it can add $6M of ARR in its first year.”
Procurement evidenceSix design partners have used a prototype since June. Two said they would pay for it; the other four already use a dedicated procurement tool and said they'd need a reason to switch. Proposed price: $8 per user per month, for users who raise purchase requests. At our customers, about 10% of employees raise purchase requests.
Card spend evidenceIn QBRs with our 60 largest customers, 26 asked for virtual cards for software subscriptions, and 31 CFOs asked for real-time spend controls. Finance estimates our customers pay about $410M a year of software subscriptions on other cards or by invoice (extrapolated from the 60 QBR accounts). Our net interchange is 1.1% of card spend.
Multi-entity evidence14 of our 60 largest customers have asked for it. Between them they have 52 subsidiaries not on Harbor, and 9 of the 14 have said in writing they would add their subsidiaries once it exists. Customer Success sized it at $2.9M ARR if all 52 joined at their parents' pricing.
Integrations evidence17% of customers use Sage Intacct through a CSV export. Their gross revenue churn is 13% a year, against 7% for customers on our NetSuite integration. Intacct was cited in 31% of lost new-logo deals last year, Dynamics in 8%. 3% of customers use Dynamics.
Other asksCRO: “Five enterprise deals worth $1.2M of pipeline need SSO and SCIM, and 11 existing accounts have it on their security review list.” CTO: “The mobile app is on a framework version that loses support in late 2027, and I want the rewrite done before then.” Support: receipt auto-match is at 71%, and unmatched receipts are the most common expense ticket.
What a strong answer doesThe answer key the graders mark against

A four-quarter roadmap that protects the 31 March deadline: the migration build has to finish by about early March to leave 4 weeks of certification, which is roughly all of Cards' capacity for Q4 and most of Q1 (18 of about 18 squad-weeks available before certification). So it adds help or buffer (Platform or another squad at 1.5×, or cutting everything else from Cards) and moves virtual cards out of Q1: both v2 features start only after cut-over (virtual cards Q2, spend controls Q2–Q3). It treats the new squads as upside: on past hiring they arrive around February to April 2027 and deliver half capacity in their first quarter, so nothing committed depends on them before Q2–Q3. It makes a clear call on procurement: the $6M claim needs about 890 accounts (81% of customers) paying $6,720 a year each (70 requesters × $8 × 12), and only 2 of 6 design partners would pay, so it replaces the 30-week build with a cheap paid test with a threshold, and keeps Approvals & Policy on multi-entity and spend controls, which have stronger expansion evidence ($2.9M sized, 9 of 14 committed in writing; about $1.5M a year of interchange from moving a third of the $410M of subscriptions onto virtual cards). Intacct comes early for its churn; Dynamics later or cut; SSO and SCIM are small and justified; the mobile rewrite waits for a new squad, before the framework's support ends. It lists what changed from Tom's draft, what's not being done, and the risks (certification slipping, hiring slipping), with commitment falling with distance.

Critical failures (cap the score at 40)
  • Leaves the processor migration finishing too late for 4 weeks of certification before 31 March 2027, or schedules virtual cards or real-time spend controls before it
Case

Staff-level · v1.1 · synthetic · B2B SaaS, fintech, capacity, hard deadline, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 12 checks passed

A senior PM could use this with light edits because the evidence, capacity math, deadline protection, and procurement call are all present. The only minor cleanup is tightening some later-quarter allocation language.

Output

Harbor FY27 roadmap: expand existing customers first

For the planning offsite | Q4 2026–Q3 2027

Recommendation: protect card continuity, unlock committed subsidiary expansion, then bring more existing-customer spend onto Harbor. Treat procurement as a commercial hypothesis to validate—not as committed revenue or a committed launch.

Board outcome: raise NRR from 104% to 112% by the end of Q3 2027. This roadmap supports that goal; the available evidence does not yet establish a credible revenue bridge all the way to 112%.

1. Proposed roadmap

Planning rules

  • Commit against the five existing squads, each with 11 squad-weeks per quarter. Neither new squad is on the critical path.
  • Numbers below are squad-weeks. Named feature estimates come from the pack; remaining capacity is an allocation ceiling for rollout, validation, hardening or contingency, not an estimate for additional features.
  • Cross-squad delivery uses the stated 1.5× cost.
  • Adoption work targets expansion and retention in existing accounts. It is not a placeholder for additional module builds.
SquadQ4 2026 — Protect and unblockQ1 2027 — Launch foundationsQ2 2027 — Capture spendQ3 2027 — Scale adoption
CardsProcessor v2: 11. Own architecture, integration and critical migration work. Outcome: eliminate the card-authorisation shutdown risk.Processor v2: remaining 3, plus 8 reserved for defects, freeze, certification and cutover contingency. Target certification 12 Feb–12 Mar, cutover 15 Mar, ahead of the 31 March retirement. No virtual-card launch commitment this quarter.Virtual cards: 10, plus 1 for integration support/contingency. Launch with customers identified through the QBRs. Outcome: move software subscription spend onto Harbor cards.Up to 11: virtual-card activation, spend conversion, controls integration support and reliability. Outcome: sustained incremental card spend before the NRR measurement date.
Approvals & PolicyMulti-entity: 10; procurement commercial validation: up to 1. Outcome: policy support for subsidiary expansion; early test of willingness to pay.Up to 7: multi-entity onboarding and policy fixes; up to 4: procurement validation. Outcome: convert written subsidiary commitments and decide whether procurement merits a funded build.Spend controls: 4; up to 7: multi-entity rollout and controls policy setup. Outcome: make controls usable by finance teams and expand subsidiary adoption.Up to 11: controls rollout, policy tuning and remaining subsidiary onboarding. Outcome: adoption that expands spend rather than creating excessive declines or customer friction.
IntegrationsSage Intacct: 9, plus 2 for pilot rollout and fixes. Outcome: remove a recurring accounting workflow gap in a substantial existing-customer segment.Up to 11: Intacct rollout, reliability and measuring CSV-to-integration conversion. Outcome: reduce retention risk and support expansion.Up to 11: Intacct and multi-entity accounting rollout; unused capacity remains reserve. Outcome: accounting readiness does not block subsidiary activation.Up to 11: integration adoption and reliability against remaining account-level expansion blockers. No Dynamics commitment.
PlatformMulti-entity: 11 of 12. Outcome: build the account structure needed for subsidiary expansion.Multi-entity: remaining 1; SSO/SCIM: 5; up to 5 for rollout and hardening. Launch multi-entity after the remaining platform work and acceptance testing. Outcome: activate subsidiaries and address existing-account security reviews, with enterprise pipeline as a secondary benefit.Up to 11: multi-entity and SSO/SCIM onboarding, operational hardening and contingency. Outcome: reliable expansion across larger customer organisations.Up to 11: scale and security work tied to demonstrated adoption issues. Outcome: retain and expand larger accounts without accumulating operational risk.
ExpensesProcessor support: 6, delivering 4 Cards-equivalent weeks; receipt matching: 5 of 6. Outcome: buy migration schedule margin while progressing the most common expense support issue.Receipt matching: remaining 1; up to 10 for rollout, quality measurement and fixes. Outcome: improve auto-match from the 71% baseline and reduce unmatched-receipt tickets.Spend-controls engineering: 9, delivering 6 of the 8 Cards-equivalent weeks; up to 2 for receipt quality. Cards retains technical ownership. Outcome: develop controls alongside virtual cards rather than queueing both behind one squad.Spend-controls engineering: 3, delivering the final 2 Cards-equivalent weeks, followed by an early-quarter launch subject to acceptance. Up to 8: controls hardening, receipt quality and mobile lifecycle assessment. Outcome: give controls time to drive adoption before quarter-end.

Critical-path checks

  • Processor migration: Q4 delivers 11 Cards weeks + 4 equivalent weeks from Expenses = 15. Q1 delivers the final 3, reaching the required 18. The February certification target creates margin before the hard retirement date.
  • Certification freeze: no changes to the certified implementation during the four-week processor window. Before committing dates, confirm the precise freeze scope, book the processor slot and agree cutover acceptance criteria.
  • Multi-entity: Approvals & Policy completes its 10 weeks in Q4; Platform completes 11 + 1. Therefore this is a Q1 launch, not a Q4 launch.
  • Spend controls: Expenses spends 12 actual weeks across Q2–Q3 to deliver 8 Cards-equivalent weeks; Approvals & Policy supplies its 4 weeks in Q2. Work begins only after v2 is live.
  • New squads: plan no committed output until teams are staffed and their ramp is observed. Use initial capacity for bounded adoption, testing or hardening work; allocate larger scope only after a delivery review.

Outcome scorecard

WorkstreamMeasure that matters
Overall expansionNRR and a cohort-based revenue bridge separating expansion, contraction and churn
Multi-entitySubsidiaries activated, incremental ARR and activation time; start with the nine parents that committed in writing
Virtual cardsIncremental software spend moved onto Harbor and resulting net interchange—not cards issued
Spend controlsEnabled accounts, incremental spend unlocked, decline quality and customer friction
IntacctEligible-account adoption, CSV retirement and subsequent retention/expansion versus comparable accounts
SSO/SCIMResolution of the 11 existing-account security reviews; pipeline conversion tracked separately
Receipt matchingMatch rate, incorrect matches and unmatched-receipt tickets per expense
Procurement discoveryPaid commitments at the proposed pricing, requester-seat counts and a credible reason to switch

CS, Sales and Finance should turn these into named-account activation targets and a revenue bridge at the offsite. We should not invent conversion forecasts from the current pack.

2. The case for the plan

Why this order

First, protect the revenue engine. Every card authorisation uses an API retiring on 31 March, and interchange represents 58% of revenue. Tom’s migration occupies 18 of Cards’ 22 available weeks across Q4 and Q1, but must also leave four calendar weeks for certification. Quarterly capacity alone does not prove that schedule is safe.

Borrowing six Expenses weeks moves four Cards-equivalent weeks into Q4. That leaves only three planned migration weeks in Q1 and creates room for defects and certification. Receipt matching slips slightly; card continuity takes precedence.

Second, prioritise expansion with identifiable buyers. Multi-entity has the strongest evidence of a concrete expansion action: nine parents have committed in writing. The full 52-subsidiary opportunity is $2.9M ARR, but we cannot treat that total as committed or infer how many subsidiaries the nine parents represent. Launch in Q1, then sell and onboard—not merely ship.

Third, pursue spend expansion in parallel. Virtual cards have a direct connection to software spend currently outside Harbor. Controls address the most frequently expressed card need in the QBRs and may make finance teams comfortable moving more spend onto Harbor. Borrowing Expenses capacity costs four additional squad-weeks versus specialist delivery, but avoids serialising both products through Cards and supports an early-Q3 controls launch.

The $410M software-spend estimate implies a $4.51M annual net-interchange ceiling at 1.1%, before accounting for capture rates, timing or spend that cannot move from invoice to card. It is an opportunity, not a forecast. Both features may affect the same spend; we will not double-count it.

Fourth, remove retention and expansion friction. Intacct serves 17% of existing customers, versus 3% for Dynamics. The 13% versus 7% churn difference is an association, not proof the integration causes better retention, but it supports prioritising Intacct. SSO/SCIM is relatively small and addresses 11 existing-account security reviews. Receipt matching targets the largest expense support problem.

What changes from Tom’s draft

  • Hiring no longer funds commitments. Historical time to the first full sprint was five to eight months, with half-capacity delivery in the first quarter. September approvals do not justify a full squad in Q1 and another in Q2. Further, new squads taking specialist work incur the stated transfer cost.
  • Migration gains early borrowed capacity and an explicit certification window. Virtual cards no longer competes with the deadline in Q1.
  • Multi-entity launches in Q1. Platform’s 12-week estimate exceeds one quarter’s 11-week capacity.
  • Procurement moves from committed launch to commercial validation. Its revenue claim is not supported by current evidence.
  • Controls gets borrowed delivery capacity. This creates an earlier adoption window without relying on hiring.
  • Dynamics and the mobile rewrite leave the committed horizon. The freed capacity supports existing-customer outcomes and schedule resilience.

Why procurement is not the anchor

At the proposed price, an average customer has 70 requester seats and contributes:

70 × $8 × 12 = $6,720 ARR.

Reaching $6M ARR requires approximately 893 average-sized customers—81% of today’s customer base. Only two of six design partners said they would pay; neither is described as a signed commitment. Four already have a competing tool.

The estimated build is 30 Approvals & Policy weeks plus six Integrations weeks, before any additional cross-squad cost. That is too large a bet to anchor this year’s NRR target without stronger evidence.

Validation should seek paid commitments, actual seat counts and evidence of a differentiated switching reason. It is not an unbudgeted production build. A successful gate triggers an explicit investment decision, identifying displaced work and a capacity-backed launch date.

What we are not doing

We are not committing to procurement launch or its $6M claim, Dynamics, or a 22-week mobile rewrite in this horizon. We are also not inventing new feature scope to consume every available week.

The mobile support deadline remains real. By Q3, the CTO must establish the exact end-of-support date and a funded maintenance, upgrade or rewrite path. If continued safe operation requires starting earlier, it returns as a mandatory trade-off—not an assumed new-squad assignment.

Risks that could change the plan

The largest execution risks are processor certification availability, migration work that cannot be delegated, and underestimated multi-entity or controls complexity. Validate the processor work split immediately; any threat to certification displaces discretionary work first.

The largest commercial risk is insufficient expansion capture. At an illustrative $86M opening cohort, an eight-point NRR improvement represents $6.88M more retained/expanded ARR than a 104% outcome. Finance must reconcile the actual cohort and baseline. Even the full multi-entity and software-card opportunity totals only $7.41M before timing, conversion and overlap effects. This plan does not justify promising 112%.

Review the revenue bridge after Q1 launches. Weak activation should shift capacity toward removing proven adoption blockers. Strong paid procurement evidence, an earlier mobile security deadline, or demonstrated new-squad capacity can justify a replan—but none should silently become an assumption today.

Grades and run details

Decision model 83 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly35%
  • passAddresses the actual decision99%
  • partialRespects explicit constraints26%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims58%
  • failProduces the required deliverable52%
  • passProtects the processor deadline94%
  • passPlans on the squads we actually have99%
  • partialMakes the call on procurement37%
  • passOutcomes, with certainty that falls with distance92%
  • passFits the stated capacity54%
  • passSequences around dependencies56%
Run
Run
#1
Time to output
2.0 min
Submitted
1 Oct 2026

Needs more than a quick edit8 of 12 checks passed

A PM would need to add real buffer or help for the v2 migration and define a procurement validation threshold before taking this to the exec team.

Output

Harbor FY27 roadmap — Q4 2026 to Q3 2027

Planning basis: Each existing squad has about 11 roadmap squad-weeks per quarter. Figures below are the candidate-work estimates; capacity not assigned to a feature is deliberately held for launch, customer activation, fixes and contingency. New squads are not counted at full capacity until their ramp is proven.

QuarterSquadWork and outcome
Q4 2026CardsProcessor v2 migration — 11 of 18 weeks. Start the mandatory migration and establish the path to retire v1 before the processor deadline.
Approvals & PolicyMulti-entity — 10 of 10 weeks. Build parent/subsidiary approval and policy support so existing customers can bring more entities onto Harbor.
PlatformMulti-entity — 11 of 12 weeks. Build the platform foundation for subsidiaries under a parent account.
IntegrationsSage Intacct — 9 weeks. Replace CSV workarounds to reduce a retention risk among existing customers.
ExpensesReceipt matching — 6 weeks. Improve the most common expense-ticket driver; use remaining capacity for fixes and customer feedback.
Q1 2027CardsFinish v2 — 7 weeks. Target implementation complete by the end of February, leaving March for the processor’s four-week certification. No card-side changes during certification.
Approvals & PolicyMulti-entity launch and customer activation. Support the first subsidiary rollouts. Use remaining capacity for a time-bounded procurement willingness-to-pay and switching validation; do not build the module.
PlatformFinish multi-entity — 1 week; SSO/SCIM — 5 weeks. Complete the entity foundation, then address security blockers affecting existing accounts and enterprise pipeline.
IntegrationsMicrosoft Dynamics — 11 of 12 weeks. Address a cited reason for lost new-logo deals; complete the final week in Q2.
ExpensesMobile rewrite — 11 of 22 weeks. Start early to reduce the risk of running out of time before the current framework loses support in late 2027.
Q2 2027CardsVirtual cards for software subscriptions — 10 weeks. After v2 certification, give customers a way to move subscription spend onto Harbor cards and capture more interchange.
Approvals & PolicyReal-time spend controls — 4 weeks. Build the policy logic for point-of-sale declines, in step with Cards’ Q3 work.
PlatformSSO/SCIM and multi-entity activation. Support adoption and resolve launch issues; no new major platform feature is precommitted.
IntegrationsFinish Dynamics — 1 week; customer rollout. Complete the integration and support implementation.
ExpensesFinish mobile rewrite — 11 weeks. Complete the rewrite by Q2, ahead of the framework support deadline.
Q3 2027CardsReal-time spend controls — 8 weeks. Complete and release with Approvals & Policy; help finance teams enforce policy at the point of sale.
Approvals & PolicyRelease and adoption for real-time controls; multi-entity activation. Focus on turning shipped capabilities into customer usage and expansion.
PlatformActivation and reliability for multi-entity and SSO/SCIM. Prioritize customer rollout and fixes over an unvalidated new feature.
IntegrationsIntacct and Dynamics adoption. Help customers get live and measure retention and deal impact before choosing further integration work.
ExpensesMobile rollout and stabilization; measure receipt-match performance. Address launch issues and confirm the support-ticket outcome.

New Squads A and B: Plan for onboarding and ramp, not full feature capacity. When productive, use them for bounded rollout, quality or activation work. No committed launch depends on their availability.

Case for the plan

Why this order

The board’s priority is to move NRR from 104% to 112% by the end of Q3, primarily through expansion in existing customers. The plan therefore favors credible opportunities to retain customers or deepen their use of Harbor, while keeping some room for new-logo needs.

First, v2 is a hard dependency and a hard deadline. All authorisations still run through v1, which retires on 31 March. The 18 Cards squad-weeks consume 11 weeks in Q4 and seven in Q1. That leaves time to finish implementation before the processor’s four-week certification and to keep the required freeze during testing. Virtual cards and real-time controls must wait until certification is complete.

Next, multi-entity has concrete expansion evidence: 14 of the 60 largest customers asked, and nine have said in writing that they would add subsidiaries. The $2.9M ARR estimate assumes all 52 subsidiaries join, so it is an upside case, not a commitment. Completing the Platform and Approvals & Policy work across Q4 and Q1 creates a path to test that demand with real rollouts.

Virtual cards are the first post-migration Cards feature because 26 large customers asked for them and Finance estimates $410M of annual subscription spend is currently elsewhere. At Harbor’s 1.1% net interchange, full capture would represent about $4.5M in annual interchange revenue, but that is a ceiling, not a forecast: the spend estimate is extrapolated, and adoption and capture are unknown. Real-time controls follow, with coordinated work across Cards and Approvals & Policy. The 31 CFO requests are strong demand evidence, but the plan avoids trying to fit both features into Cards’ 11-week quarterly capacity.

Intacct comes before Dynamics. Intacct is used by 17% of customers through CSV, and those customers have higher observed gross revenue churn than customers on NetSuite. That comparison is correlational, but it supports addressing a retention risk. Dynamics has lower existing use (3%) and a more limited new-logo case, but the 8% of lost deals citing it justifies completing the 12-week integration after Intacct.

SSO/SCIM is a relatively small, five-week Platform investment tied to five enterprise deals representing $1.2M of pipeline and 11 existing security reviews. It supports commercial outcomes without displacing the larger expansion bets.

What changed from Tom’s draft

  • Removed procurement from committed build and launch. Six design partners are not enough evidence for the board’s $6M first-year ARR claim: only two said they would pay, and four already use another tool. At the proposed price, an average 700-employee customer would generate about $6,720 ARR if 10% of employees were buyers. Reaching $6M would require roughly 893 such customer equivalents. We should validate willingness to switch and pricing before committing 30 Approvals & Policy and six Integrations squad-weeks.
  • Corrected the hiring assumption. Prior squads took five to eight months from approval to a first full sprint and delivered about half capacity in their first quarter. The plan does not count either new squad at full capacity in Q1 or Q2, or make a feature dependent on them.
  • Made the v2 critical path explicit. The migration spans Q4 and Q1, with implementation targeted ahead of March certification. Tom’s draft did not protect that freeze clearly enough.
  • Sequenced Cards work to fit capacity. Virtual cards and real-time controls require 18 Cards squad-weeks together, more than one quarter’s 11-week capacity. They are sequenced across Q2 and Q3.
  • Started the mobile rewrite with the existing Expenses squad. The work is 22 weeks, so starting in Q1 targets completion in Q2 without relying on an unproven new squad.

What we are not doing

We are not committing to build procurement, starting a mobile rewrite with a new squad at assumed full capacity, or taking on another major integration or platform feature in Q3 without evidence from the launches above. New-logo work remains in the plan, but expansion and retention come first.

Risks that could change the plan

  • Processor timing: Any migration slip puts card authorisations at risk. Track delivery and certification weekly; protect the Q1 card schedule and do not pull dependent features forward.
  • Expansion may not materialize: Multi-entity interest and card requests do not guarantee adoption or spend migration. Track subsidiary activation, card spend moved, and controls usage—not just releases.
  • Commercial estimates are uncertain: The $2.9M multi-entity estimate and $4.5M card-spend ceiling are not forecasts. Procurement’s $6M claim is especially unsupported by current evidence.
  • Mobile has little slack: The 22-week rewrite fills two quarters of Expenses capacity. Scope growth, delivery delays or an earlier-than-expected framework deadline would require replanning.
  • Hiring, integration and cross-squad delivery may slip: Ramp timing is uncertain; the Intacct churn difference may not be caused by the integration; and spend controls depend on coordinated Cards and Approvals & Policy work.

This roadmap gives Harbor credible bets aligned to the NRR goal, but it does not establish that the 112% target is guaranteed. We should review expansion and retention indicators at each quarterly planning checkpoint and change the plan if customer evidence or delivery risk materially shifts.

Grades and run details

Decision model 79 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly56%
  • passAddresses the actual decision96%
  • partialRespects explicit constraints16%
  • passIdentifies material uncertainty89%
  • passAvoids unsupported claims64%
  • failProduces the required deliverable9%
  • partialProtects the processor deadline35%
  • passPlans on the squads we actually have98%
  • partialMakes the call on procurement84%
  • passOutcomes, with certainty that falls with distance72%
  • passFits the stated capacity62%
  • passSequences around dependencies54%
Run
Run
#1
API response time
1.9 min
Submitted
1 Oct 2026

Usable with a quick edit10 of 12 checks passed

The output is well-structured, evidence-based, and addresses all required elements; a PM could use it with minor wording edits without re-checking evidence or reasoning.

Output

Part 1: Roadmap, Q4 2026 to Q3 2027

Planning basis: Five existing squads at 11 squad-weeks a quarter (55 a quarter). The two new squads are counted as zero in the commitments (see the case). Lending between squads is costed at 1.5×. Figures in brackets are squad-weeks of the named squad.

Q4 2026: Protect the card rail, start entities

SquadWorkOutcome served
Cardsv2 migration (11). Ring-fenced, nothing else.Protect continuity of card processing (~$50M interchange)
IntegrationsLend to v2 migration (6, ≈4 Cards-weeks). Intacct (5 of 9).Protect the deadline; cut churn in Intacct customers
PlatformMulti-entity (11 of 12)More entities
Approvals & PolicyMulti-entity (10). Procurement pricing test and paid-pilot outreach (1).More entities; more modules
ExpensesReceipt matching (6). Mobile options and architecture, including an in-place upgrade check (5).Fewer support tickets; keep the mobile app supportable

Q1 2027: Cut over, ship entities, start controls

SquadWorkOutcome served
CardsFinish v2 (3), code complete by late January. Certification and change freeze in February. Cutover by end of February. Real-time spend controls build (8), released after cutover.Protect processing; more spend on Harbor cards
Approvals & PolicyReal-time spend controls policy side (4). Procurement build (7). Gate: procurement go/no-go at end of Q1.Card spend; more modules
PlatformFinish multi-entity (1), GA in January. SSO and SCIM (5). Reserve (5).More entities; retention and new-logo pipeline
IntegrationsFinish Intacct (4), GA. Procurement ERP sync (3). Reserve (4).Lower churn; more modules
ExpensesLend to procurement (5). Mobile rewrite (6).More modules; app supportability

Q2 2027: Virtual cards, procurement beta

SquadWorkOutcome served
CardsVirtual cards for software subscriptions (10). Reserve (1).More spend on Harbor cards
Approvals & PolicyProcurement (11), beta with design partners by end of Q2More modules
PlatformLend to procurement (9, ≈6 A&P-weeks). Reserve (2).More modules
IntegrationsProcurement ERP sync (3). Dynamics starts only if capacity remains (stretch).More modules
ExpensesLend to procurement (4). Mobile rewrite (7).More modules; app supportability
New squad A (if in seat)Bounded procurement items, ≈5 weeks at half capacity. Upside only.

Q3 2027: Launch and land

SquadWorkOutcome served
Approvals & PolicyProcurement GA and hardening (11), target JulyMore modules
ExpensesFinish and ship mobile (4). Reserve (7).App supported before the framework deadline
Cards, Platform, IntegrationsUncommitted. Post-launch iteration on controls and virtual cards, multi-entity rollout follow-ups, Dynamics. Allocated at the Q2 review.
New squadsSquad A: procurement. Squad B: Dynamics or the first FY28 area. Charters set at the Q2 review.

About 60 of 220 squad-weeks are deliberately uncommitted, mostly in Q3.

---

Part 2: The case (≈1,150 words)

The short version

The order is set by one thing we can't negotiate and by the evidence for everything else. The processor retires v1 on 31 March, every authorisation runs through it, and interchange is 58% of revenue. Cards does nothing else until cutover. After that, work is ordered by evidence and cost per squad-week.

The plan gets our exit run-rate to roughly 109–111% NRR, not reliably 112%. We should tell the board that now rather than in Q3.

Why this order

1. v2 first, with real slack. Tom's draft has Cards doing 18 weeks at 11 a quarter, so build finishes around late February. Four weeks of certification then ends around 25 March, about a week before a deadline that would stop all card authorisation. My plan ring-fences Cards in Q4 and has Integrations, which does API work daily, lend 6 weeks (≈4 Cards-weeks). Code is complete in late January, certification runs in February, and we keep about five weeks of slack.

2. Multi-entity next. It has the best evidence we hold. Nine of 14 requesting customers said in writing they'd add subsidiaries, which means ~33 of the 52 subsidiaries, or ~$1.9M ARR. Customer Success's $2.9M assumes all 52 join. The cost is 22 weeks, and it is a direct "more entities" lever.

3. Intacct, then SSO. 17% of customers (~187) use Intacct via CSV and churn at 13% against 7% on NetSuite. If the integration closes that gap, it retains ~$0.9M a year for 9 weeks of work. That assumes average ARR of $78k and that the gap is causal, which is unproven. Intacct also appears in 31% of lost deals. SSO and SCIM costs 5 weeks and covers $1.2M of pipeline plus 11 existing accounts' security reviews. Dynamics is 12 weeks for 3% of customers and 8% of lost deals, and it serves new-logo growth, so it is stretch.

4. Controls before virtual cards. Both need v2 live, so both follow cutover. 31 CFOs asked for real-time controls against 26 for virtual cards. The 4 Approvals & Policy weeks for controls fit in Q1, before procurement takes that squad's capacity. Virtual cards is the only card item we can size: $410M × 10–25% capture × 1.1% ≈ $0.5–1.1M. The $410M is extrapolated from our 60 largest accounts, so it is probably high. This is a judgment call. Swapping the order is cheap, but the Approvals & Policy slot would collide with the procurement build.

5. Procurement: build it, resize the promise. At $8 × 10% of a 700-person company's employees, a customer pays $6,720 a year. Even 100% attach across 1,100 customers is $7.4M, so $6M needs ~890 customers (81%). Two of six design partners would pay (33%), and four already own a procurement tool. At 10–33% attach, year-one ARR is $0.7–2.5M, and with GA in Q3 only a small part lands before Q3 ends. It is still worth building: it is our only new paid module, it recurs, and it matches the board's "more modules" lever. But it is not a path to 112% on its own.

The arithmetic: 8 points of NRR on $86M is ~$6.9M. Sized items total roughly $4–6M of exit run-rate: entities $1.9M, Intacct $0.9M, virtual cards $0.5–1.1M, and procurement $0.7–2.5M (mostly after Q3). Controls and SSO retention are unsized upside. Recognised NRR at Q3 will lag run-rate because late launches haven't ramped.

What I changed from Tom's draft

  • New squads count as zero. Last year's three squads took 5, 6 and 8 months from approval to first full sprint, then delivered half capacity in their first quarter. From a September approval, squad A is likely in seat between February and May and squad B later. A realistic ~5 weeks in Q2 and ~11 in Q3 for A, plus ~5 in Q3 for B, is about 22 weeks, against 55 in the draft. Even at Tom's capacity, procurement (36 weeks) couldn't launch in Q2, because squad A would have 22 weeks by the end of Q2.
  • Procurement moves to Approvals & Policy. They own the domain and, in the draft, have no work after Q4. Expenses and Platform lend 9 weeks each (≈6 A&P-weeks each) to get a beta by the end of Q2 and GA in Q3. New squads join as upside.
  • Mobile moves to Expenses. The draft puts it on squad B, which may not exist until Q3. Expenses has ~38 uncommitted weeks, and the rewrite ships in Q3, ahead of the late-2027 framework cutoff.
  • Controls before virtual cards (point 4 above).
  • Intacct GA moves from Q4 to Q1 to fund the v2 help. The cost is roughly one quarter of a ~$0.9M-a-year benefit, about $0.2M.
  • Q3's "everyone else: adoption" is replaced by a held reserve, because adoption isn't squad work and Q3 launches can't move FY27 NRR much.
  • A procurement checkpoint is added at the end of Q1 (see risks).

What we're not doing

  • Dynamics, beyond a stretch start in Q2–Q3.
  • A Q2 procurement launch, or putting the $6M figure in front of the board.
  • Any new Cards scope before cutover.
  • A mobile rewrite staffed by a squad that doesn't exist yet.
  • Planning on capacity from squads we haven't hired.

Risks that could change the plan

  1. Processor timeline. If v2 isn't code-complete by 31 January, Platform joins Cards and controls and virtual cards slip a quarter. Two asks for this month: a written fallback or extension from the processor, and confirmation that feature work behind flags is allowed during the certification freeze (my assumption). Book the certification slot now.
  2. Procurement demand. At the end-of-Q1 gate I propose requiring at least 10 signed paid commitments at $8 from ~30 accounts approached (≈$67k ARR). If we miss it, we cut scope and redirect Approvals & Policy and the lenders to multi-entity rollout and card-spend work. The threshold is my proposal and is open to debate.
  3. The NRR target itself. Exec decision: hold 112% as a Q3 recognised-NRR target, or restate it as an exit run-rate target, which I think is achievable only with upside from controls and procurement.
  4. Evidence quality. The $410M is extrapolated from our largest accounts. The Intacct churn gap is a correlation. The multi-entity sizing rests on 14 of 60 customers. Any could be off by a factor of two.
  5. Lending efficiency. The 1.5× penalty is an average. Expenses and Platform have less procurement context, and if they're slower than assumed, GA slips into Q4.
  6. Hiring. If squad A lands in February, procurement GA can move earlier. If squad B lands late, nothing in the plan changes.
  7. Mobile date. The CTO should confirm the exact end-of-support date and whether an in-place upgrade avoids a 22-week rewrite. Q4's scoping work answers this.

Grades and run details

Decision model 79 · LLM judge 13 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly14%
  • passAddresses the actual decision93%
  • passRespects explicit constraints30%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims25%
  • failProduces the required deliverable8%
  • passProtects the processor deadline75%
  • passPlans on the squads we actually have99%
  • passMakes the call on procurement71%
  • passOutcomes, with certainty that falls with distance93%
  • passFits the stated capacity44%
  • passSequences around dependencies51%
Run
Run
#1
API response time
3.4 min
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyRightRightMixed
GPT-6 Astra · ChatGPT

The output’s factual statements and arithmetic are drawn from the supplied context, and it labels uncertain causal and revenue claims appropriately.

GPT-6 Luna · API

The output's factual statements about Harbor, evidence, estimates, and arithmetic are supported by the supplied context.

Sonnet 5.5 · API

Every factual claim about the current situation is taken directly from the brief or supplied context, with no invented facts.

Respects explicit constraintsRightWrongRight
GPT-6 Astra · ChatGPT

It provides a quarter-by-quarter roadmap and a case under 1,200 words, respects the 11 squad-week capacity, 1.5× cross-squad cost, and processor certification constraint, and does not rely on unhired squads.

GPT-6 Luna · API

It does not enforce the processor-deadline constraint with real buffer or help, and the procurement validation lacks a threshold that would justify the full build.

Sonnet 5.5 · API

The roadmap is by quarter with squads and outcomes, the case is under 1,200 words, and all requested elements (order, changes, not doing, risks) are included.

Protects the processor deadlineRightWrongRight
GPT-6 Astra · ChatGPT

It finishes the migration with a February-to-March certification window before 31 March, adds buffer and borrowed capacity, and starts virtual cards and spend controls only after cut-over.

GPT-6 Luna · API

It finishes the migration with no Cards slack or added help, relying on an exact end-February target rather than protecting the four-week certification window.

Sonnet 5.5 · API

The migration finishes code-complete by late January, certification runs in February, cutover by end of February, with slack before 31 March; virtual cards and spend controls start only after cutover.

Makes the call on procurementRightWrongRight
GPT-6 Astra · ChatGPT

It checks Tom’s $6M claim against pricing and customer-base arithmetic, weighs the weak design-partner evidence, commits Approvals & Policy to better-evidenced work, and proposes a paid validation gate.

GPT-6 Luna · API

It correctly challenges the $6M claim but proposes only time-bounded validation without a threshold that would justify the full procurement build.

Sonnet 5.5 · API

The output checks the $6M claim against pricing and adoption evidence, shows it implies ~81% attach, weighs it against 2 of 6 design partners, commits Approvals & Policy to better-evidenced work, and proposes a paid test with a threshold.

All mixed 1

Produces the required deliverableMixedMixedMixed
GPT-6 Astra · ChatGPT

The roadmap and case are present, executive-readable, and actionable with only light edits.

GPT-6 Luna · API

It provides both a quarter-by-quarter roadmap with outcomes and a case under 1,200 words for the exec team.

Sonnet 5.5 · API

The roadmap and case are complete, in the requested form for the exec team, within the word limit, and usable with light edits.

All got right 7

Addresses the actual decisionRightRightRight
GPT-6 Astra · ChatGPT

It commits early to protecting card continuity, launching multi-entity, validating procurement rather than building it, and deferring Dynamics and mobile, with clear replan triggers.

GPT-6 Luna · API

It commits to a clear four-quarter roadmap and procurement call, and says quarterly evidence checkpoints and risks could change the plan.

Sonnet 5.5 · API

The output commits to a clear roadmap and case, makes an explicit call on procurement with a gate, and states what would change the plan.

Identifies material uncertaintyRightRightRight
GPT-6 Astra · ChatGPT

It names certification, hiring, migration delegation, multi-entity/controls complexity, procurement willingness to pay, and revenue-bridge uncertainty, with resolution or replan conditions.

GPT-6 Luna · API

It names key unknowns such as certification slip, hiring ramp, adoption, and correlational churn evidence, and says they would be reviewed at quarterly checkpoints.

Sonnet 5.5 · API

The output names specific unknowns (processor timeline, procurement demand, NRR target, evidence quality, lending efficiency, hiring, mobile date) and says how they would be resolved or change the plan.

Avoids unsupported claimsRightRightRight
GPT-6 Astra · ChatGPT

It avoids presenting procurement revenue, integration causation, or the 112% bridge as established facts and labels them as hypotheses or ceilings.

GPT-6 Luna · API

It labels the $4.5M interchange figure as a ceiling, the churn comparison as correlational, and the $2.9M multi-entity figure as upside rather than forecast.

Sonnet 5.5 · API

Interpretations and forecasts are clearly labelled as such (e.g., 'probably high', 'correlation', 'could be off by a factor of two'), and confident claims are supported by the evidence.

Plans on the squads we actually haveRightRightRight
GPT-6 Astra · ChatGPT

It explicitly excludes the two new squads from committed critical-path work and treats their capacity as upside after observed ramp.

GPT-6 Luna · API

It uses the hiring history to avoid counting new squads at full capacity and commits no important launch to them.

Sonnet 5.5 · API

The new squads are treated as zero in commitments, with realistic arrival and half capacity based on past hiring, and nothing important depends on them before Q2-Q3.

Outcomes, with certainty that falls with distanceRightRightRight
GPT-6 Astra · ChatGPT

Each workstream has an outcome, near-term items are specific, and later-quarter scope is deliberately looser and tied to adoption or contingency.

GPT-6 Luna · API

Each item names an outcome or problem, and later-quarter items are deliberately looser than near-term commitments.

Sonnet 5.5 · API

Every roadmap item names its outcome or problem, near-term items are specific with dates, and later items are deliberately looser (e.g., Q3 uncommitted, allocated at Q2 review).

Fits the stated capacityRightRightRight
GPT-6 Astra · ChatGPT

The committed squad-week sums fit the stated 11-week quarterly capacity, use the 1.5× cross-squad cost, and identify deferred work such as Dynamics and mobile.

GPT-6 Luna · API

The listed squad-weeks fit the stated 11-week quarterly capacity, and the output names deferred or non-committed work.

Sonnet 5.5 · API

The committed work sums to 55 squad-weeks per quarter with slack, the sums are checkable, and it names what was cut or deferred (Dynamics, Q2 procurement launch, etc.).

Sequences around dependenciesRightRightRight
GPT-6 Astra · ChatGPT

It sequences virtual cards and spend controls after v2, multi-entity after Platform and Approvals work, and procurement build after validation, naming the key dependencies.

GPT-6 Luna · API

It places virtual cards and spend controls after v2 certification and names the key deadline, hiring, and cross-squad dependencies.

Sonnet 5.5 · API

All dependencies are respected (v2 before virtual cards and controls, hiring before new squad work, procurement gate), and the ones driving the order are named.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 89% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI89.6100.02None
2GPT-6 AstrawithChatGPT91.787.52None
3GPT-6.1 SolwithAPI91.784.62None
4GPT-6 LunawithAPI85.076.32None
5Opus 5.5withClaude80.565.12None
6Gemini 3.8 FlashwithAPI70.151.92None
7Gemini 3.5 Flash-LitewithGemini28.48.322 capped

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review