Usable with a quick edit11 of 12 checks passed
The output is well-structured, evidence-based, and addresses all required elements; a PM could use it with minor wording edits without re-checking evidence or reasoning.
Output
Part 1: Roadmap, Q4 2026 to Q3 2027
Planning basis: Five existing squads at 11 squad-weeks a quarter (55 a quarter). The two new squads are counted as zero in the commitments (see the case). Lending between squads is costed at 1.5×. Figures in brackets are squad-weeks of the named squad.
Q4 2026: Protect the card rail, start entities
| Squad | Work | Outcome served |
|---|---|---|
| Cards | v2 migration (11). Ring-fenced, nothing else. | Protect continuity of card processing (~$50M interchange) |
| Integrations | Lend to v2 migration (6, ≈4 Cards-weeks). Intacct (5 of 9). | Protect the deadline; cut churn in Intacct customers |
| Platform | Multi-entity (11 of 12) | More entities |
| Approvals & Policy | Multi-entity (10). Procurement pricing test and paid-pilot outreach (1). | More entities; more modules |
| Expenses | Receipt matching (6). Mobile options and architecture, including an in-place upgrade check (5). | Fewer support tickets; keep the mobile app supportable |
Q1 2027: Cut over, ship entities, start controls
| Squad | Work | Outcome served |
|---|---|---|
| Cards | Finish v2 (3), code complete by late January. Certification and change freeze in February. Cutover by end of February. Real-time spend controls build (8), released after cutover. | Protect processing; more spend on Harbor cards |
| Approvals & Policy | Real-time spend controls policy side (4). Procurement build (7). Gate: procurement go/no-go at end of Q1. | Card spend; more modules |
| Platform | Finish multi-entity (1), GA in January. SSO and SCIM (5). Reserve (5). | More entities; retention and new-logo pipeline |
| Integrations | Finish Intacct (4), GA. Procurement ERP sync (3). Reserve (4). | Lower churn; more modules |
| Expenses | Lend to procurement (5). Mobile rewrite (6). | More modules; app supportability |
Q2 2027: Virtual cards, procurement beta
| Squad | Work | Outcome served |
|---|---|---|
| Cards | Virtual cards for software subscriptions (10). Reserve (1). | More spend on Harbor cards |
| Approvals & Policy | Procurement (11), beta with design partners by end of Q2 | More modules |
| Platform | Lend to procurement (9, ≈6 A&P-weeks). Reserve (2). | More modules |
| Integrations | Procurement ERP sync (3). Dynamics starts only if capacity remains (stretch). | More modules |
| Expenses | Lend to procurement (4). Mobile rewrite (7). | More modules; app supportability |
| New squad A (if in seat) | Bounded procurement items, ≈5 weeks at half capacity. Upside only. |
Q3 2027: Launch and land
| Squad | Work | Outcome served |
|---|---|---|
| Approvals & Policy | Procurement GA and hardening (11), target July | More modules |
| Expenses | Finish and ship mobile (4). Reserve (7). | App supported before the framework deadline |
| Cards, Platform, Integrations | Uncommitted. Post-launch iteration on controls and virtual cards, multi-entity rollout follow-ups, Dynamics. Allocated at the Q2 review. | |
| New squads | Squad A: procurement. Squad B: Dynamics or the first FY28 area. Charters set at the Q2 review. |
About 60 of 220 squad-weeks are deliberately uncommitted, mostly in Q3.
---
Part 2: The case (≈1,150 words)
The short version
The order is set by one thing we can't negotiate and by the evidence for everything else. The processor retires v1 on 31 March, every authorisation runs through it, and interchange is 58% of revenue. Cards does nothing else until cutover. After that, work is ordered by evidence and cost per squad-week.
The plan gets our exit run-rate to roughly 109–111% NRR, not reliably 112%. We should tell the board that now rather than in Q3.
Why this order
1. v2 first, with real slack. Tom's draft has Cards doing 18 weeks at 11 a quarter, so build finishes around late February. Four weeks of certification then ends around 25 March, about a week before a deadline that would stop all card authorisation. My plan ring-fences Cards in Q4 and has Integrations, which does API work daily, lend 6 weeks (≈4 Cards-weeks). Code is complete in late January, certification runs in February, and we keep about five weeks of slack.
2. Multi-entity next. It has the best evidence we hold. Nine of 14 requesting customers said in writing they'd add subsidiaries, which means ~33 of the 52 subsidiaries, or ~$1.9M ARR. Customer Success's $2.9M assumes all 52 join. The cost is 22 weeks, and it is a direct "more entities" lever.
3. Intacct, then SSO. 17% of customers (~187) use Intacct via CSV and churn at 13% against 7% on NetSuite. If the integration closes that gap, it retains ~$0.9M a year for 9 weeks of work. That assumes average ARR of $78k and that the gap is causal, which is unproven. Intacct also appears in 31% of lost deals. SSO and SCIM costs 5 weeks and covers $1.2M of pipeline plus 11 existing accounts' security reviews. Dynamics is 12 weeks for 3% of customers and 8% of lost deals, and it serves new-logo growth, so it is stretch.
4. Controls before virtual cards. Both need v2 live, so both follow cutover. 31 CFOs asked for real-time controls against 26 for virtual cards. The 4 Approvals & Policy weeks for controls fit in Q1, before procurement takes that squad's capacity. Virtual cards is the only card item we can size: $410M × 10–25% capture × 1.1% ≈ $0.5–1.1M. The $410M is extrapolated from our 60 largest accounts, so it is probably high. This is a judgment call. Swapping the order is cheap, but the Approvals & Policy slot would collide with the procurement build.
5. Procurement: build it, resize the promise. At $8 × 10% of a 700-person company's employees, a customer pays $6,720 a year. Even 100% attach across 1,100 customers is $7.4M, so $6M needs ~890 customers (81%). Two of six design partners would pay (33%), and four already own a procurement tool. At 10–33% attach, year-one ARR is $0.7–2.5M, and with GA in Q3 only a small part lands before Q3 ends. It is still worth building: it is our only new paid module, it recurs, and it matches the board's "more modules" lever. But it is not a path to 112% on its own.
The arithmetic: 8 points of NRR on $86M is ~$6.9M. Sized items total roughly $4–6M of exit run-rate: entities $1.9M, Intacct $0.9M, virtual cards $0.5–1.1M, and procurement $0.7–2.5M (mostly after Q3). Controls and SSO retention are unsized upside. Recognised NRR at Q3 will lag run-rate because late launches haven't ramped.
What I changed from Tom's draft
- New squads count as zero. Last year's three squads took 5, 6 and 8 months from approval to first full sprint, then delivered half capacity in their first quarter. From a September approval, squad A is likely in seat between February and May and squad B later. A realistic ~5 weeks in Q2 and ~11 in Q3 for A, plus ~5 in Q3 for B, is about 22 weeks, against 55 in the draft. Even at Tom's capacity, procurement (36 weeks) couldn't launch in Q2, because squad A would have 22 weeks by the end of Q2.
- Procurement moves to Approvals & Policy. They own the domain and, in the draft, have no work after Q4. Expenses and Platform lend 9 weeks each (≈6 A&P-weeks each) to get a beta by the end of Q2 and GA in Q3. New squads join as upside.
- Mobile moves to Expenses. The draft puts it on squad B, which may not exist until Q3. Expenses has ~38 uncommitted weeks, and the rewrite ships in Q3, ahead of the late-2027 framework cutoff.
- Controls before virtual cards (point 4 above).
- Intacct GA moves from Q4 to Q1 to fund the v2 help. The cost is roughly one quarter of a ~$0.9M-a-year benefit, about $0.2M.
- Q3's "everyone else: adoption" is replaced by a held reserve, because adoption isn't squad work and Q3 launches can't move FY27 NRR much.
- A procurement checkpoint is added at the end of Q1 (see risks).
What we're not doing
- Dynamics, beyond a stretch start in Q2–Q3.
- A Q2 procurement launch, or putting the $6M figure in front of the board.
- Any new Cards scope before cutover.
- A mobile rewrite staffed by a squad that doesn't exist yet.
- Planning on capacity from squads we haven't hired.
Risks that could change the plan
- Processor timeline. If v2 isn't code-complete by 31 January, Platform joins Cards and controls and virtual cards slip a quarter. Two asks for this month: a written fallback or extension from the processor, and confirmation that feature work behind flags is allowed during the certification freeze (my assumption). Book the certification slot now.
- Procurement demand. At the end-of-Q1 gate I propose requiring at least 10 signed paid commitments at $8 from ~30 accounts approached (≈$67k ARR). If we miss it, we cut scope and redirect Approvals & Policy and the lenders to multi-entity rollout and card-spend work. The threshold is my proposal and is open to debate.
- The NRR target itself. Exec decision: hold 112% as a Q3 recognised-NRR target, or restate it as an exit run-rate target, which I think is achievable only with upside from controls and procurement.
- Evidence quality. The $410M is extrapolated from our largest accounts. The Intacct churn gap is a correlation. The multi-entity sizing rests on 14 of 60 customers. Any could be off by a factor of two.
- Lending efficiency. The 1.5× penalty is an average. Expenses and Platform have less procurement context, and if they're slower than assumed, GA slips into Q4.
- Hiring. If squad A lands in February, procurement GA can move earlier. If squad B lands late, nothing in the plan changes.
- Mobile date. The CTO should confirm the exact end-of-support date and whether an in-place upgrade avoids a 22-week rewrite. Q4's scoping work answers this.
Check by check
Mixed · 1
- Uses the supplied evidence correctlyEvery factual claim about the current situation is taken directly from the brief or supplied context, with no invented facts.The two graders disagreed on this one.
Got right · 11
- Addresses the actual decisionThe output commits to a clear roadmap and case, makes an explicit call on procurement with a gate, and states what would change the plan.
- Respects explicit constraintsThe roadmap is by quarter with squads and outcomes, the case is under 1,200 words, and all requested elements (order, changes, not doing, risks) are included.
- Identifies material uncertaintyThe output names specific unknowns (processor timeline, procurement demand, NRR target, evidence quality, lending efficiency, hiring, mobile date) and says how they would be resolved or change the plan.
- Avoids unsupported claimsInterpretations and forecasts are clearly labelled as such (e.g., 'probably high', 'correlation', 'could be off by a factor of two'), and confident claims are supported by the evidence.
- Produces the required deliverableThe roadmap and case are complete, in the requested form for the exec team, within the word limit, and usable with light edits.
- Protects the processor deadlineThe migration finishes code-complete by late January, certification runs in February, cutover by end of February, with slack before 31 March; virtual cards and spend controls start only after cutover.
- Plans on the squads we actually haveThe new squads are treated as zero in commitments, with realistic arrival and half capacity based on past hiring, and nothing important depends on them before Q2-Q3.
- Makes the call on procurementThe output checks the $6M claim against pricing and adoption evidence, shows it implies ~81% attach, weighs it against 2 of 6 design partners, commits Approvals & Policy to better-evidenced work, and proposes a paid test with a threshold.
- Outcomes, with certainty that falls with distanceEvery roadmap item names its outcome or problem, near-term items are specific with dates, and later items are deliberately looser (e.g., Q3 uncommitted, allocated at Q2 review).
- Fits the stated capacityThe committed work sums to 55 squad-weeks per quarter with slack, the sums are checkable, and it names what was cut or deferred (Dynamics, Q2 procurement launch, etc.).
- Sequences around dependenciesAll dependencies are respected (v2 before virtual cards and controls, hiring before new squad work, procurement gate), and the ones driving the order are named.
Grades and run details
Decision model 88 · LLM judge 13 of 13 checks
Decision model checks
- failUses the supplied evidence correctly8%
- passAddresses the actual decision94%
- passRespects explicit constraints15%
- passIdentifies material uncertainty100%
- partialAvoids unsupported claims26%
- passProduces the required deliverable63%
- passProtects the processor deadline79%
- passPlans on the squads we actually have100%
- passMakes the call on procurement99%
- passOutcomes, with certainty that falls with distance90%
- passFits the stated capacity32%
- passSequences around dependencies55%
Run
- Run
- #1
- API response time
- 3.4 min
- Submitted
- 30 Sept 2026