Tasks / Define

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 6 graded outputs by 3 models. 67% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits to a clear two-quarter roadmap and named deferrals, and states the evidence or conditions that would change later choices.
    GPT-6.1 Sol · API · Two squads, eight asks, one half
  2. Identifies material uncertainty100% pass
    It names material unknowns such as reminder causation, recovered support time, dashboard adoption, timely cancellations, and Physitrack demand, and says how they would be resolved.
    GPT-6.1 Sol · API · Two squads, eight asks, one half
  3. Outcomes, with certainty that falls with distance100% pass
    Items are framed by outcomes, near-term work is specific, and later work is deliberately looser pending evidence.
    GPT-6.1 Sol · API · Two squads, eight asks, one half

Where it slips

  1. Fits the stated capacity67% pass
    Committed Q1/Q2 Approvals & Policy work exceeds 11 squad-weeks per quarter, and the conditional procurement plan displaces multi-entity and spend-controls work.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  2. Respects explicit constraints67% pass
    It exceeds the 1,200-word limit for the case and proposes procurement work that would consume the Approvals & Policy squad's multi-entity and spend-controls capacity.
    GPT-6.1 Sol · API · A year of spend management, with a hard deadline
  3. Every commitment serves the goals75% pass
    Physitrack integration is committed in Q1 but is not tied to either CEO goal (no-shows or multi-site churn) or to freeing capacity, and the criterion requires such items to be deferred with a reason.
    Opus 5.5 · Claude · Two squads, eight asks, one half

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Harbor. Tom Achebe, our CEO, has shared his draft roadmap for the next four quarters (Q4 2026 to Q3 2027), and asked you to turn it into the roadmap we take to next week's planning offsite. The exec team (CEO, CFO, CRO and CTO) will read it beforehand. Write: 1. The roadmap itself, by quarter or by now, next and later: what each squad works on and the outcome each item serves. 2. The case for it, in no more than 1,200 words: why this order, what you changed from Tom's draft and why, what we're not doing, and the risks that could change the plan. The pack below is everything we have. Not all of it matters equally.

About HarborSpend management (corporate cards, expenses and approvals) for companies with 200 to 2,000 employees. 1,100 customers, $86M ARR, average 700 employees per customer. 58% of revenue is interchange on card spend; the rest is subscription. Net revenue retention (NRR) is 104%.
FY27 goal (approved by the board, September 2026)Raise NRR to 112% by the end of Q3 2027 by expanding inside existing customers: more spend on Harbor cards, more entities and more modules. New-logo growth matters, but it comes second this year.
Squads and capacityFive squads: Cards, Expenses, Approvals & Policy, Integrations and Platform. After support and on-call, each has about 11 squad-weeks of roadmap capacity a quarter. Squads can take work from another squad's area, but it takes them about 1.5× as long.
HiringTwo new squads were approved in September 2026. Tom's draft counts the first from Q1 2027 and the second from Q2 2027, both at full capacity. Last year, our three approved squads took 5, 6 and 8 months from approval to their first full sprint, and each delivered about half its capacity in its first quarter.
Card processor deadlineOur card processor retires its v1 API on 31 March 2027. Every card authorisation we process runs through v1 today. Migrating to v2 is about 18 squad-weeks of Cards work, and the processor must then run 4 weeks of certification testing before we can cut over. Certification is the processor's time, not ours, but nothing can change on our side while it runs.
Candidate work (estimates in squad-weeks, by owning squad)1. Processor v2 migration: Cards 18. Hard deadline above. 2. Virtual cards for software subscriptions: Cards 10. Needs v2. 3. Real-time spend controls (declines out-of-policy spend at the point of sale): Cards 8, Approvals & Policy 4. Needs v2. 4. Procurement module (purchase requests and approvals, a new paid add-on): Approvals & Policy 30, Integrations 6. 5. Multi-entity support (subsidiaries under one parent account): Platform 12, Approvals & Policy 10. 6. Sage Intacct integration: Integrations 9. 7. Microsoft Dynamics integration: Integrations 12. 8. SSO and SCIM provisioning: Platform 5. 9. Receipt-matching improvements: Expenses 6. 10. Mobile app rewrite: Expenses 22.
Tom's draft roadmapQ4 2026: Cards: v2 migration. Approvals & Policy: multi-entity. Integrations: Intacct. Platform: multi-entity. Expenses: receipt matching. Q1 2027: Cards: finish migration, start virtual cards. New squad A: procurement module. Integrations: Dynamics. Platform: SSO and SCIM. Q2 2027: Cards: virtual cards, spend controls. New squad A: procurement launch. New squad B: mobile rewrite. Q3 2027: Cards: spend controls. Everyone else: procurement adoption, mobile launch. Tom's note: “Procurement is our path to 112%. I've told the board it can add $6M of ARR in its first year.”
Procurement evidenceSix design partners have used a prototype since June. Two said they would pay for it; the other four already use a dedicated procurement tool and said they'd need a reason to switch. Proposed price: $8 per user per month, for users who raise purchase requests. At our customers, about 10% of employees raise purchase requests.
Card spend evidenceIn QBRs with our 60 largest customers, 26 asked for virtual cards for software subscriptions, and 31 CFOs asked for real-time spend controls. Finance estimates our customers pay about $410M a year of software subscriptions on other cards or by invoice (extrapolated from the 60 QBR accounts). Our net interchange is 1.1% of card spend.
Multi-entity evidence14 of our 60 largest customers have asked for it. Between them they have 52 subsidiaries not on Harbor, and 9 of the 14 have said in writing they would add their subsidiaries once it exists. Customer Success sized it at $2.9M ARR if all 52 joined at their parents' pricing.
Integrations evidence17% of customers use Sage Intacct through a CSV export. Their gross revenue churn is 13% a year, against 7% for customers on our NetSuite integration. Intacct was cited in 31% of lost new-logo deals last year, Dynamics in 8%. 3% of customers use Dynamics.
Other asksCRO: “Five enterprise deals worth $1.2M of pipeline need SSO and SCIM, and 11 existing accounts have it on their security review list.” CTO: “The mobile app is on a framework version that loses support in late 2027, and I want the rewrite done before then.” Support: receipt auto-match is at 71%, and unmatched receipts are the most common expense ticket.
What a strong answer does

A four-quarter roadmap that protects the 31 March deadline: the migration build has to finish by about early March to leave 4 weeks of certification, which is roughly all of Cards' capacity for Q4 and most of Q1 (18 of about 18 squad-weeks available before certification). So it adds help or buffer (Platform or another squad at 1.5×, or cutting everything else from Cards) and moves virtual cards out of Q1: both v2 features start only after cut-over (virtual cards Q2, spend controls Q2–Q3). It treats the new squads as upside: on past hiring they arrive around February to April 2027 and deliver half capacity in their first quarter, so nothing committed depends on them before Q2–Q3. It makes a clear call on procurement: the $6M claim needs about 890 accounts (81% of customers) paying $6,720 a year each (70 requesters × $8 × 12), and only 2 of 6 design partners would pay, so it replaces the 30-week build with a cheap paid test with a threshold, and keeps Approvals & Policy on multi-entity and spend controls, which have stronger expansion evidence ($2.9M sized, 9 of 14 committed in writing; about $1.5M a year of interchange from moving a third of the $410M of subscriptions onto virtual cards). Intacct comes early for its churn; Dynamics later or cut; SSO and SCIM are small and justified; the mobile rewrite waits for a new squad, before the framework's support ends. It lists what changed from Tom's draft, what's not being done, and the risks (certification slipping, hiring slipping), with commitment falling with distance.

Critical failures (cap the score)
  • Leaves the processor migration finishing too late for 4 weeks of certification before 31 March 2027, or schedules virtual cards or real-time spend controls before it
Case

v1.0 · synthetic · B2B SaaS, fintech, capacity, hard deadline, staff-level

Needs more than a quick edit9 of 12 checks passed

A PM would need to shorten the case and rebuild the Approvals & Policy capacity plan before using it. The core deadline and procurement logic are strong, but the capacity conflicts are material.

Output

Harbor FY27 roadmap — planning-offsite proposal

Q4 2026–Q3 2027

Objective: Raise NRR from 104% to 112% through expansion in existing customers. Protect card continuity first, then deliver the best-supported expansion opportunities. Treat procurement as a gated investment—not a committed $6M forecast.

1. Roadmap

Planning assumptions

  • Each existing squad has 11 squad-weeks per quarter, after support and on-call.
  • Neither new squad is required to deliver this plan. Hiring is upside, not committed capacity.
  • Numbers below are squad-weeks. Adoption work and reserves are timeboxed allocations, not additional feature estimates.
  • Procurement production work proceeds only if the commercial gate below passes. Otherwise, its allocations go to expansion activation and remain available for replanning.
SquadQ4 2026Q1 2027Q2 2027Q3 2027
CardsProcessor v2 migration: 11. Protect all card revenue and unblock new card capabilities.Finish migration: 5. Complete certification and cut over before 31 March. Reserve: 6 for deadline contingency and cutover stabilization; no planned feature changes during certification.Virtual cards for software subscriptions: 10; rollout reserve: 1. Capture subscription spend currently on other cards or invoices.Real-time spend controls: 8; rollout/reserve: 3. Give CFOs confidence to put more spend on Harbor.
Approvals & PolicyMulti-entity: 10. Enable subsidiary expansion. Procurement validation: 1, supported by PM, Sales and Finance. Test willingness to pay before committing production capacity.Multi-entity activation: 1. Conditional procurement build: 10. Establish the purchase-request and approval workflow for a paid add-on.Spend controls: 4. Prepare policy capabilities for the Q3 Cards release. Conditional procurement build: 7.Conditional procurement completion and launch: 11. Deliver the add-on if the gate passes; otherwise focus on multi-entity and controls adoption.
IntegrationsSage Intacct: 9; launch/activation: 2. Replace CSV workflows and address a retention risk.Intacct activation and reliability: up to 11. Move existing CSV customers onto the integration, prioritized by revenue and renewal risk.Conditional procurement integration work: 6. Intacct activation/reserve: 5.Intacct adoption and, if launched, procurement onboarding: up to 11. Turn shipped capabilities into retained and expanded revenue.
PlatformMigration assistance: 3, equivalent to 2 Cards squad-weeks at the cross-squad rate. Multi-entity: 8. Create processor schedule margin while advancing subsidiary support.Finish multi-entity: 4. SSO/SCIM: 5. Activation/reserve: 2. Launch subsidiary expansion and remove security blockers for existing accounts and enterprise deals.Conditional procurement assistance: 3, equivalent to 2 Approvals & Policy squad-weeks. Multi-entity/security activation and reserve: 8.Multi-entity and identity reliability/activation: up to 11. Support subsidiary onboarding and secure expansion.
ExpensesReceipt matching: 6; measurement and rollout: 5. Improve the 71% match rate and reduce the largest expense-support burden.Receipt-matching follow-through and mobile rewrite preparation: up to 11. Measure ticket reduction and prepare a safe migration; no additional feature scope assumed.Mobile rewrite: 11. Replace the framework approaching end of support.Finish mobile rewrite: 11, including release work within the estimate. Target completion before the late-2027 support deadline.

Delivery gates

Processor gate — non-negotiable - Allocate identifiable, independently executable migration work to Platform in Q4. - Complete migration code by early/mid-February, then freeze Harbor-side changes for the processor’s four-week certification. - Target cutover in mid-March, leaving contingency before 31 March. - If certification or implementation slips, pause card feature work and reallocate capacity immediately.

Procurement gate — decision by the end of Q4 Approve production funding only with: - Signed paid-pilot commitments at a validated price and requestor count—not general expressions of interest. - Evidence that customers with dedicated procurement tools will switch, or a clearly defined segment that does not require displacement. - A bottom-up expansion pipeline and pilot success criteria covering adoption, willingness to pay and implementation effort.

If the gate passes, the plan supplies the full estimate without hiring: 28 Approvals & Policy weeks + 3 Platform weeks at 1.5× = 30 equivalent weeks, plus 6 Integrations weeks. Target a Q3 launch, not a Q2 launch. If it fails or arrives late, do not start the full build; return to the exec team with revised scope and timing.

Outcome scorecard

Finance, Product and CS should maintain a monthly, existing-customer expansion bridge:

InvestmentPrimary outcome to track
Processor migrationSuccessful certified cutover; no processor-driven interruption
Multi-entitySubsidiaries contracted and live; incremental ARR
Virtual cardsIncremental software spend moved to Harbor; net interchange
Spend controlsAdoption among requesting customers; subsequent spend expansion
IntacctCSV customers activated; renewal and churn outcomes
SSO/SCIMExisting-account security blockers resolved; expansion unlocked
Receipt matchingMatch rate and unmatched-receipt ticket volume
ProcurementPaid pilots, active requestors and contracted incremental ARR
Mobile rewriteSafe release before framework support ends

Do not count pipeline, enabled subsidiaries or estimated spend as realized NRR.

---

2. The case for this plan

Why this order

First, protect the business we already have. Every card authorization depends on the retiring processor API, and interchange represents 58% of revenue. The migration is not an ordinary roadmap item. Tom’s allocation of 11 Cards weeks in Q4 leaves seven weeks in Q1, followed by four calendar weeks of certification. That is too little schedule margin once effective capacity and the certification freeze are considered.

Moving three Platform weeks into Q4 migration work produces two equivalent Cards weeks. Cards then has five implementation weeks remaining in Q1. This buys a realistic certification window and a cutover buffer. We should not schedule virtual cards into that buffer.

Next, pursue expansion with the strongest customer evidence. Multi-entity has nine written commitments among 14 requesting customers. The identified opportunity is $2.9M ARR across 52 subsidiaries, although that is a ceiling—not a forecast. We should validate pricing, rollout requirements and which subsidiaries are covered by the nine commitments before booking expected revenue.

Virtual cards address a substantial identified spend pool. At 1.1% net interchange, capturing all $410M of estimated software spend would generate approximately $4.5M annually. Capturing 25–50% would generate roughly $1.1M–$2.3M, before considering rollout timing. These are scenarios, not forecasts: the spend estimate is extrapolated, and invoice spend may not be readily cardable.

Virtual cards precede real-time controls because they offer a direct, measurable spend-capture opportunity. Controls follow to broaden CFO confidence and adoption. If customer testing shows controls are a prerequisite for moving software spend, we should reverse their order.

Retention and security are part of the expansion strategy. Intacct serves 17% of customers; its CSV cohort has materially higher gross revenue churn than the NetSuite cohort. That does not prove integration causes the difference, but it supports prioritizing Intacct over Dynamics. SSO/SCIM is relatively small and addresses security reviews at 11 existing accounts, as well as new-logo pipeline. Receipt matching similarly offers a bounded investment against a known support burden.

What changed from Tom’s draft

  1. Removed assumed hiring capacity. Last year’s squads took five to eight months to reach a first full sprint and delivered roughly half capacity in their first quarter. Full-capacity squads in Q1 and Q2 are not a dependable planning assumption.
  1. Made multi-entity delivery feasible. Its Platform estimate is 12 weeks, so it cannot fit wholly into an 11-week Q4. With migration assistance included, Platform delivers eight weeks in Q4 and four in Q1.
  1. Gated procurement and moved any launch to Q3. Two of six design partners expressing willingness to pay is insufficient evidence for a full production commitment, particularly when four already have procurement tools.

At the proposed price, an average customer has about 70 requestors and generates $6,720 annually. Reaching $6M requires approximately 893 average-sized customers—81% of our current base. Larger customers or different packaging could change that calculation, but the prototype evidence does not support Tom’s forecast. A first-year ARR claim also should not be treated as revenue available by Q3.

  1. Funded the mobile rewrite with the existing Expenses squad. It is necessary lifecycle work, not the leading NRR investment. Scheduling its 22 weeks across Q2 and Q3 meets the known support horizon without depending on a new squad.
  1. Added explicit activation capacity. Shipping is not expansion. Subsidiary onboarding, integration migration and card-spend conversion require attention after launch.

What we are not doing

  • Dynamics this year: only 3% of customers use it, versus 17% on Intacct; its new-logo evidence is also weaker.
  • An unconditional procurement build or $6M revenue commitment.
  • A Q2 procurement launch funded by unstaffed squads.
  • New mobile feature scope beyond the rewrite.
  • Filling every reserve with another launch. Remaining capacity is deliberately available for activation, technical uncertainty and measured opportunities.

Risks that could change the plan

Processor execution: Cross-squad assistance must be genuinely separable work. If it is not, the CTO must secure qualified help or displace other work immediately. Certification failure overrides all discretionary card work.

Commercial conversion: Written multi-entity intent is not a contract; estimated software spend is not captured spend. CS and Finance must validate both before treating them as the expansion bridge.

Procurement: Passing the commercial gate earns production funding, not a $6M forecast. Implementation complexity or weak paid-pilot adoption could still stop or narrow the investment.

Retention assumptions: Intacct’s churn association may reflect customer mix rather than integration quality. Measure adoption and renewal outcomes rather than assuming the six-point difference is recoverable.

NRR sufficiency: Moving from 104% to 112% represents roughly $6.9M of additional annualized cohort revenue against an $86M reference base, subject to Finance’s actual cohort calculation. The evidence does not yet prove this roadmap closes that gap. Before the offsite, Finance and CS should build a customer-level bridge covering expansion, contraction, churn and timing, without double-counting opportunities.

Hiring and technical estimates: Earlier hiring can accelerate a validated bottleneck. Estimate changes or a nearer mobile-support deadline can force reprioritization. Neither should quietly expand the committed scope.

Check by check

Got wrong · 1
  • Respects explicit constraintsIt exceeds the 1,200-word limit for the case and proposes procurement work that would consume the Approvals & Policy squad's multi-entity and spend-controls capacity.
Mixed · 2
  • Produces the required deliverableThe roadmap and case are present, but the case is too long and the roadmap has capacity conflicts that would require major rework.The two graders disagreed on this one.
  • Fits the stated capacityCommitted Q1/Q2 Approvals & Policy work exceeds 11 squad-weeks per quarter, and the conditional procurement plan displaces multi-entity and spend-controls work.The two graders disagreed on this one.
Got right · 9
  • Uses the supplied evidence correctlyThe output's current-situation facts and arithmetic are drawn from the supplied context, with speculative items labelled as scenarios or assumptions.
  • Addresses the actual decisionIt commits to a clear order: protect the processor migration, then virtual cards and spend controls, with procurement gated and Dynamics deferred, and states the gate that would change procurement.
  • Identifies material uncertaintyIt names material unknowns such as certification slippage, cross-squad separability, procurement willingness to pay, Intacct churn causality, and NRR sufficiency, with resolution paths.
  • Avoids unsupported claimsIt labels causal and revenue claims as scenarios, ceilings, or assumptions rather than established facts.
  • Protects the processor deadlineIt finishes migration by early/mid-February, freezes for four weeks of certification, targets mid-March cut-over, and starts v2 features only in Q2/Q3.
  • Plans on the squads we actually haveIt explicitly treats the two new squads as upside and commits no essential work to them before Q2/Q3.
  • Makes the call on procurementIt checks Tom's $6M claim against pricing and adoption evidence, requires a paid-pilot gate, and defers any full build to Q3.
  • Outcomes, with certainty that falls with distanceEach item names an outcome, and later quarters are looser with reserves and conditional allocations.
  • Sequences around dependenciesIt sequences virtual cards and spend controls after v2 cut-over and places procurement behind a commercial gate.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly37%
  • passAddresses the actual decision96%
  • partialRespects explicit constraints37%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims68%
  • passProduces the required deliverable52%
  • passProtects the processor deadline94%
  • passPlans on the squads we actually have91%
  • passMakes the call on procurement100%
  • passOutcomes, with certainty that falls with distance87%
  • passFits the stated capacity26%
  • passSequences around dependencies44%
Run
Run
#1
API response time
2.1 min
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1Sonnet 5.5withAPI84.7100.02None
2GPT-6.1 SolwithAPI97.984.62None
3Opus 5.5withClaude82.465.12None

About the task

The PM job

Turning strategy into a sequenced plan.

Why it matters

A roadmap is where strategy meets capacity. Dated feature lists turn guesses into promises.

What good looks like

  • Items are problems or outcomes, not just features
  • Sequencing reflects dependencies
  • Explicit trade-offs
  • Commitment falls with distance

Deliberately not measured

  • Gantt formatting
Capability tested

Sequencing under constraints

The failure we’re looking for

A dated wishlist sorted by excitement

Grading

Decision model and LLM judge, calibrated against a blind PM review