Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  2. Identifies material uncertainty100% pass
    It names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  3. Uses the interviews faithfully100% pass
    All quotes are accurate and correctly attributed to A05, A08, A04, A07, and A12.
    GPT-6 Luna · API · An AI SDR for small agencies

Where it slips

  1. A cheap test that can actually read out64% pass
    The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.
    GPT-6.1 Sol · API · An AI SDR for small agencies
  2. Re-estimates the revenue correctly71% pass
    It shows working but lands on a central $2.3M and $1M–$6M range, not roughly the $5–10M range required by the grading rubric, and its $9.1M pilot upper bound is not used as the main re-estimate.
    Sonnet 5.5 · API · The CEO's embedded-payments bet
  3. Avoids unsupported claims71% pass
    The memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
    Opus 5.5 · Claude · The CEO's embedded-payments bet

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Fieldline. Our CEO wants to commit three squads for three quarters to Fieldline Pay, embedded payments for our customers' invoices, and has told the board it could add $24M of annual revenue within two years. Priya Raman, our CPO, has asked you to write the strongest challenge to the plan as a pre-read for next week's exec offsite. The CEO, CFO and Head of Sales will all read it. Write a memo of no more than 1,200 words that: 1. Names the single assumption the plan most depends on that the evidence does not support, and shows why, using the numbers in the pack. 2. Re-estimates the revenue from the supplied data, showing your working, as a range. 3. Says what we would need to see to be proved wrong, and the cheapest test that would show it within six weeks. 4. Says what, if anything, we should do instead or how the bet should change. The pack below is everything we have. Some of it matters more than the rest.

What the model was given11 items: About Fieldline, CEO's memo to the exec team (excerpt), CFO's revenue model, Invoicing data (last 12 months, all customers), How invoices are paid today, Payments partner term sheet (unit economics), Pilot (4 months), Sales notes on the largest accounts, Customer interviews (22 customers, last quarter), Competitor, Engineering
About FieldlineField-service software for trades businesses (plumbing, HVAC, electrical): scheduling, dispatch, quotes and invoicing. 2,300 customers, $41.0M ARR, average $17,800 per customer. Series C; the next raise is planned in about 14 months.
CEO's memo to the exec team (excerpt)“Every invoice our customers send is money we don't touch. Our customers invoiced $4.83B last year through Fieldline. If we process those payments, we become part of how they get paid, not just how they schedule. The model is simple: 70% of customers adopt within 18 months, the average customer invoices $2.1M a year, and we keep a 0.7% blended net take. That is $24M of new annual revenue by month 24, more than half our current ARR, and it makes the next raise a very different conversation. Tradesly has shown it works: payments are now 22% of their revenue. I've told the board I believe payments can be 35% of our revenue by 2028. I want three squads on this from next quarter, which means pausing the scheduling rewrite.”
CFO's revenue modelCustomers: 2,300. Adoption by month 18: 70% (1,610 customers). Annual invoiced value per customer: $2.1M (total invoiced $4.83B ÷ 2,300). Blended net take rate: 0.7%. Month-24 revenue run-rate: 1,610 × $2.1M × 0.7% = $23.7M. CFO's note: “Adoption and take rate are the CEO's assumptions. I haven't stress-tested them.”
Invoicing data (last 12 months, all customers)Total invoiced value: $4.83B. Mean per customer: $2.1M. Median per customer: $640k. The largest 5% of customers (115) account for 48% of invoiced value ($2.32B). By job type, residential jobs are 42% of invoiced value and commercial jobs (property managers, facilities contracts) are 58%.
How invoices are paid todayFrom the 690 customers (30%) who record the payment method in Fieldline, by share of invoice value: card 19%, ACH/bank transfer 44%, check 31%, cash 6%. Card share is 38% of residential invoice value and 5% of commercial. Average commercial invoice: $3,800; average residential invoice: $410. Commercial clients pay on net-45 or net-60 terms; the average commercial invoice is paid 52 days after it is sent.
Payments partner term sheet (unit economics)Card: the customer is charged 2.9% + $0.30 per payment; our all-in cost (interchange, network, partner fee) is about 2.2%, so we net about 0.7% of card value. ACH: the customer is charged a flat $2.00 per payment; our cost is $0.40. Checks and cash earn nothing unless the payer switches to card or ACH. The partner handles licensing, KYC and risk; they have approved our application.
Pilot (4 months)62 customers invited, 38 adopted (61%). Pilot customers' invoice value is 64% residential (the customer base is 42%). With pay-by-link on every invoice, card share of invoices paid through Fieldline Pay rose to 71% for residential and 6% for commercial. Net payments revenue, annualised: $214k across the 38 customers ($5,630 per customer per year).
Sales notes on the largest accountsOf the 115 largest customers, 71 have multi-year contracts with an existing payment processor, most running to 2028. Head of Sales, in Slack: “None of the big ones will move processors before their contracts end, and their property-manager clients will not pay 2.9% on a $3,800 invoice. They'll pay by ACH or check like they always have.”
Customer interviews (22 customers, last quarter)17 of 22 named getting paid on commercial jobs as their biggest cash problem (“I'm floating $200k of payroll while property managers sit on invoices for two months”). 9 said they would pay a fee to be paid faster. 6 said they won't offer card payment because clients fight the surcharge. Of the 14 customers with mostly commercial work, 11 said their clients require ACH or check.
CompetitorTradesly launched embedded payments in 2025 and says payments are now 22% of its revenue. Its published case studies are almost all home-services businesses; it describes its customers as about 80% residential.
EngineeringThe plan needs three squads for three quarters. The scheduling rewrite it would pause is the top-cited reason for churn among customers with more than 50 technicians (lost ARR from that segment last year: $1.9M).
What a strong answer doesThe answer key the graders mark against

Names the load-bearing assumption as the 0.7% blended net take: that is the card margin applied to all invoiced value, but only about 19% of value is paid by card (5% of commercial, which is 58% of value), ACH nets about 0.04% on a $3,800 invoice, and checks and cash earn nothing. Re-estimates with the supplied data: even at the pilot's improved mix (about a third of value on card), the blended take is roughly 0.2–0.25%, so the run-rate is about $5–9M, not $24M (the pilot's $5,630 per customer × 1,610 = $9.1M is an upper bound because the pilot was residential-heavy; the 71 locked-in large accounts, with 48% of value in the top 5%, cut it further). Notes that $24M needs essentially all invoice value on card. Treats Tradesly as a residential-mix result, not proof, and engineering capacity, licensing and competition as secondary. Keeps what the evidence supports: card payments work for residential invoices, and the strongest pull is commercial cash flow (17 of 22), so it reshapes the bet toward faster payout or financing on commercial invoices, which earns on the volume the processing model can't. Proposes a cheap test with a threshold that would prove the challenge wrong (for example, commercial card share well above 20% with pay-by-link in a six-week cohort, or uptake of paid instant payout), and is written for a room that includes the CEO: firm on the numbers, not personal.

Critical failures (cap the score at 40)
  • Accepts the $24M projection or the 0.7% blended net take rate as sound
  • Invents market sizes, competitor figures or customer statistics that are not in the pack
Case

Staff-level · v1.7 · synthetic · B2B SaaS, fintech, strategy, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: Invents market sizes, competitor figures or customer statistics that are not in the pack. Score capped at 40.

Needs more than a quick edit6 of 11 checks passed

A PM would need to recheck the locked-GMV arithmetic and the revenue model assumptions before using it. The memo is also too long for the stated constraint.

Output

MEMORANDUM

TO: Exec Team (CEO, CFO, CPO, Head of Sales) FROM: Staff Product Manager DATE: October 24, 2024 SUBJECT: Strategic Challenge & Alternative Plan: Fieldline Pay

---

Executive Summary

Committing three squads for three quarters to Fieldline Pay under the current $24M revenue projection risks our 14-month Series raise.

The $24M model assumes a 0.7% blended net take rate across our total $4.83B GMV. In reality, our invoice mix (58% commercial) and payment fee structure yield an effective take rate of roughly 0.15%–0.28% on accessible volume. Realistic Month-24 payments revenue is $4.8M–$7.8M—an 68%–80% shortfall against the board commitment.

Simultaneously, pausing the scheduling rewrite directly imperils our core ARR by aggravating churn in our highest-value segment (which lost $1.9M ARR last year). We must right-size this bet, protect core retention, and re-orient our fintech strategy around commercial cash flow.

---

1. The Fatal Assumption: A 0.7% Blended Net Take Rate on Total GMV

The $24M model hinges on earning 0.7% net margin on $3.38B of adopted GMV (70% of $4.83B). The data contradicts this in three ways:

#### A. Card Economics Do Not Apply to Commercial GMV (58% of Volume) Our partner agreement gives us 0.7% on card volume (2.9% + $0.30 fee less 2.2% cost), but only a flat $1.60 net per ACH transaction ($2.00 fee less $0.40 cost), and $0 on checks/cash. * Commercial jobs represent 58% of our GMV ($2.80B) with an average invoice of $3,800. * Property managers will not pay a 2.9% card surcharge ($110.20 per invoice). In our pilot, commercial card adoption barely moved: from 5% baseline to just 6%. * Monopolized by ACH, a $3,800 commercial invoice earns us a flat $1.60. That is an effective take rate of 0.042% (4.2 basis points)—not 70 basis points.

#### B. The Tradesly Comparison is a False Equivalency Tradesly's 22% payments revenue share is driven by a customer base that is ~80% residential. Fieldline’s base is 58% commercial. Residential customers pay $410 tickets where card convenience wins (pilot card share hit 71%). Commercial clients operate on net-45/60 terms and require ACH or check (11 of 14 interviewed confirmed this). Applying residential payment dynamics to a commercial invoice base is fundamentally flawed.

#### C. 30% of Total GMV Is Contractually Inaccessible Until 2028 The largest 5% of our customers (115 accounts) control 48% of total invoiced volume ($2.32B). Sales confirmed that 71 of these 115 accounts are locked into multi-year processor contracts running to 2028. That removes ~$1.43B in GMV from near-term adoption.

---

2. Revenue Re-Estimation (Month-24 Run-Rate)

#### Addressable GMV & Unit Take Rates * Total GMV: $4.83B * Less Contractually Locked Tier-1 GMV: -$1.43B (71/115 largest accounts) * Net Addressable GMV: $3.40B * Residential (42%): $1.43B (avg ticket $410 $\rightarrow$ ~3.48M invoices) * Commercial (58%): $1.97B (avg ticket $3,800 $\rightarrow$ ~518k invoices)

#### Expected Net Revenue per $1M GMV (Based on Pilot Performance): * Residential ($1M GMV): * 71% Card ($710k $\times$ 0.70%) = $4,970 * 25% ACH ($250k $\div$ $410 = 610 txns $\times$ $1.60) = $976 * 4% Check/Cash = $0 * Residential Blended Take Rate: ~0.59% * Commercial ($1M GMV): * 6% Card ($60k $\times$ 0.70%) = $420 * 60% ACH ($600k $\div$ $3,800 = 158 txns $\times$ $1.60) = $253 * 34% Check = $0 * Commercial Blended Take Rate: ~0.067% (6.7 bps) * Portfolio Blended Take Rate on Accessible GMV: $(0.42 \times 0.59\%) + (0.58 \times 0.067\%) =$ 0.287%

``` +------------------------------------+------------------+------------------+------------------+ | Metric | Conservative | Base Case | Aggressive | +------------------------------------+------------------+------------------+------------------+ | Accessible GMV Adoption | 40% ($1.36B) | 55% ($1.87B) | 70% ($2.38B) | | Res / Comm Effective Take Rate | 0.22% (Card dips)| 0.28% (Pilot mix)| 0.33% (Card lift)| +------------------------------------+------------------+------------------+------------------+ | Re-estimated Month-24 Revenue | $3.0M | $5.2M | $7.8M | +------------------------------------+------------------+------------------+------------------+ | Shortfall vs. $24M Plan | -$21.0M (-87%) | -$18.8M (-78%) | -$16.2M (-68%) | +------------------------------------+------------------+------------------+------------------+ ``` (Sanity check against pilot: 38 pilot customers generated $214k/yr = $5,630/customer. If 1,610 customers adopted at that run-rate, total revenue would be $9.06M. However, the pilot was 64% residential vs. 42% in our real base; adjusting for our actual commercial mix brings the pilot run-rate directly into our $5M–$7M range.)

---

3. Falsification Criteria & The 6-Week Test

#### What Would Prove This Challenge Wrong: 1. Commercial payers are willing to pay credit card surcharges (2.9%) at scale on invoices over $2,000, lifting commercial card share above 25%. 2. Mid-market/enterprise customers are willing to buy out or break existing processor contracts to consolidate inside Fieldline. 3. Commercial customers will accept basis-point pricing on ACH (e.g., 0.5% capped at $25) rather than a flat transaction fee.

#### The Cheapest 6-Week Test ($0 Eng, 1 Ops/Sales Lead): Instead of committing three engineering squads, run a zero-code commercial pricing test: 1. Target Group: 40 non-pilot customers with heavy commercial mix (HVAC/commercial plumbing) invoicing property managers. 2. Execution: Issue invoices using off-the-shelf payment links provided directly by our payments partner (dashboard-configured, zero Fieldline engineering). 3. Intervention: Enable Pay-by-Link on 500 commercial invoices ($2,000+) across two cohorts: * Cohort A (Standard): 2.9% fee passed to payer or absorbed by merchant. * Cohort B (Convenience ACH): $5 flat fee or 0.5% fee for direct bank integration. 4. Success Threshold to Validate Plan: If card volume on commercial exceeds 20% or ACH basis-point monetization is accepted without client backlash, the CEO's revenue model becomes viable. If commercial card adoption remains under 8%, the thesis is conclusively disproven.

---

4. Recommended Strategic Pivot

Pausing the scheduling rewrite is an unacceptable risk. That rewrite addresses the #1 driver of churn among customers with >50 technicians—a segment that cost us $1.9M in ARR last year alone. Churning enterprise SaaS ARR to pursue low-margin payment processing right before a Series raise will impair our valuation multiple.

We should adjust our strategy as follows:

``` +-----------------------------------------------------------------------------------------------+ | Proposed Engineering & Resource Allocation | +-------------------------------------+---------------------------------------------------------+ | Current Plan (Paused Scheduling) | Recommended Allocation | +-------------------------------------+---------------------------------------------------------+ | Fieldline Pay: 3 Squads (9 mos) | Scheduling Rewrite: 2 Squads (Protects $1.9M+ ARR Churn)| | Scheduling Rewrite: 0 Squads (Paused)| Fieldline Pay (Residential MVP): 1 Squad (6 mos) | | Commercial Discovery: None | Product/Finance: Commercial Working Capital Discovery | +-------------------------------------+---------------------------------------------------------+ ```

#### 1. Allocate 1 Squad to Fieldline Pay (Residential Focus) Package an MVP pay-by-link flow exclusively for residential jobs ($410 avg invoice, 71% card share). One squad can ship this within two quarters using out-of-the-box partner UI. This captures the high-margin, viable portion of payments (~$3.0M–$4.0M ARR) without over-investing.

#### 2. Fully Fund the Scheduling Rewrite (2 Squads) Protect the core $41.0M SaaS business. Retaining our 50+ tech accounts protects enterprise ARR and maintains our clean SaaS growth story for the Series raise in 14 months.

#### 3. Monetize Commercial Cash Flow, Not Processing (The Real Opportunity) 17 of 22 interviewed customers cited commercial float (waiting 52 days on net-45/60 while funding $200k payroll) as their single biggest problem. Nine stated they would gladly pay a fee to get paid faster. * Rather than fighting over a $1.60 ACH fee, evaluate accelerated invoice settlement (e.g., 1.5%–2.0% fee for Net-2 payout vs. Net-52, balance-sheet funded via partner). * Earning 1.5% on even 20% of our $2.8B commercial GMV yields $8.4M in high-margin fintech ARR by solving an acute customer pain point.

Summary

Let us present the board with a defensible plan: secure our core SaaS ARR, deploy 1 squad to capture ~$3M–$4M in residential payments, and test commercial invoice financing to solve our customers' true cash flow crisis.

Claims the judge couldn’t find in the brief

  • 30% of total GMV, or about $1.43B, is contractually inaccessible until 2028 because 71 of the 115 largest accounts are locked in.
  • Accessible GMV can be split 42% residential and 58% commercial, and pilot card shares can be combined with assumed ACH/check splits to produce blended take rates.

Grades and run details

Decision model 91 · LLM judge 6 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly21%
  • passAddresses the actual decision97%
  • partialRespects explicit constraints47%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims33%
  • passProduces the required deliverable74%
  • passFinds the load-bearing assumption100%
  • passRe-estimates the revenue correctly44%
  • passSays what would prove it wrong100%
  • passReshapes the bet from the evidence100%
  • passTigers, not paper tigers90%
Run
Run
#1
API response time
51 s
Submitted
1 Oct 2026

Usable with a quick edit11 of 11 checks passed

The memo is well-reasoned, evidence-backed, and directly usable for the offsite; only light editorial polish would be needed, with no factual or logical gaps to fill.

Output

To: Priya Raman and the executive team Subject: Fieldline Pay: validate the revenue engine before committing three squads

Recommendation: do not approve three squads for three quarters. Authorise a six-week commercial validation using the existing pilot and protect the scheduling rewrite. Payments may be a useful extension, but the supplied evidence supports an illustrative annual revenue range of $3.3M–$7.2M1, with substantial uncertainty. It does not support underwriting $24M or making that figure central to our next raise.

The assumption the plan depends on

The unsupported assumption is that customer adoption translates into monetising those customers’ entire invoice value at a 0.7% net yield.

The model multiplies customer count by average invoicing and card economics. Each component hides a material constraint:

  • Adoption is not payment capture. Enabling Fieldline Pay does not mean every payer switches. In the reporting sample, 81% of invoice value is currently paid without cards.
  • 0.7% is card economics, not blended economics. ACH earns $1.60 per payment. On a $3,800 commercial invoice, that is 0.042%, around one-seventeenth of the assumed yield. Checks and cash earn nothing.
  • Customer count is a poor proxy for accessible volume. The largest 5% generate 48% of invoicing. Of those 115 customers, 71 have processor contracts running mostly to 2028. Seventy per cent adoption among smaller customers could leave much of the revenue pool untouched.

Our mix compounds the problem. Commercial work represents 58% of invoice value. Commercial card share barely changed in the pilot, from 5% to 6%, despite pay-by-link on every invoice. Residential card share increased to 71%, but the pilot was 64% residential against 42% across Fieldline.

The customer problem is also different from the proposed solution. Seventeen of 22 interviewees identified slow commercial payment as their biggest cash problem. Offering another payment method does not establish that property managers will abandon net-45 or net-60 terms.

Tradesly shows payments can work for a business with approximately 80% residential customers. Its 22% revenue contribution does not validate our volume, adoption or margin assumptions.

What the supplied numbers support

Use:

Annual revenue = eligible invoice value × adopting share of that value × realised net yield.

The following are scenarios, not statistical bounds. Neither the exact invoicing of the 71 contracted accounts nor representative payment capture is supplied.

InputLower scenarioUpper scenario
Eligible annual invoice value$2.51B: exclude all largest 115 accounts$3.40B: exclude the 71 contracted accounts, assuming equal volume within the top 115
Adoption by invoice value61%, matching pilot customer adoption70%, matching the plan
Card share by valueCurrent reported 19%33.3%: 42% × 71% residential + 58% × 6% commercial
ACH treatmentCurrent reported 44% of valueEvery remaining payment becomes ACH
Modelled net yield0.216%0.304%
Annual revenue$3.3M$7.2M

Working:

Lower yield:

`19% × 0.7% + 44% × $1.60 × (42% ÷ $410 + 58% ÷ $3,800) = 0.216%`

`$2.51B × 61% × 0.216% = $3.3M`

Upper yield:

`33.3% × 0.7% + $1.60 × (42% × 29% ÷ $410 + 58% × 94% ÷ $3,800) = 0.304%`

`$3.40B × 70% × 0.304% = $7.2M`

These calculations assume eligible customers retain the overall residential/commercial mix and that adoption rates by customer also apply by value. The lower scenario assumes 44% ACH within both job types. Both assume full capture of the payment streams modelled. The upper scenario additionally interprets the pilot’s card shares as shares of value; if they are transaction counts, that extrapolation is unsupported. It also assumes complete conversion of checks and cash to ACH.

Consequently, $3.3M is not a floor. The reporting sample covers only 30% of customers, and actual volume concentration or capture could reduce revenue further.

A separate check points in the same direction: scaling the pilot’s $5,630 annual revenue per adopter to 61–70% of 2,300 customers produces $7.9M–$9.1M, before correcting its residential bias or account availability.

On the upper scenario’s eligible volume and 70% adoption, $23.7M would require approximately 1.0% net revenue on every adopted invoice dollar, above even the quoted card margin. Growth could alter this, but the pack supplies no growth assumptions to bridge the gap.

These are annualised net payments revenues under the partner terms, not profit or contracted software ARR.

What would prove this challenge wrong

We would need evidence that accessible customers generate materially more net revenue per total invoiced dollar than these scenarios imply, at sufficient committed volume to support $24M. That could come from much greater accessible volume, different economics or a demonstrably valuable paid service. Enthusiasm and payment-feature activation would not suffice.

The cheapest six-week test is to extend the existing pilot, with analyst, sales and limited engineering support:

  1. Establish the accessible denominator. Match the 71 processor contracts to actual invoice value, expiry dates and any permitted carve-outs. Replace the equal-volume assumption with account-level evidence.
  2. Recruit a stratified cohort. Include commercial-heavy businesses and large uncontracted accounts, alongside residential firms. Count refusals in adoption results and weight findings to eligible invoice value.
  3. Measure actual money movement. Track total invoicing, activated volume, captured volume, payment methods, realised fees and net revenue. Report residential and commercial results separately. Use existing receivables approaching their due dates to observe commercial payments within six weeks.

Before recruitment, agree that expanding investment requires a volume-weighted model supported by observed payments and credible customer commitments. To justify the original commitment, that model must show a defensible route to $24M by month 24. Six weeks cannot establish mature adoption or commercial collection cycles; an inconclusive result should not trigger automatic approval.

How the bet should change

If validation confirms residential economics, pursue a narrower residential payments offering with staged staffing and revenue gates. Separately test whether commercial customers will make concrete commitments to a faster-payment proposition. Nine expressions of willingness to pay justify discovery, not a financing product.

Keep the scheduling rewrite funded. The affected segment lost $1.9M ARR last year, with scheduling the top-cited churn reason. That is not all recoverable, but pausing the work has a documented cost against a speculative upside.

We should update the board with a scenario range and validation milestones. With the raise approximately 14 months away, protecting credibility and existing revenue matters more than preserving an unsupported headline.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeFrame the range up front as illustrative full-capture scenarios, with actual capture still unmeasured, so it isn't read as a forecast.

Grades and run details

Decision model 91 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly26%
  • passAddresses the actual decision99%
  • passRespects explicit constraints55%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims48%
  • passProduces the required deliverable76%
  • passFinds the load-bearing assumption100%
  • partialRe-estimates the revenue correctly23%
  • partialSays what would prove it wrong55%
  • passReshapes the bet from the evidence96%
  • passTigers, not paper tigers92%
Run
Run
#1
Time to output
80 s
Submitted
25 Sept 2026

Usable with a quick edit11 of 11 checks passed

The pre-read is tightly argued, numerically grounded, and directly usable by a senior PM; only minor phrasing tweaks would be needed to align the bull case with pure payments revenue if desired.

Output

To: Priya Raman (CPO), CEO, CFO, Head of Sales From: Staff Product Manager, Fieldline Date: September 25, 2026 Subject: Pre-read: Strategic Challenge to the Fieldline Pay Plan ($24M Revenue Target)

---

Executive Summary

Committing three squads for three quarters to Fieldline Pay to capture $24M in new revenue within two years is a high-risk bet that relies on a structural distortion in our financial model.

The plan’s core vulnerability is not execution; it is a misallocation of our customer base's economic reality. By using the mean invoice value ($2.1M) rather than the median ($640k), the model assumes our average customer mirrors our largest commercial accounts1, while simultaneously ignoring entrenched enterprise contracts and commercial payment norms.

Executing this plan requires pausing the scheduling rewrite, which risks $1.9M in annual churn from our largest, most valuable accounts. Below is the evidence-based challenge, a revised revenue range, an inexpensive six-week test, and a recommended alternative path.

---

1. The Dependent Assumption and Why the Evidence Fails It

> The Plan's Single Dependent Assumption: That we can achieve a $23.7M–$24M revenue run-rate by month 24 by applying a blended 0.7% net take rate across a homogenous $2.1M annual invoice volume per adopted customer.

The evidence flatly refutes this assumption in three ways:

  • The Mean vs. Median Distortion: The CEO's model relies on a mean invoice value of $2.1M. However, the median customer invoices just $640k. Our distribution is heavily skewed: the top 5% of customers (115 accounts) drive 48% ($2.32B) of our total invoiced value.
  • Enterprise Lock-in: Of those 115 largest customers, 71 are locked into multi-year contracts with existing payment processors running through 2028. As the Head of Sales noted, none will migrate early.
  • The Commercial Wall: Commercial jobs comprise 58% of our invoiced value ($2.80B). The average commercial invoice is $3,800 and is paid on net-45/net-60 terms (averaging 52 days). Property managers and facility contractors will not pay a 2.9% card surcharge on a $3,800 invoice. They pay via ACH ($2.00 flat fee, netting us $1.60) or check ($0 revenue), exactly as they always have.

Applying a 0.7% card take-rate model to commercial volume is fundamentally flawed because commercial clients do not use cards (card share is only 5% for commercial, vs. 38% for residential).

---

2. Re-Estimated Revenue Range (Working Included)

To model realistic revenue, we must segment our base by customer type, recognizing that residential and commercial segments have radically different payment behaviors and take rates.

#### Step-by-Step Working: 1. Customer Base: 2,300 total customers. 2. Adoption Rate: Pilot adoption was 61% (close to the 70% target). Let's model a realistic 60% adoption rate by month 24 = $1,380$ adopting customers. 3. Segmentation Split: Based on overall volume, 42% of value is residential ($2.03B) and 58% is commercial ($2.80B). Total invoiced value = $4.83B. * Total Residential Invoiced: $2.03B ($\approx$ $882k$ per customer across 2,300) * Total Commercial Invoiced: $2.80B ($\approx$ $1.22M$ per customer across 2,300) 4. Take Rates & Behavior: * Residential: 71% card adoption via pay-by-link. Net take rate on card is 0.7%. ACH/Check take rate is near zero (flat $1.60 net on ACH, negligible volume). Effective blended take rate on residential volume $\approx$ $0.7\% \times 71\% \approx \mathbf{0.50\%}$. * Commercial: 6% card adoption, 94% ACH/Check/Terms. Card take rate is 0.7%; ACH nets a flat $1.60 per transaction (on a $3,800 invoice, $1.60 is a 0.04% effective take rate). Effective blended take rate on commercial volume $\approx$ $\mathbf{0.08\%}$.

#### The Realistic Range (Month 24 Run-Rate): * Bear Case ($3.2M ARR): Commercial customers reject card fees entirely, sticking strictly to ACH/checks; residential adoption stalls at 45% due to surcharge pushback. * Base Case ($6.1M ARR): 60% overall adoption. Residential volume yields a 0.50% blended take ($2.03B $\times$ 60% adoption $\times$ 0.50% = $6.09M). Commercial yields minimal flat-fee ACH revenue. * Bull Case ($9.8M ARR):3 70% adoption matches the CEO's target, and we successfully introduce a B2B "accelerated payout" fee (drawing on the 9 customers out of 22 who said they would pay a fee to solve their $200k payroll float).

> Result: The realistic revenue run-rate at Month 24 is $3.2M to $9.8M, falling drastically short of the $24M target.

---

3. Falsification Criteria and a Six-Week Test

#### What would prove us wrong? If a randomized cohort of commercial-heavy customers willingly adopts card payments at >20% volume despite a 2.9% surcharge, or if property managers accept automated card-on-file billing for invoices over $3,000, our commercial pessimism is unfounded.

#### The Six-Week Test ($15k budget, 1 squad for 6 weeks): * The Experiment: Launch a targeted "Fast-Pay Commercial Portal" pilot with 30 mid-market commercial customers currently handling invoices between $2,000 and $5,000. * The Mechanics: Offer them an explicit choice: continue standard net-45 terms via free ACH, or use a discounted commercial card rate (e.g., split-surcharge or 1.9% + $0.30 via a specialized B2B interchange tier)2 to get paid in 48 hours. * Success Metric: Do at least 25% of their commercial payors opt-in to the digital card rail over 6 weeks? If not, the commercial card thesis is dead.

---

4. What We Should Do Instead

Pausing the scheduling rewrite to chase an illusory $24M payments figure is a strategic trap. The scheduling rewrite is our defensive moat: losing it costs us $1.9M in high-value churn annually from customers with >50 technicians.

#### Recommended Pivot: 1. Protect Core Engineering: Keep the scheduling rewrite on track. Do not cannibalize the core product teams for a pivot that math shows yields ~$6M, not $24M. Allocate one dedicated squad (not three) to productize payments iteratively. 2. Pivot the Value Proposition: Stop trying to monetize residential card transactions where competition (Tradesly, Stripe) is fierce. Instead, solve the actual pain point identified in customer interviews: commercial cash flow float. 3. Build B2B Working Capital Solutions: Build features that help trades businesses bridge the net-45 gap (e.g., instant payout/invoice factoring models embedded in Fieldline Pay) rather than relying purely on interchange take-rates.

---

Given that our next Series C raise is in 14 months, how comfortable is the board with a realistic payments ARR of $6M built on workflow efficiency, versus an aggressive $24M gamble that risks our core enterprise retention?

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticStart again

    Source: CFO's revenue modelAnnual invoiced value per customer: $2.1M (total invoiced $4.83B ÷ 2,300).

    What we’d changeDrop the mean-versus-median argument: customers times mean invoicing correctly recovers total volume. The real flaw is applying a card-only 0.7% margin to all of it.

  2. 2Invented evidenceVerify or remove the claimSubstantial rework

    What we’d changeRemove the discounted tier: nothing in the term sheet offers it. Design the test around economics we have, and make it test the $24M claim.

  3. 3Numbers wrongRedo the arithmeticTargeted repair

    What we’d changeShow the working for the bear and bull cases. Only the base case is calculated.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly13%
  • passAddresses the actual decision75%
  • passRespects explicit constraints28%
  • passIdentifies material uncertainty74%
  • partialAvoids unsupported claims28%
  • passProduces the required deliverable55%
  • passFinds the load-bearing assumption52%
  • passRe-estimates the revenue correctly24%
  • passSays what would prove it wrong100%
  • passReshapes the bet from the evidence100%
  • passTigers, not paper tigers65%
Run
Run
#1
Time to output
14 s
Submitted
25 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyMixedRightRight
Gemini 3.8 Flash · API

The memo uses several unsupported current-state assumptions, especially the $1.43B locked GMV figure and the accessible GMV/payment-method splits.

GPT-6 Astra · ChatGPT

Every factual statement about the current situation is drawn directly from the supplied context or follows from it by arithmetic; nothing about the current situation is invented.

Gemini 3.5 Flash-Lite · Gemini

Every factual statement about the current situation is directly taken from the supplied context or derived by straightforward arithmetic, with no inventions.

Respects explicit constraintsWrongRightRight
Gemini 3.8 Flash · API

The memo appears to exceed the 1,200-word limit and includes heavy tables and formatting that make it less usable as a concise exec pre-read.

GPT-6 Astra · ChatGPT

The output is a memo under 1,200 words, addresses all four required points, and respects the reader and the decision context; it does not propose anything that would violate an explicit constraint.

Gemini 3.5 Flash-Lite · Gemini

The deliverable is a memo under 1200 words, addressed to the specified readers, and respects the four numbered requirements.

Avoids unsupported claimsWrongRightRight
Gemini 3.8 Flash · API

It presents several estimates and causes as established facts, including the locked-GMV amount, accessible mix, and that pilot adjustment lands directly in the $5M–$7M range.

GPT-6 Astra · ChatGPT

Scenarios are clearly labeled as such, assumptions are stated explicitly, and no forecast or cause is presented as established fact without evidence.

Gemini 3.5 Flash-Lite · Gemini

Interpretations such as "structural distortion" are clearly argued from the data, and forecasts are presented as cases with assumptions, not as established fact.

Produces the required deliverableMixedRightRight
Gemini 3.8 Flash · API

It is a memo with the required sections, but it is too long and contains material arithmetic assumptions that would need rechecking before use.

GPT-6 Astra · ChatGPT

A complete memo in the right format, word count within limit, addressed to the CPO and exec team, with sections that answer each prompt point; the reader can act on it directly.

Gemini 3.5 Flash-Lite · Gemini

The memo follows the requested format, is well within the word limit, and contains all four required elements in a way the executive audience could act on with light edits.

Re-estimates the revenue correctlyMixedRightRight
Gemini 3.8 Flash · API

It lands in the right range, but the working relies on unsupported locked-GMV and payment-mix assumptions rather than a clean recalculation from supplied data.

GPT-6 Astra · ChatGPT

It recalculates revenue as $3.3M–$7.2M using the payment mix and unit economics, shows working, notes the pilot-based upper bound of ~$9M, and explains why even that is an overestimate.

Gemini 3.5 Flash-Lite · Gemini

It segments customers by residential/commercial, applies realistic take rates from the pilot and unit economics, shows its arithmetic, and lands in a $3.2 M–$9.8 M range, far below $24 M, noting that the pilot's residential-heavy mix would give a higher bound.

All got right 6

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

It commits early to right-sizing the bet, funding the scheduling rewrite, limiting Fieldline Pay to residential, and testing commercial cash-flow monetization.

GPT-6 Astra · ChatGPT

The memo commits early to 'do not approve three squads', frames it for the executive team, and states exactly what evidence would change that (a defensible route to $24M from a volume-weighted model).

Gemini 3.5 Flash-Lite · Gemini

The memo commits early to challenging the plan, recommends scaling back to one squad and pivoting to commercial cash-flow solutions, and specifies a condition (≥25% commercial card opt-in) that would change its assessment, all framed for the CEO, CFO and Head of Sales.

Identifies material uncertaintyRightRightRight
Gemini 3.8 Flash · API

It names the key unknowns—commercial card acceptance, contract breakability, and ACH basis-point pricing—and gives six-week thresholds.

GPT-6 Astra · ChatGPT

It names the key unknowns (contractual lock-in, actual captured volume, commercial card uptake), bounds them with scenario analysis, and describes a test that would resolve them and change the call.

Gemini 3.5 Flash-Lite · Gemini

The memo pinpoints commercial card adoption as the critical unknown, bounds the revenue range, and proposes a concrete six-week test with a clear threshold that would resolve whether its commercial pessimism is wrong.

Finds the load-bearing assumptionRightRightRight
Gemini 3.8 Flash · API

It correctly makes the 0.7% blended take rate applied to all invoice value the central challenge and supports it with payment-mix and unit-economics figures.

GPT-6 Astra · ChatGPT

The central challenge is the mismatch between the 0.7% card-based take rate and the actual payment mix, substantiated with ACH/card percentages and unit economics from the pack.

Gemini 3.5 Flash-Lite · Gemini

The memo makes the mismatch between the assumed 0.7% blended net take rate and the actual payment-method mix (mostly ACH/check for commercial) its central challenge, exactly as the strong answer expects.

Says what would prove it wrongRightRightRight
Gemini 3.8 Flash · API

It specifies a cheap six-week commercial pay-by-link test with concrete thresholds that would change the conclusion.

GPT-6 Astra · ChatGPT

It proposes a concrete six-week cohort test with specific measurements and a clear threshold: a volume-weighted model that shows a defensible path to $24M before expanding investment.

Gemini 3.5 Flash-Lite · Gemini

It describes a six-week test with 30 mid-market commercial customers, a discounted card option, and the explicit threshold 'at least 25% of their commercial payors opt-in' that would prove its skepticism wrong.

Reshapes the bet from the evidenceRightRightRight
Gemini 3.8 Flash · API

It keeps residential card payments and pivots the commercial opportunity toward faster payout/working capital, grounded in interview evidence.

GPT-6 Astra · ChatGPT

It recommends a narrower residential payments offering, a separate test for commercial faster-payment demand, and keeping the scheduling rewrite funded, directly grounded in the residential-mix evidence and commercial cash-flow pain.

Gemini 3.5 Flash-Lite · Gemini

It leverages the evidence that residential card payments work and that commercial cash-flow pain is acute, recommending a smaller investment, a pivot to faster payout/factoring solutions, and protecting the scheduling rewrite.

Tigers, not paper tigersRightRightRight
Gemini 3.8 Flash · API

It ranks the take-rate and contract-access issues as material, treats Tradesly as a residential-mix comparison, and surfaces the churn/valuation risk of pausing the scheduling rewrite.

GPT-6 Astra · ChatGPT

It dismisses the Tradesly comparison as irrelevant for a different mix, elevates the take-rate mismatch as the central risk, and surfaces the hidden cost of pausing the scheduling rewrite ($1.9M lost ARR) that the CEO's plan avoided.

Gemini 3.5 Flash-Lite · Gemini

It distinguishes the real killer (commercial card adoption) from execution risk, dismisses the residential-only comparison to Tradesly as misleading because of mix, and calls out the unspoken cost of pausing the scheduling rewrite ($1.9 M churn).

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.095.82None
2GPT-6 AstrawithChatGPT93.295.82None
3GPT-6 LunawithAPI97.791.72None
4Sonnet 5.5withAPI90.975.02None
5Gemini 3.5 Flash-LitewithGemini88.675.021 capped
6Opus 5.5withClaude93.291.721 capped
7Gemini 3.8 FlashwithAPI93.237.521 capped

About the task

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Separates the risks that could kill it from the ones that only look scary
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill · Staff level: a company bet with a long evidence pack

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.