Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  2. Identifies material uncertainty100% pass
    It names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  3. Uses the interviews faithfully100% pass
    All quotes are accurate and correctly attributed to A05, A08, A04, A07, and A12.
    GPT-6 Luna · API · An AI SDR for small agencies

Where it slips

  1. A cheap test that can actually read out64% pass
    The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.
    GPT-6.1 Sol · API · An AI SDR for small agencies
  2. Re-estimates the revenue correctly71% pass
    It shows working but lands on a central $2.3M and $1M–$6M range, not roughly the $5–10M range required by the grading rubric, and its $9.1M pilot upper bound is not used as the main re-estimate.
    Sonnet 5.5 · API · The CEO's embedded-payments bet
  3. Avoids unsupported claims71% pass
    The memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
    Opus 5.5 · Claude · The CEO's embedded-payments bet

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Fieldline. Our CEO wants to commit three squads for three quarters to Fieldline Pay, embedded payments for our customers' invoices, and has told the board it could add $24M of annual revenue within two years. Priya Raman, our CPO, has asked you to write the strongest challenge to the plan as a pre-read for next week's exec offsite. The CEO, CFO and Head of Sales will all read it. Write a memo of no more than 1,200 words that: 1. Names the single assumption the plan most depends on that the evidence does not support, and shows why, using the numbers in the pack. 2. Re-estimates the revenue from the supplied data, showing your working, as a range. 3. Says what we would need to see to be proved wrong, and the cheapest test that would show it within six weeks. 4. Says what, if anything, we should do instead or how the bet should change. The pack below is everything we have. Some of it matters more than the rest.

What the model was given11 items: About Fieldline, CEO's memo to the exec team (excerpt), CFO's revenue model, Invoicing data (last 12 months, all customers), How invoices are paid today, Payments partner term sheet (unit economics), Pilot (4 months), Sales notes on the largest accounts, Customer interviews (22 customers, last quarter), Competitor, Engineering
About FieldlineField-service software for trades businesses (plumbing, HVAC, electrical): scheduling, dispatch, quotes and invoicing. 2,300 customers, $41.0M ARR, average $17,800 per customer. Series C; the next raise is planned in about 14 months.
CEO's memo to the exec team (excerpt)“Every invoice our customers send is money we don't touch. Our customers invoiced $4.83B last year through Fieldline. If we process those payments, we become part of how they get paid, not just how they schedule. The model is simple: 70% of customers adopt within 18 months, the average customer invoices $2.1M a year, and we keep a 0.7% blended net take. That is $24M of new annual revenue by month 24, more than half our current ARR, and it makes the next raise a very different conversation. Tradesly has shown it works: payments are now 22% of their revenue. I've told the board I believe payments can be 35% of our revenue by 2028. I want three squads on this from next quarter, which means pausing the scheduling rewrite.”
CFO's revenue modelCustomers: 2,300. Adoption by month 18: 70% (1,610 customers). Annual invoiced value per customer: $2.1M (total invoiced $4.83B ÷ 2,300). Blended net take rate: 0.7%. Month-24 revenue run-rate: 1,610 × $2.1M × 0.7% = $23.7M. CFO's note: “Adoption and take rate are the CEO's assumptions. I haven't stress-tested them.”
Invoicing data (last 12 months, all customers)Total invoiced value: $4.83B. Mean per customer: $2.1M. Median per customer: $640k. The largest 5% of customers (115) account for 48% of invoiced value ($2.32B). By job type, residential jobs are 42% of invoiced value and commercial jobs (property managers, facilities contracts) are 58%.
How invoices are paid todayFrom the 690 customers (30%) who record the payment method in Fieldline, by share of invoice value: card 19%, ACH/bank transfer 44%, check 31%, cash 6%. Card share is 38% of residential invoice value and 5% of commercial. Average commercial invoice: $3,800; average residential invoice: $410. Commercial clients pay on net-45 or net-60 terms; the average commercial invoice is paid 52 days after it is sent.
Payments partner term sheet (unit economics)Card: the customer is charged 2.9% + $0.30 per payment; our all-in cost (interchange, network, partner fee) is about 2.2%, so we net about 0.7% of card value. ACH: the customer is charged a flat $2.00 per payment; our cost is $0.40. Checks and cash earn nothing unless the payer switches to card or ACH. The partner handles licensing, KYC and risk; they have approved our application.
Pilot (4 months)62 customers invited, 38 adopted (61%). Pilot customers' invoice value is 64% residential (the customer base is 42%). With pay-by-link on every invoice, card share of invoices paid through Fieldline Pay rose to 71% for residential and 6% for commercial. Net payments revenue, annualised: $214k across the 38 customers ($5,630 per customer per year).
Sales notes on the largest accountsOf the 115 largest customers, 71 have multi-year contracts with an existing payment processor, most running to 2028. Head of Sales, in Slack: “None of the big ones will move processors before their contracts end, and their property-manager clients will not pay 2.9% on a $3,800 invoice. They'll pay by ACH or check like they always have.”
Customer interviews (22 customers, last quarter)17 of 22 named getting paid on commercial jobs as their biggest cash problem (“I'm floating $200k of payroll while property managers sit on invoices for two months”). 9 said they would pay a fee to be paid faster. 6 said they won't offer card payment because clients fight the surcharge. Of the 14 customers with mostly commercial work, 11 said their clients require ACH or check.
CompetitorTradesly launched embedded payments in 2025 and says payments are now 22% of its revenue. Its published case studies are almost all home-services businesses; it describes its customers as about 80% residential.
EngineeringThe plan needs three squads for three quarters. The scheduling rewrite it would pause is the top-cited reason for churn among customers with more than 50 technicians (lost ARR from that segment last year: $1.9M).
What a strong answer doesThe answer key the graders mark against

Names the load-bearing assumption as the 0.7% blended net take: that is the card margin applied to all invoiced value, but only about 19% of value is paid by card (5% of commercial, which is 58% of value), ACH nets about 0.04% on a $3,800 invoice, and checks and cash earn nothing. Re-estimates with the supplied data: even at the pilot's improved mix (about a third of value on card), the blended take is roughly 0.2–0.25%, so the run-rate is about $5–9M, not $24M (the pilot's $5,630 per customer × 1,610 = $9.1M is an upper bound because the pilot was residential-heavy; the 71 locked-in large accounts, with 48% of value in the top 5%, cut it further). Notes that $24M needs essentially all invoice value on card. Treats Tradesly as a residential-mix result, not proof, and engineering capacity, licensing and competition as secondary. Keeps what the evidence supports: card payments work for residential invoices, and the strongest pull is commercial cash flow (17 of 22), so it reshapes the bet toward faster payout or financing on commercial invoices, which earns on the volume the processing model can't. Proposes a cheap test with a threshold that would prove the challenge wrong (for example, commercial card share well above 20% with pay-by-link in a six-week cohort, or uptake of paid instant payout), and is written for a room that includes the CEO: firm on the numbers, not personal.

Critical failures (cap the score at 40)
  • Accepts the $24M projection or the 0.7% blended net take rate as sound
  • Invents market sizes, competitor figures or customer statistics that are not in the pack
Case

Staff-level · v1.7 · synthetic · B2B SaaS, fintech, strategy, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 11 checks passed

The pre-read is tightly argued, numerically grounded, and directly usable by a senior PM; only minor phrasing tweaks would be needed to align the bull case with pure payments revenue if desired.

Output

To: Priya Raman (CPO), CEO, CFO, Head of Sales From: Staff Product Manager, Fieldline Date: September 25, 2026 Subject: Pre-read: Strategic Challenge to the Fieldline Pay Plan ($24M Revenue Target)

---

Executive Summary

Committing three squads for three quarters to Fieldline Pay to capture $24M in new revenue within two years is a high-risk bet that relies on a structural distortion in our financial model.

The plan’s core vulnerability is not execution; it is a misallocation of our customer base's economic reality. By using the mean invoice value ($2.1M) rather than the median ($640k), the model assumes our average customer mirrors our largest commercial accounts1, while simultaneously ignoring entrenched enterprise contracts and commercial payment norms.

Executing this plan requires pausing the scheduling rewrite, which risks $1.9M in annual churn from our largest, most valuable accounts. Below is the evidence-based challenge, a revised revenue range, an inexpensive six-week test, and a recommended alternative path.

---

1. The Dependent Assumption and Why the Evidence Fails It

> The Plan's Single Dependent Assumption: That we can achieve a $23.7M–$24M revenue run-rate by month 24 by applying a blended 0.7% net take rate across a homogenous $2.1M annual invoice volume per adopted customer.

The evidence flatly refutes this assumption in three ways:

  • The Mean vs. Median Distortion: The CEO's model relies on a mean invoice value of $2.1M. However, the median customer invoices just $640k. Our distribution is heavily skewed: the top 5% of customers (115 accounts) drive 48% ($2.32B) of our total invoiced value.
  • Enterprise Lock-in: Of those 115 largest customers, 71 are locked into multi-year contracts with existing payment processors running through 2028. As the Head of Sales noted, none will migrate early.
  • The Commercial Wall: Commercial jobs comprise 58% of our invoiced value ($2.80B). The average commercial invoice is $3,800 and is paid on net-45/net-60 terms (averaging 52 days). Property managers and facility contractors will not pay a 2.9% card surcharge on a $3,800 invoice. They pay via ACH ($2.00 flat fee, netting us $1.60) or check ($0 revenue), exactly as they always have.

Applying a 0.7% card take-rate model to commercial volume is fundamentally flawed because commercial clients do not use cards (card share is only 5% for commercial, vs. 38% for residential).

---

2. Re-Estimated Revenue Range (Working Included)

To model realistic revenue, we must segment our base by customer type, recognizing that residential and commercial segments have radically different payment behaviors and take rates.

#### Step-by-Step Working: 1. Customer Base: 2,300 total customers. 2. Adoption Rate: Pilot adoption was 61% (close to the 70% target). Let's model a realistic 60% adoption rate by month 24 = $1,380$ adopting customers. 3. Segmentation Split: Based on overall volume, 42% of value is residential ($2.03B) and 58% is commercial ($2.80B). Total invoiced value = $4.83B. * Total Residential Invoiced: $2.03B ($\approx$ $882k$ per customer across 2,300) * Total Commercial Invoiced: $2.80B ($\approx$ $1.22M$ per customer across 2,300) 4. Take Rates & Behavior: * Residential: 71% card adoption via pay-by-link. Net take rate on card is 0.7%. ACH/Check take rate is near zero (flat $1.60 net on ACH, negligible volume). Effective blended take rate on residential volume $\approx$ $0.7\% \times 71\% \approx \mathbf{0.50\%}$. * Commercial: 6% card adoption, 94% ACH/Check/Terms. Card take rate is 0.7%; ACH nets a flat $1.60 per transaction (on a $3,800 invoice, $1.60 is a 0.04% effective take rate). Effective blended take rate on commercial volume $\approx$ $\mathbf{0.08\%}$.

#### The Realistic Range (Month 24 Run-Rate): * Bear Case ($3.2M ARR): Commercial customers reject card fees entirely, sticking strictly to ACH/checks; residential adoption stalls at 45% due to surcharge pushback. * Base Case ($6.1M ARR): 60% overall adoption. Residential volume yields a 0.50% blended take ($2.03B $\times$ 60% adoption $\times$ 0.50% = $6.09M). Commercial yields minimal flat-fee ACH revenue. * Bull Case ($9.8M ARR):3 70% adoption matches the CEO's target, and we successfully introduce a B2B "accelerated payout" fee (drawing on the 9 customers out of 22 who said they would pay a fee to solve their $200k payroll float).

> Result: The realistic revenue run-rate at Month 24 is $3.2M to $9.8M, falling drastically short of the $24M target.

---

3. Falsification Criteria and a Six-Week Test

#### What would prove us wrong? If a randomized cohort of commercial-heavy customers willingly adopts card payments at >20% volume despite a 2.9% surcharge, or if property managers accept automated card-on-file billing for invoices over $3,000, our commercial pessimism is unfounded.

#### The Six-Week Test ($15k budget, 1 squad for 6 weeks): * The Experiment: Launch a targeted "Fast-Pay Commercial Portal" pilot with 30 mid-market commercial customers currently handling invoices between $2,000 and $5,000. * The Mechanics: Offer them an explicit choice: continue standard net-45 terms via free ACH, or use a discounted commercial card rate (e.g., split-surcharge or 1.9% + $0.30 via a specialized B2B interchange tier)2 to get paid in 48 hours. * Success Metric: Do at least 25% of their commercial payors opt-in to the digital card rail over 6 weeks? If not, the commercial card thesis is dead.

---

4. What We Should Do Instead

Pausing the scheduling rewrite to chase an illusory $24M payments figure is a strategic trap. The scheduling rewrite is our defensive moat: losing it costs us $1.9M in high-value churn annually from customers with >50 technicians.

#### Recommended Pivot: 1. Protect Core Engineering: Keep the scheduling rewrite on track. Do not cannibalize the core product teams for a pivot that math shows yields ~$6M, not $24M. Allocate one dedicated squad (not three) to productize payments iteratively. 2. Pivot the Value Proposition: Stop trying to monetize residential card transactions where competition (Tradesly, Stripe) is fierce. Instead, solve the actual pain point identified in customer interviews: commercial cash flow float. 3. Build B2B Working Capital Solutions: Build features that help trades businesses bridge the net-45 gap (e.g., instant payout/invoice factoring models embedded in Fieldline Pay) rather than relying purely on interchange take-rates.

---

Given that our next Series C raise is in 14 months, how comfortable is the board with a realistic payments ARR of $6M built on workflow efficiency, versus an aggressive $24M gamble that risks our core enterprise retention?

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticStart again

    Source: CFO's revenue modelAnnual invoiced value per customer: $2.1M (total invoiced $4.83B ÷ 2,300).

    What we’d changeDrop the mean-versus-median argument: customers times mean invoicing correctly recovers total volume. The real flaw is applying a card-only 0.7% margin to all of it.

  2. 2Invented evidenceVerify or remove the claimSubstantial rework

    What we’d changeRemove the discounted tier: nothing in the term sheet offers it. Design the test around economics we have, and make it test the $24M claim.

  3. 3Numbers wrongRedo the arithmeticTargeted repair

    What we’d changeShow the working for the bear and bull cases. Only the base case is calculated.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly13%
  • passAddresses the actual decision75%
  • passRespects explicit constraints28%
  • passIdentifies material uncertainty74%
  • partialAvoids unsupported claims28%
  • passProduces the required deliverable55%
  • passFinds the load-bearing assumption52%
  • passRe-estimates the revenue correctly24%
  • passSays what would prove it wrong100%
  • passReshapes the bet from the evidence100%
  • passTigers, not paper tigers65%
Run
Run
#1
Time to output
14 s
Submitted
25 Sept 2026

Usable with a quick edit10 of 11 checks passed

A PM could use it with light edits, mainly tightening the revenue range to align with the expected $5–10M upper-bound framing. The evidence and reasoning are otherwise strong and actionable.

Output

To: Priya Raman (for exec offsite pre-read) From: Staff PM Re: The strongest case against committing three squads for three quarters to Fieldline Pay

Bottom line

Payments is a real opportunity, but the $24M is not supported by our own data. Re-estimated from the pack, month-24 run-rate is roughly $1M–$6M, central case about $2.3M. Even the most generous extrapolation of the pilot gives about $9M. The plan would pause the scheduling rewrite, which is tied to $1.9M of churned ARR a year, to fund a payments line that at its central case earns about the same amount after nine squad-quarters. I recommend a smaller, staged bet.

1. The assumption the plan depends on most

The plan assumes a dollar of invoice value is worth 0.7% to us. It is not, because 0.7% is the margin on card payments only.

The term sheet says we net about 0.7% on card value, but ACH nets $1.60 per payment ($2.00 charged, $0.40 cost), and checks and cash earn nothing. The CEO and CFO models apply the card margin to all $4.83B of invoices. What matters is how those invoices get paid.

  • Commercial jobs are 58% of invoiced value, and only 5% of that is paid by card. The average commercial invoice is $3,800, so ACH nets $1.60 / $3,800 = 0.04%.
  • Residential jobs are 42% of value, and the pilot lifted card share to 71%.

Yield per dollar of volume that actually flows through Fieldline Pay, using pilot card shares and assuming the non-card remainder moves by ACH:

CardACHYield
Residential ($410 avg)71% × 0.70% = 0.50%29% × ($1.60/$410 = 0.39%) = 0.11%0.61%
Commercial ($3,800 avg)6% × 0.70% = 0.04%94% × 0.042% = 0.04%0.08%
Whole base (42/58 mix)0.30%

Even if every invoice dollar went through Fieldline Pay, the blended yield would be 0.30%, less than half the modelled 0.7%.

The evidence is consistent with this: - 11 of 14 mostly-commercial customers say their clients require ACH or check. - 6 of 22 won't offer card at all. - Head of Sales expects property managers to refuse 2.9% on a $3,800 invoice.

Pilot revenue is $5,630 per adopting customer, against the $14,700 the CEO's model implies ($2.1M × 0.7%). That is 38% of the claim, from a pilot that was 64% residential, against 42% for the base. This is the best case for the mix.

Two supporting errors compound this: - The $2.1M is a mean. The median customer invoices $640k. The top 115 customers carry 48% of volume, and 71 of them are on multi-year processor contracts, mostly to 2028. Sales expects none to move. - Tradesly is not a comparable. It is about 80% residential, and we are 42%. Its 22% revenue share says little about a base that is mostly commercial by value.

2. Re-estimate of month-24 run-rate

Formula: addressable volume × adoption × share of adopters' invoice value routed through Pay × blended yield.

Addressable volume. I remove the locked-in accounts, assuming the 71 are typical of the top 115: 71/115 × $2.32B ≈ $1.43B. That leaves $3.40B.

LowBaseHigh
Adoption (by volume)30%45%61% (pilot rate)
Share of adopter volume routed30%50%70%
Blended yield0.25%0.30%0.40%
Run-rate$0.8M$2.3M$5.8M
  • Adoption. The pilot's 61% came from invited customers over four months, so I treat it as a ceiling.
  • Routing. The CEO's model implicitly assumes 100%. Customers who adopt still get paid by check and cash on many invoices, and payers choose the method.
  • Yield. Base is the whole-base mix. High approaches the pilot's residential-heavy 0.42%. Low assumes the locked-in accounts are commercial-heavy, so the remainder skews slightly more residential and the low case is harsher than it needs to be.

Cross-check against the pilot. The base case implies about $2,300 per adopting customer, against $5,630 in a residential-skewed, hand-picked pilot. That is plausible. The most generous version applies the pilot's $5,630 to 1,610 customers, giving $9.1M. That ignores the lock-ins and the pilot's mix, and it is still 38% of the claim.

Implication for the board. At $2–6M, payments would be roughly 5–12% of revenue, not 35% by 2028. Reaching $24M would need about a 3% yield on the volume in my base case, which is not available on this term sheet.

Caveats: - Payment-method data comes from only 690 self-reporting customers (30%). - The pilot ran four months. - I have not modelled ramp, disputes or chargebacks, or upside from a paid fast-payout product. That last one is the most interesting missing piece.

3. What would prove me wrong, and the cheapest test

To vindicate the plan, we would need to show that commercial volume can yield at least ~0.5% net, with at least 60% of adopters' volume routed through Pay. That would put the total near $15M or more. On a $3,800 invoice, 0.5% is about $19 net per invoice. That requires either much higher commercial card share, or a fee people will pay, such as faster payout.

A six-week constraint: commercial invoices take 52 days on average to be paid, longer than the test window. We cannot wait to observe commercial cash-in, so the test must use leading indicators and data we already hold.

The test (roughly one PM, one analyst and one engineer, using existing pilot plumbing): 1. Week 1: re-cut the existing pilot. Split the 38 customers by commercial share and compute yield per dollar of volume for the commercial-heavy ones. Their older invoices are already paid, so this costs nothing and tells us quickly whether 0.08% is right. 2. Weeks 1–6: a commercial-heavy cohort. Recruit about 20 mostly-commercial, non-locked-in customers, drawn from the 11 who said their clients require ACH or check. Turn on pay-by-link and measure the share of payers who choose card and the share of invoice value routed via link. 3. Weeks 2–6: price a fast-pay offer. Offer these customers accelerated payout, funded by a fee, as a real priced offer with sign-up. Nine of 22 said they would pay to be paid faster, and this is the only route I see to meaningful commercial yield. Measure conversion at the stated price.

Thresholds, set now: commercial card share above 25%, or fast-pay uptake implying at least 0.3% net on routed commercial volume. Below both, the base case stands and the bet should be sized to it.

4. What we should do instead

Don't commit three squads for three quarters, and don't pause the scheduling rewrite. The rewrite addresses the top-cited churn reason in the >50-technician segment, which lost $1.9M ARR last year. Those are also the accounts with the most volume and the ones payments can't reach. Cutting the rewrite risks retention at the top of the base to chase payments revenue the top can't deliver.

A staged alternative: 1. Fund one squad for two quarters, focused on residential-heavy and small customers, where the pilot works (71% card share, ~0.6% yield). Ship pay-by-link, card and ACH broadly, and track net revenue per adopting customer and routed share. 2. Run the six-week commercial test above. Its result decides whether a second and third squad get funded for a commercial cash-flow product (fast payout or early pay). That product fits the real pain: 17 of 22 customers named getting paid on commercial jobs as their biggest cash problem. 3. Set gates. Add squads only when the test or the first quarter's data shows blended yield above 0.4% on routed volume. Revisit the locked-in accounts as contracts approach expiry in 2028. 4. Reset the board narrative from "$24M / 35% of revenue" to a $2–6M range with a clear path to more if commercial monetisation works. For the raise in about 14 months, a measured payments attach rate and net revenue per customer will be more credible to investors than a projection that unravels in diligence.

The CEO's instinct that payments could make Fieldline stickier is sound, and the pilot's 61% adoption is a good sign. But the current plan assumes we can monetise the volume we don't touch at the rate we monetise the volume we do, and our own numbers say otherwise.

Grades and run details

Decision model 91 · LLM judge 11 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly49%
  • passAddresses the actual decision100%
  • passRespects explicit constraints29%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims32%
  • passProduces the required deliverable89%
  • passFinds the load-bearing assumption100%
  • failRe-estimates the revenue correctly27%
  • passSays what would prove it wrong100%
  • passReshapes the bet from the evidence100%
  • passTigers, not paper tigers99%
Run
Run
#1
API response time
70 s
Submitted
29 Sept 2026

Needs more than a quick edit10 of 11 checks passed

The memo is factually strong and gives a clear go/no-go call, but its 'instead' is generic and misses the commercial cash-flow/faster-payout opportunity from the interviews. A PM would need to add that alternative reasoning rather than just edit wording.

Output

To: Priya Raman Subject: Fieldline Pay: challenge the $24M case before committing three squads

The strongest case against the current plan is not that embedded payments cannot work. It is that the $24M forecast treats invoice volume as if it were payment volume earning a card-like margin.

The critical unsupported assumption: 70% of customers will adopt and route enough of their invoices through Fieldline Pay to earn roughly 0.7% on that value. The 0.7% net take is specifically the card margin. ACH earns $1.60 per payment; checks and cash earn nothing. The CFO’s arithmetic is correct only if essentially every dollar from every adopting customer earns the card margin.

That is a poor fit with our mix. Commercial work is 58% of invoiced value; commercial clients typically require ACH or check, pay after 52 days, and may not accept a card surcharge on a $3,800 invoice. Eleven of 14 interviewed mostly-commercial customers said their clients require ACH or check. By contrast, the pilot’s customers were 64% residential, against 42% for Fieldline overall. And 71 of the 115 largest customers have processor contracts running mostly to 2028; the largest 5% of customers account for 48% of invoiced value. The pilot’s 61% adoption is encouraging, but it does not establish broad adoption or broad monetization. Tradesly’s 22% revenue share is also a weak comparator: its customer base is about 80% residential.

A more defensible annualized revenue range

The following estimates are run-rate scenarios, not a forecast. They use the supplied 61% pilot adoption rate and 70% plan assumption as bounds, and make explicit that the take depends on payment type.

Card revenue: Current recorded payment methods show 19% of invoice value paid by card. In the pilot, the reported card shares imply a 33.3% value-weighted card share if applied to our mix: 42% residential × 71% + 58% commercial × 6% = 33.3%.

At a 0.7% net card take:

  • Lower case: $4.83B × 61% adoption × 19% card share × 0.7% = $3.9M
  • Upper case: $4.83B × 70% adoption × 33.3% card share × 0.7% = $7.9M

The upper case assumes pilot card behavior transfers to the full customer mix, despite the pilot’s residential skew.

ACH revenue: The recorded-method sample reports ACH at 44% of invoice value, but value share does not tell us the number of $2 ACH payments—and the pilot does not report ACH usage. As an illustrative bound, if that 44% share transferred to adopters, ACH revenue would be about $0.5M–$2.1M after the $0.40 cost: the lower end assumes ACH dollars are mostly $3,800 commercial invoices; the upper end allows the pilot’s card mix and available non-card residential volume. This is uncertain upside, not a reliable forecast.

Together, that supports roughly $4.5M–$10M of annualized revenue, before accounting for contract lock-in or slower adoption among large accounts.1 The range is still conditional on transferring today’s payment mix and pilot behavior to adopters. As a cross-check, the pilot’s $214K annualized revenue across 38 customers is $5,630 per customer; scaling that to 61%–70% adoption gives $7.9M–$9.1M. That sits within the range, but is not a clean base case because the pilot skewed residential.

The $24M case requires $14,700 per adopting customer per year ($2.1M × 0.7%). That is about 2.6× the pilot’s $5,630, while assuming adoption and monetization across a materially more commercial customer base.

What would change my mind—and the cheapest six-week test

I would support a larger bet if a representative cohort showed both credible adoption and enough net revenue per adopter to approach the $14,700 annualized level the board case requires—especially among commercial customers and larger accounts not locked into existing processors. A high signup rate alone would not prove the economics.

Run a six-week, no-build test using the existing pay-by-link and partner setup. Recruit a small, stratified cohort across residential and commercial customers, including larger unlocked accounts. Track offers accepted, invoice value actually routed, card versus ACH mix, net revenue after costs, and payment timing. Report results by segment; do not annualize a handful of unusually large or residential-heavy customers as if they were representative. Include a holdout or compare with each customer’s recent invoices to distinguish new payment volume from existing behavior.

Recommendation

Do not commit three squads for three quarters or pause the scheduling rewrite on the current evidence. The rewrite addresses the top-cited churn reason among customers with more than 50 technicians, a segment that lost $1.9M ARR last year. Run the six-week test with a small team and existing capabilities, then make a stage-gated decision. If commercial routing and net revenue per adopter are strong, fund the next increment. If not, preserve the scheduling investment and narrow Pay to the segments where the data shows customers can—and will—use it.

What a PM had to fix

  1. 1OtherFix itTargeted repair

    What we’d changeThese assumptions produce an illustrative $4.5M–$10M envelope, not a validated revenue range. They assume customer adoption translates proportionally into invoice volume and the modelled payments route through Pay. Rebuild by segment: accessible invoice value × volume-weighted adoption × routing share × payment-method economics, excluding contracted volume.

Grades and run details

Decision model 95 · LLM judge 10 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly14%
  • passAddresses the actual decision93%
  • passRespects explicit constraints55%
  • passIdentifies material uncertainty97%
  • passAvoids unsupported claims42%
  • passProduces the required deliverable53%
  • passFinds the load-bearing assumption98%
  • passRe-estimates the revenue correctly83%
  • passSays what would prove it wrong78%
  • partialReshapes the bet from the evidence53%
  • passTigers, not paper tigers89%
Run
Run
#1
API response time
73 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 2

Re-estimates the revenue correctlyRightWrongRight
Gemini 3.5 Flash-Lite · Gemini

It segments customers by residential/commercial, applies realistic take rates from the pilot and unit economics, shows its arithmetic, and lands in a $3.2 M–$9.8 M range, far below $24 M, noting that the pilot's residential-heavy mix would give a higher bound.

Sonnet 5.5 · API

It shows working but lands on a central $2.3M and $1M–$6M range, not roughly the $5–10M range required by the grading rubric, and its $9.1M pilot upper bound is not used as the main re-estimate.

GPT-6 Luna · API

It recalculates revenue from the payment mix and pilot, lands in the $4.5-10M range, and explains why the pilot is an upper bound.

Reshapes the bet from the evidenceRightRightWrong
Gemini 3.5 Flash-Lite · Gemini

It leverages the evidence that residential card payments work and that commercial cash-flow pain is acute, recommending a smaller investment, a pivot to faster payout/factoring solutions, and protecting the scheduling rewrite.

Sonnet 5.5 · API

It keeps residential card payments and proposes a commercial cash-flow/fast-payout alternative grounded in interview and pilot evidence.

GPT-6 Luna · API

The recommendation is a generic stage-gate/narrow-to-segments move and does not propose the specific alternative the evidence supports, such as residential card focus or commercial cash-flow/faster payout.

All got right 9

Uses the supplied evidence correctlyRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Every factual statement about the current situation is directly taken from the supplied context or derived by straightforward arithmetic, with no inventions.

Sonnet 5.5 · API

The memo’s factual statements about the current situation are drawn from the pack or follow by arithmetic, with assumptions labelled.

GPT-6 Luna · API

All facts and arithmetic trace to the supplied pack, with no invented current-state figures found.

Addresses the actual decisionRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The memo commits early to challenging the plan, recommends scaling back to one squad and pivoting to commercial cash-flow solutions, and specifies a condition (≥25% commercial card opt-in) that would change its assessment, all framed for the CEO, CFO and Head of Sales.

Sonnet 5.5 · API

It clearly recommends not committing three squads and instead funding a staged one-squad bet, with gates and a six-week test that would change the call.

GPT-6 Luna · API

The memo commits clearly to not committing now, proposes a stage-gated test, and states what would change the call.

Respects explicit constraintsRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The deliverable is a memo under 1200 words, addressed to the specified readers, and respects the four numbered requirements.

Sonnet 5.5 · API

It is a memo under 1,200 words, addressed to Priya for the exec offsite, and covers all four requested elements.

GPT-6 Luna · API

It is a memo addressed to Priya, within the requested length, and covers the four required elements.

Identifies material uncertaintyRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The memo pinpoints commercial card adoption as the critical unknown, bounds the revenue range, and proposes a concrete six-week test with a clear threshold that would resolve whether its commercial pessimism is wrong.

Sonnet 5.5 · API

It names material unknowns such as payment-method reporting coverage, pilot mix, routing share, commercial card acceptance, and fast-pay uptake, and proposes tests to resolve them.

GPT-6 Luna · API

It names payment-mix, adoption, contract lock-in, and pilot skew as unknowns and says how a six-week test would resolve them.

Avoids unsupported claimsRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Interpretations such as "structural distortion" are clearly argued from the data, and forecasts are presented as cases with assumptions, not as established fact.

Sonnet 5.5 · API

Most interpretations are labelled as assumptions or estimates; the few rhetorical characterisations do not drive the core analysis.

GPT-6 Luna · API

Estimates are labelled as scenarios/bounds, and generalizations are tied to interview and sales evidence rather than asserted as fact.

Produces the required deliverableRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The memo follows the requested format, is well within the word limit, and contains all four required elements in a way the executive audience could act on with light edits.

Sonnet 5.5 · API

The deliverable is complete, usable, and written for the named executive audience within the requested length.

GPT-6 Luna · API

The memo is a complete, usable pre-read for the exec audience with a clear recommendation and supporting analysis.

Finds the load-bearing assumptionRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The memo makes the mismatch between the assumed 0.7% blended net take rate and the actual payment-method mix (mostly ACH/check for commercial) its central challenge, exactly as the strong answer expects.

Sonnet 5.5 · API

The central challenge is explicitly the 0.7% blended take rate being applied to all invoice value despite the actual payment mix and unit economics.

GPT-6 Luna · API

It makes the take-rate/payment-mix mismatch the central challenge, using the card, ACH, and check/cash unit economics from the pack.

Says what would prove it wrongRightRightRight
Gemini 3.5 Flash-Lite · Gemini

It describes a six-week test with 30 mid-market commercial customers, a discounted card option, and the explicit threshold 'at least 25% of their commercial payors opt-in' that would prove its skepticism wrong.

Sonnet 5.5 · API

It specifies a six-week commercial cohort and fast-pay test with concrete thresholds, including commercial card share above 25% or 0.3% net on routed commercial volume.

GPT-6 Luna · API

It specifies a six-week stratified cohort test with net revenue per adopter approaching $14,700 as the threshold to change its mind.

Tigers, not paper tigersRightRightRight
Gemini 3.5 Flash-Lite · Gemini

It distinguishes the real killer (commercial card adoption) from execution risk, dismisses the residential-only comparison to Tradesly as misleading because of mix, and calls out the unspoken cost of pausing the scheduling rewrite ($1.9 M churn).

Sonnet 5.5 · API

It ranks the take-rate/payment-mix flaw as primary, treats Tradesly and lock-ins as supporting evidence, and surfaces the churn/rewrite tradeoff as the uncomfortable issue.

GPT-6 Luna · API

It triages the take-rate mismatch as the sink, dismisses Tradesly as a weak comparator, and surfaces the scheduling-rewrite churn risk.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.095.82None
2GPT-6 AstrawithChatGPT93.295.82None
3GPT-6 LunawithAPI97.791.72None
4Sonnet 5.5withAPI90.975.02None
5Gemini 3.5 Flash-LitewithGemini88.675.021 capped
6Opus 5.5withClaude93.291.721 capped
7Gemini 3.8 FlashwithAPI93.237.521 capped

About the task

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Separates the risks that could kill it from the ones that only look scary
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill · Staff level: a company bet with a long evidence pack

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.