Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  2. Identifies material uncertainty100% pass
    It names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  3. Finds the load-bearing assumption100% pass
    It makes the 0.7% blended take applied to all invoice value the central challenge, backed by payment-mix and unit-economics figures.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet

Where it slips

  1. A cheap test that can actually read out67% pass
    The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.
    GPT-6.1 Sol · API · An AI SDR for small agencies
  2. Avoids unsupported claims77% pass
    The memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
    Opus 5.5 · Claude · The CEO's embedded-payments bet
  3. Uses the supplied evidence correctly79% pass
    The memo presents the pilot blended net take of ~0.38% as an achieved fact, but it is not supported by the pack and is an assumption-driven extrapolation.
    Opus 5.5 · Claude · The CEO's embedded-payments bet

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Fieldline. Our CEO wants to commit three squads for three quarters to Fieldline Pay, embedded payments for our customers' invoices, and has told the board it could add $24M of annual revenue within two years. Priya Raman, our CPO, has asked you to write the strongest challenge to the plan as a pre-read for next week's exec offsite. The CEO, CFO and Head of Sales will all read it. Write a memo of no more than 1,200 words that: 1. Names the single assumption the plan most depends on that the evidence does not support, and shows why, using the numbers in the pack. 2. Re-estimates the revenue from the supplied data, showing your working, as a range. 3. Says what we would need to see to be proved wrong, and the cheapest test that would show it within six weeks. 4. Says what, if anything, we should do instead or how the bet should change. The pack below is everything we have. Some of it matters more than the rest.

About FieldlineField-service software for trades businesses (plumbing, HVAC, electrical): scheduling, dispatch, quotes and invoicing. 2,300 customers, $41.0M ARR, average $17,800 per customer. Series C; the next raise is planned in about 14 months.
CEO's memo to the exec team (excerpt)“Every invoice our customers send is money we don't touch. Our customers invoiced $4.83B last year through Fieldline. If we process those payments, we become part of how they get paid, not just how they schedule. The model is simple: 70% of customers adopt within 18 months, the average customer invoices $2.1M a year, and we keep a 0.7% blended net take. That is $24M of new annual revenue by month 24, more than half our current ARR, and it makes the next raise a very different conversation. Tradesly has shown it works: payments are now 22% of their revenue. I've told the board I believe payments can be 35% of our revenue by 2028. I want three squads on this from next quarter, which means pausing the scheduling rewrite.”
CFO's revenue modelCustomers: 2,300. Adoption by month 18: 70% (1,610 customers). Annual invoiced value per customer: $2.1M (total invoiced $4.83B ÷ 2,300). Blended net take rate: 0.7%. Month-24 revenue run-rate: 1,610 × $2.1M × 0.7% = $23.7M. CFO's note: “Adoption and take rate are the CEO's assumptions. I haven't stress-tested them.”
Invoicing data (last 12 months, all customers)Total invoiced value: $4.83B. Mean per customer: $2.1M. Median per customer: $640k. The largest 5% of customers (115) account for 48% of invoiced value ($2.32B). By job type, residential jobs are 42% of invoiced value and commercial jobs (property managers, facilities contracts) are 58%.
How invoices are paid todayFrom the 690 customers (30%) who record the payment method in Fieldline, by share of invoice value: card 19%, ACH/bank transfer 44%, check 31%, cash 6%. Card share is 38% of residential invoice value and 5% of commercial. Average commercial invoice: $3,800; average residential invoice: $410. Commercial clients pay on net-45 or net-60 terms; the average commercial invoice is paid 52 days after it is sent.
Payments partner term sheet (unit economics)Card: the customer is charged 2.9% + $0.30 per payment; our all-in cost (interchange, network, partner fee) is about 2.2%, so we net about 0.7% of card value. ACH: the customer is charged a flat $2.00 per payment; our cost is $0.40. Checks and cash earn nothing unless the payer switches to card or ACH. The partner handles licensing, KYC and risk; they have approved our application.
Pilot (4 months)62 customers invited, 38 adopted (61%). Pilot customers' invoice value is 64% residential (the customer base is 42%). With pay-by-link on every invoice, card share of invoices paid through Fieldline Pay rose to 71% for residential and 6% for commercial. Net payments revenue, annualised: $214k across the 38 customers ($5,630 per customer per year).
Sales notes on the largest accountsOf the 115 largest customers, 71 have multi-year contracts with an existing payment processor, most running to 2028. Head of Sales, in Slack: “None of the big ones will move processors before their contracts end, and their property-manager clients will not pay 2.9% on a $3,800 invoice. They'll pay by ACH or check like they always have.”
Customer interviews (22 customers, last quarter)17 of 22 named getting paid on commercial jobs as their biggest cash problem (“I'm floating $200k of payroll while property managers sit on invoices for two months”). 9 said they would pay a fee to be paid faster. 6 said they won't offer card payment because clients fight the surcharge. Of the 14 customers with mostly commercial work, 11 said their clients require ACH or check.
CompetitorTradesly launched embedded payments in 2025 and says payments are now 22% of its revenue. Its published case studies are almost all home-services businesses; it describes its customers as about 80% residential.
EngineeringThe plan needs three squads for three quarters. The scheduling rewrite it would pause is the top-cited reason for churn among customers with more than 50 technicians (lost ARR from that segment last year: $1.9M).
What a strong answer does

Names the load-bearing assumption as the 0.7% blended net take: that is the card margin applied to all invoiced value, but only about 19% of value is paid by card (5% of commercial, which is 58% of value), ACH nets about 0.04% on a $3,800 invoice, and checks and cash earn nothing. Re-estimates with the supplied data: even at the pilot's improved mix (about a third of value on card), the blended take is roughly 0.2–0.25%, so the run-rate is about $5–9M, not $24M (the pilot's $5,630 per customer × 1,610 = $9.1M is an upper bound because the pilot was residential-heavy; the 71 locked-in large accounts, with 48% of value in the top 5%, cut it further). Notes that $24M needs essentially all invoice value on card. Treats Tradesly as a residential-mix result, not proof, and engineering capacity, licensing and competition as secondary. Keeps what the evidence supports: card payments work for residential invoices, and the strongest pull is commercial cash flow (17 of 22), so it reshapes the bet toward faster payout or financing on commercial invoices, which earns on the volume the processing model can't. Proposes a cheap test with a threshold that would prove the challenge wrong (for example, commercial card share well above 20% with pay-by-link in a six-week cohort, or uptake of paid instant payout), and is written for a room that includes the CEO: firm on the numbers, not personal.

Critical failures (cap the score)
  • Accepts the $24M projection or the 0.7% blended net take rate as sound
  • Invents market sizes, competitor figures or customer statistics that are not in the pack
Case

v1.5 · synthetic · B2B SaaS, fintech, strategy, staff-level

Usable with a quick edit9 of 11 checks passed

The one unsupported figure (0.38% pilot take) is used only in a secondary cross-check and can be removed with a light edit; the core analysis and recommendations are solid and data-grounded.

Output

Fieldline Pay: the case against the plan as written

Sep 25, 2026 · @Martin Slaney · Pre-read for exec offsite

Bottom line

The $24M figure applies a card-only margin to invoice value that is mostly not paid by card. Re-estimated from our own data, Fieldline Pay is a $2–7M business at month 24, central case about $4M. That is worth building, but not with three squads for three quarters, and not by pausing the rewrite that cost us $1.9M ARR last year.

1. The assumption that breaks: a 0.7% blended take

The 0.7% is not a blended rate. It is our net margin on card payments (2.9% charged, \~2.2% cost). The model applies it to all $4.83B.

  • Card is 19% of invoice value today. ACH is 44%, check 31%, cash 6%.
  • ACH nets a flat $1.60 per payment. On a $3,800 commercial invoice that is 0.04%; on a $410 residential invoice, 0.39%.
  • Checks and cash earn nothing.
  • Commercial is 58% of value, and its card share is 5%. With pay-by-link on every invoice, the pilot moved it to 6%. Sales and 11 of 14 commercial-heavy interviewees say clients require ACH or check.

Worked through at pilot-level card adoption (assuming half of non-card value moves to ACH):

SegmentShare of valueCard share (pilot)Net take
Residential42%71%\~0.55%
Commercial58%6%\~0.065%
Blended, our mix100%\~0.27%

At today's payment mix, with no card uplift, blended take is about 0.15%.3

The $24M needs 0.71% across 70% of customers: every processed dollar on card. Even 100% adoption of all $4.83B would need 0.50%, above what the residential-skewed pilot achieved (\~0.38%).

Two things compound it:

  • Adoption by customers is not adoption by value. The top 115 customers hold 48% of value; 71 are locked into processor contracts, mostly to 2028. Pro rata, that is \~$1.43B (30% of value) out of reach in the window. The pilot's 61% came from an invited, 64%-residential cohort.
  • Tradesly is not a comparable. It is \~80% residential; we are 42%. Its 22% shows residential payments work, which our pilot confirms. It says nothing about commercial.

2. Re-estimate: $2–7M at month 24

Addressable value = $4.83B less about $1.43B locked = about $3.4B. Revenue = addressable value × value-weighted adoption × blended take.

CaseAdoption (by value)Blended takeMonth-24 run-rate
Low30%0.15% (card uplift stalls)$1.5M
Central45%0.27% (pilot behaviour, our mix)$4.1M
High60%0.35% (unlocked base skews residential)$7.1M

Cross-check from the pilot: $5,630 per customer a year on a 64%-residential cohort. Re-weighted to our mix (0.27% ÷ 0.38%) that is \~$4,000. Across 1,000–1,560 adopters (45–70% of the 2,229 unlocked customers), it gives $4.0–6.2M, consistent with the central-to-high cases.

Central case is \~10% of current ARR, not the 35% of revenue told to the board. Against it: three squads for three quarters, and the rewrite delayed in a segment that lost $1.9M ARR last year. At the central case, year-two payments revenue roughly covers two years of that churn if it continues.

3. What would prove this wrong, and the cheapest test

I am wrong if commercial invoice value can be moved onto rails we earn on. The threshold: a commercially representative cohort showing blended net take of 0.5% or more through Pay. Below 0.3%, the $24M is off the table.2

Six-week test, configuration only, no new build:

  1. Enable Pay for 30 customers with at least 60% commercial invoice value, none from the pilot.
  2. Attach pay-links to open invoices sent in the last 60 days, not just new ones. Commercial invoices take 52 days to pay; new invoices alone won't settle inside six weeks.
  3. Split the cohort on ACH price: half at the current flat $2, half at 0.5% capped at $25. Flat pricing is why commercial ACH earns $1.60 on a $3,800 invoice instead of $19.
  4. Make payment-method capture mandatory for the cohort, so we stop relying on the 30% who record it.
  5. Measure value-weighted method mix and net take by segment, plus opt-outs and payer complaints.

Alongside it, at no cost: account managers ask the 44 unlocked top-115 accounts whether they would move processors in 2027. More than ten yeses changes the adoption maths.

4. What to do instead

Shrink the bet, point it where the evidence is, and gate the rest.

  • Keep the scheduling rewrite. It protects known ARR. Payments is unproven in the segment that holds most of the value.
  • One squad on Pay for residential work now. The pilot shows it works there (71% card share). Ship pay-by-link, default on.
  • Reprice ACH from flat $2 to a capped percentage, subject to the test. It is the largest lever on commercial take and needs no new build.
  • Go after the real problem. 17 of 22 customers named getting paid on commercial jobs as their biggest cash problem; invoices take 52 days; 9 would pay to be paid faster. An early-payout product (for example, 1.5% to get paid in two days) earns $57 on a $3,800 invoice against $1.60 on ACH.1 It earns on the 58% of value that card never will. It needs a capital partner; scope it in the same six weeks with a priced offer to 10 commercial customers.
  • Gate at week six. If commercial take clears 0.5% or the early-payout offer converts, add squads then.
  • Reset the board number now. We raise in about 14 months, when Pay will show perhaps a $2–3M run-rate. Walking the board from $24M to a range today is cheaper than explaining the gap in a data room.

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticQuick edit

    What we’d changeThe $57 is gross. Net it of the cost of capital and defaults before comparing it with ACH's net $1.60.

  2. 2Test or gate too weakTighten the testQuick edit

    What we’d changeSay which claim the test settles, and whether 0.3% means commercial or blended take.

  3. 3Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeLabel modelled figures as modelled, so the CFO doesn't read them as observed.

Check by check

Got wrong · 1
  • Avoids unsupported claimsThe memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
Mixed · 1
  • Uses the supplied evidence correctlyThe memo presents the pilot blended net take of ~0.38% as an achieved fact, but it is not supported by the pack and is an assumption-driven extrapolation.The two graders disagreed on this one.
Got right · 9
  • Addresses the actual decisionIt commits to a clear call (shrink the bet, one squad on residential, early-payout) and states what would change the call (commercial blended take >0.5%).
  • Respects explicit constraintsIt is a memo for the CEO, CFO and Head of Sales, within the word limit, and addresses all four requested points.
  • Identifies material uncertaintyIt names the key unknowns (commercial card take, locked account adoption) and proposes a six-week test with a specific threshold to resolve them.
  • Produces the required deliverableThe memo is a complete, action-ready pre-read that covers the challenge, re-estimate, test and alternative, and could be used by the execs with light edits.
  • Finds the load-bearing assumptionIt makes the mismatch between the 0.7% card-only net take and the actual payment mix the central challenge, and backs it with the payment-method and unit-economics data.
  • Re-estimates the revenue correctlyIt recalculates from the payment mix and pilot behavior, showing work, yielding a central ~$4M and an upper bound of $7.1M, well below $24M.
  • Says what would prove it wrongIt specifies a six-week cohort test with a commercial blended take threshold of 0.5% and a secondary light-touch inquiry for locked accounts.
  • Reshapes the bet from the evidenceIt grounds the recommendation in what works (residential card payments) and the strongest evidence (commercial cash-flow pain), proposing a narrower bet with an early-payout product.
  • Tigers, not paper tigersIt ranks the take-rate mismatch as the killer risk, dismisses the Tradesly comp as residential-only, and names the uncomfortable board reset needed now.

Claims the judge couldn’t find in the brief

  • The residential-skewed pilot achieved a blended net take of ~0.38%.

Grades and run details

Decision model 91 · LLM judge 10 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly8%
  • passAddresses the actual decision100%
  • passRespects explicit constraints50%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims22%
  • passProduces the required deliverable92%
  • passFinds the load-bearing assumption100%
  • partialRe-estimates the revenue correctly32%
  • passSays what would prove it wrong100%
  • passReshapes the bet from the evidence100%
  • passTigers, not paper tigers99%
Run
Run
#1
Time to output
2.0 min
Submitted
25 Sept 2026

Usable with a quick edit10 of 11 checks passed

The pre-read is tightly argued, numerically grounded, and directly usable by a senior PM; only minor phrasing tweaks would be needed to align the bull case with pure payments revenue if desired.

Output

To: Priya Raman (CPO), CEO, CFO, Head of Sales From: Staff Product Manager, Fieldline Date: September 25, 2026 Subject: Pre-read: Strategic Challenge to the Fieldline Pay Plan ($24M Revenue Target)

---

Executive Summary

Committing three squads for three quarters to Fieldline Pay to capture $24M in new revenue within two years is a high-risk bet that relies on a structural distortion in our financial model.

The plan’s core vulnerability is not execution; it is a misallocation of our customer base's economic reality. By using the mean invoice value ($2.1M) rather than the median ($640k), the model assumes our average customer mirrors our largest commercial accounts1, while simultaneously ignoring entrenched enterprise contracts and commercial payment norms.

Executing this plan requires pausing the scheduling rewrite, which risks $1.9M in annual churn from our largest, most valuable accounts. Below is the evidence-based challenge, a revised revenue range, an inexpensive six-week test, and a recommended alternative path.

---

1. The Dependent Assumption and Why the Evidence Fails It

> The Plan's Single Dependent Assumption: That we can achieve a $23.7M–$24M revenue run-rate by month 24 by applying a blended 0.7% net take rate across a homogenous $2.1M annual invoice volume per adopted customer.

The evidence flatly refutes this assumption in three ways:

  • The Mean vs. Median Distortion: The CEO's model relies on a mean invoice value of $2.1M. However, the median customer invoices just $640k. Our distribution is heavily skewed: the top 5% of customers (115 accounts) drive 48% ($2.32B) of our total invoiced value.
  • Enterprise Lock-in: Of those 115 largest customers, 71 are locked into multi-year contracts with existing payment processors running through 2028. As the Head of Sales noted, none will migrate early.
  • The Commercial Wall: Commercial jobs comprise 58% of our invoiced value ($2.80B). The average commercial invoice is $3,800 and is paid on net-45/net-60 terms (averaging 52 days). Property managers and facility contractors will not pay a 2.9% card surcharge on a $3,800 invoice. They pay via ACH ($2.00 flat fee, netting us $1.60) or check ($0 revenue), exactly as they always have.

Applying a 0.7% card take-rate model to commercial volume is fundamentally flawed because commercial clients do not use cards (card share is only 5% for commercial, vs. 38% for residential).

---

2. Re-Estimated Revenue Range (Working Included)

To model realistic revenue, we must segment our base by customer type, recognizing that residential and commercial segments have radically different payment behaviors and take rates.

#### Step-by-Step Working: 1. Customer Base: 2,300 total customers. 2. Adoption Rate: Pilot adoption was 61% (close to the 70% target). Let's model a realistic 60% adoption rate by month 24 = $1,380$ adopting customers. 3. Segmentation Split: Based on overall volume, 42% of value is residential ($2.03B) and 58% is commercial ($2.80B). Total invoiced value = $4.83B. * Total Residential Invoiced: $2.03B ($\approx$ $882k$ per customer across 2,300) * Total Commercial Invoiced: $2.80B ($\approx$ $1.22M$ per customer across 2,300) 4. Take Rates & Behavior: * Residential: 71% card adoption via pay-by-link. Net take rate on card is 0.7%. ACH/Check take rate is near zero (flat $1.60 net on ACH, negligible volume). Effective blended take rate on residential volume $\approx$ $0.7\% \times 71\% \approx \mathbf{0.50\%}$. * Commercial: 6% card adoption, 94% ACH/Check/Terms. Card take rate is 0.7%; ACH nets a flat $1.60 per transaction (on a $3,800 invoice, $1.60 is a 0.04% effective take rate). Effective blended take rate on commercial volume $\approx$ $\mathbf{0.08\%}$.

#### The Realistic Range (Month 24 Run-Rate): * Bear Case ($3.2M ARR): Commercial customers reject card fees entirely, sticking strictly to ACH/checks; residential adoption stalls at 45% due to surcharge pushback. * Base Case ($6.1M ARR): 60% overall adoption. Residential volume yields a 0.50% blended take ($2.03B $\times$ 60% adoption $\times$ 0.50% = $6.09M). Commercial yields minimal flat-fee ACH revenue. * Bull Case ($9.8M ARR):3 70% adoption matches the CEO's target, and we successfully introduce a B2B "accelerated payout" fee (drawing on the 9 customers out of 22 who said they would pay a fee to solve their $200k payroll float).

> Result: The realistic revenue run-rate at Month 24 is $3.2M to $9.8M, falling drastically short of the $24M target.

---

3. Falsification Criteria and a Six-Week Test

#### What would prove us wrong? If a randomized cohort of commercial-heavy customers willingly adopts card payments at >20% volume despite a 2.9% surcharge, or if property managers accept automated card-on-file billing for invoices over $3,000, our commercial pessimism is unfounded.

#### The Six-Week Test ($15k budget, 1 squad for 6 weeks): * The Experiment: Launch a targeted "Fast-Pay Commercial Portal" pilot with 30 mid-market commercial customers currently handling invoices between $2,000 and $5,000. * The Mechanics: Offer them an explicit choice: continue standard net-45 terms via free ACH, or use a discounted commercial card rate (e.g., split-surcharge or 1.9% + $0.30 via a specialized B2B interchange tier)2 to get paid in 48 hours. * Success Metric: Do at least 25% of their commercial payors opt-in to the digital card rail over 6 weeks? If not, the commercial card thesis is dead.

---

4. What We Should Do Instead

Pausing the scheduling rewrite to chase an illusory $24M payments figure is a strategic trap. The scheduling rewrite is our defensive moat: losing it costs us $1.9M in high-value churn annually from customers with >50 technicians.

#### Recommended Pivot: 1. Protect Core Engineering: Keep the scheduling rewrite on track. Do not cannibalize the core product teams for a pivot that math shows yields ~$6M, not $24M. Allocate one dedicated squad (not three) to productize payments iteratively. 2. Pivot the Value Proposition: Stop trying to monetize residential card transactions where competition (Tradesly, Stripe) is fierce. Instead, solve the actual pain point identified in customer interviews: commercial cash flow float. 3. Build B2B Working Capital Solutions: Build features that help trades businesses bridge the net-45 gap (e.g., instant payout/invoice factoring models embedded in Fieldline Pay) rather than relying purely on interchange take-rates.

---

Given that our next Series C raise is in 14 months, how comfortable is the board with a realistic payments ARR of $6M built on workflow efficiency, versus an aggressive $24M gamble that risks our core enterprise retention?

What a PM had to fix

  1. 1Numbers wrongRedo the arithmeticStart again

    Source: CFO's revenue modelAnnual invoiced value per customer: $2.1M (total invoiced $4.83B ÷ 2,300).

    What we’d changeDrop the mean-versus-median argument: customers times mean invoicing correctly recovers total volume. The real flaw is applying a card-only 0.7% margin to all of it.

  2. 2Invented evidenceVerify or remove the claimSubstantial rework

    What we’d changeRemove the discounted tier: nothing in the term sheet offers it. Design the test around economics we have, and make it test the $24M claim.

  3. 3Numbers wrongRedo the arithmeticTargeted repair

    What we’d changeShow the working for the bear and bull cases. Only the base case is calculated.

Check by check

Mixed · 1
  • Uses the supplied evidence correctlyEvery factual statement about the current situation is directly taken from the supplied context or derived by straightforward arithmetic, with no inventions.The two graders disagreed on this one.
Got right · 10
  • Addresses the actual decisionThe memo commits early to challenging the plan, recommends scaling back to one squad and pivoting to commercial cash-flow solutions, and specifies a condition (≥25% commercial card opt-in) that would change its assessment, all framed for the CEO, CFO and Head of Sales.
  • Respects explicit constraintsThe deliverable is a memo under 1200 words, addressed to the specified readers, and respects the four numbered requirements.
  • Identifies material uncertaintyThe memo pinpoints commercial card adoption as the critical unknown, bounds the revenue range, and proposes a concrete six-week test with a clear threshold that would resolve whether its commercial pessimism is wrong.
  • Avoids unsupported claimsInterpretations such as "structural distortion" are clearly argued from the data, and forecasts are presented as cases with assumptions, not as established fact.
  • Produces the required deliverableThe memo follows the requested format, is well within the word limit, and contains all four required elements in a way the executive audience could act on with light edits.
  • Finds the load-bearing assumptionThe memo makes the mismatch between the assumed 0.7% blended net take rate and the actual payment-method mix (mostly ACH/check for commercial) its central challenge, exactly as the strong answer expects.
  • Re-estimates the revenue correctlyIt segments customers by residential/commercial, applies realistic take rates from the pilot and unit economics, shows its arithmetic, and lands in a $3.2 M–$9.8 M range, far below $24 M, noting that the pilot's residential-heavy mix would give a higher bound.
  • Says what would prove it wrongIt describes a six-week test with 30 mid-market commercial customers, a discounted card option, and the explicit threshold 'at least 25% of their commercial payors opt-in' that would prove its skepticism wrong.
  • Reshapes the bet from the evidenceIt leverages the evidence that residential card payments work and that commercial cash-flow pain is acute, recommending a smaller investment, a pivot to faster payout/factoring solutions, and protecting the scheduling rewrite.
  • Tigers, not paper tigersIt distinguishes the real killer (commercial card adoption) from execution risk, dismisses the residential-only comparison to Tradesly as misleading because of mix, and calls out the unspoken cost of pausing the scheduling rewrite ($1.9 M churn).

Grades and run details

Decision model 86 · LLM judge 12 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly21%
  • passAddresses the actual decision74%
  • passRespects explicit constraints26%
  • passIdentifies material uncertainty68%
  • partialAvoids unsupported claims28%
  • passProduces the required deliverable48%
  • passFinds the load-bearing assumption50%
  • passRe-estimates the revenue correctly32%
  • passSays what would prove it wrong100%
  • passReshapes the bet from the evidence100%
  • passTigers, not paper tigers56%
Run
Run
#1
Time to output
14 s
Submitted
25 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.095.82None
2GPT-6 AstrawithChatGPT95.595.82None
3GPT-6 LunawithAPI93.291.72None
4Sonnet 5.5withAPI93.275.02None
5Gemini 3.5 Flash-LitewithGemini81.875.021 capped
6Opus 5.5withClaude90.991.721 capped

About the task

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Separates the risks that could kill it from the ones that only look scary
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill · Staff level: a company bet with a long evidence pack

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.