Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  2. Identifies material uncertainty100% pass
    It names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  3. Finds the load-bearing assumption100% pass
    It makes the 0.7% blended take applied to all invoice value the central challenge, backed by payment-mix and unit-economics figures.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet

Where it slips

  1. A cheap test that can actually read out67% pass
    The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.
    GPT-6.1 Sol · API · An AI SDR for small agencies
  2. Avoids unsupported claims77% pass
    The memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
    Opus 5.5 · Claude · The CEO's embedded-payments bet
  3. Uses the supplied evidence correctly79% pass
    The memo presents the pilot blended net take of ~0.38% as an achieved fact, but it is not supported by the pack and is an assumption-driven extrapolation.
    Opus 5.5 · Claude · The CEO's embedded-payments bet

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Fieldline. Our CEO wants to commit three squads for three quarters to Fieldline Pay, embedded payments for our customers' invoices, and has told the board it could add $24M of annual revenue within two years. Priya Raman, our CPO, has asked you to write the strongest challenge to the plan as a pre-read for next week's exec offsite. The CEO, CFO and Head of Sales will all read it. Write a memo of no more than 1,200 words that: 1. Names the single assumption the plan most depends on that the evidence does not support, and shows why, using the numbers in the pack. 2. Re-estimates the revenue from the supplied data, showing your working, as a range. 3. Says what we would need to see to be proved wrong, and the cheapest test that would show it within six weeks. 4. Says what, if anything, we should do instead or how the bet should change. The pack below is everything we have. Some of it matters more than the rest.

About FieldlineField-service software for trades businesses (plumbing, HVAC, electrical): scheduling, dispatch, quotes and invoicing. 2,300 customers, $41.0M ARR, average $17,800 per customer. Series C; the next raise is planned in about 14 months.
CEO's memo to the exec team (excerpt)“Every invoice our customers send is money we don't touch. Our customers invoiced $4.83B last year through Fieldline. If we process those payments, we become part of how they get paid, not just how they schedule. The model is simple: 70% of customers adopt within 18 months, the average customer invoices $2.1M a year, and we keep a 0.7% blended net take. That is $24M of new annual revenue by month 24, more than half our current ARR, and it makes the next raise a very different conversation. Tradesly has shown it works: payments are now 22% of their revenue. I've told the board I believe payments can be 35% of our revenue by 2028. I want three squads on this from next quarter, which means pausing the scheduling rewrite.”
CFO's revenue modelCustomers: 2,300. Adoption by month 18: 70% (1,610 customers). Annual invoiced value per customer: $2.1M (total invoiced $4.83B ÷ 2,300). Blended net take rate: 0.7%. Month-24 revenue run-rate: 1,610 × $2.1M × 0.7% = $23.7M. CFO's note: “Adoption and take rate are the CEO's assumptions. I haven't stress-tested them.”
Invoicing data (last 12 months, all customers)Total invoiced value: $4.83B. Mean per customer: $2.1M. Median per customer: $640k. The largest 5% of customers (115) account for 48% of invoiced value ($2.32B). By job type, residential jobs are 42% of invoiced value and commercial jobs (property managers, facilities contracts) are 58%.
How invoices are paid todayFrom the 690 customers (30%) who record the payment method in Fieldline, by share of invoice value: card 19%, ACH/bank transfer 44%, check 31%, cash 6%. Card share is 38% of residential invoice value and 5% of commercial. Average commercial invoice: $3,800; average residential invoice: $410. Commercial clients pay on net-45 or net-60 terms; the average commercial invoice is paid 52 days after it is sent.
Payments partner term sheet (unit economics)Card: the customer is charged 2.9% + $0.30 per payment; our all-in cost (interchange, network, partner fee) is about 2.2%, so we net about 0.7% of card value. ACH: the customer is charged a flat $2.00 per payment; our cost is $0.40. Checks and cash earn nothing unless the payer switches to card or ACH. The partner handles licensing, KYC and risk; they have approved our application.
Pilot (4 months)62 customers invited, 38 adopted (61%). Pilot customers' invoice value is 64% residential (the customer base is 42%). With pay-by-link on every invoice, card share of invoices paid through Fieldline Pay rose to 71% for residential and 6% for commercial. Net payments revenue, annualised: $214k across the 38 customers ($5,630 per customer per year).
Sales notes on the largest accountsOf the 115 largest customers, 71 have multi-year contracts with an existing payment processor, most running to 2028. Head of Sales, in Slack: “None of the big ones will move processors before their contracts end, and their property-manager clients will not pay 2.9% on a $3,800 invoice. They'll pay by ACH or check like they always have.”
Customer interviews (22 customers, last quarter)17 of 22 named getting paid on commercial jobs as their biggest cash problem (“I'm floating $200k of payroll while property managers sit on invoices for two months”). 9 said they would pay a fee to be paid faster. 6 said they won't offer card payment because clients fight the surcharge. Of the 14 customers with mostly commercial work, 11 said their clients require ACH or check.
CompetitorTradesly launched embedded payments in 2025 and says payments are now 22% of its revenue. Its published case studies are almost all home-services businesses; it describes its customers as about 80% residential.
EngineeringThe plan needs three squads for three quarters. The scheduling rewrite it would pause is the top-cited reason for churn among customers with more than 50 technicians (lost ARR from that segment last year: $1.9M).
What a strong answer does

Names the load-bearing assumption as the 0.7% blended net take: that is the card margin applied to all invoiced value, but only about 19% of value is paid by card (5% of commercial, which is 58% of value), ACH nets about 0.04% on a $3,800 invoice, and checks and cash earn nothing. Re-estimates with the supplied data: even at the pilot's improved mix (about a third of value on card), the blended take is roughly 0.2–0.25%, so the run-rate is about $5–9M, not $24M (the pilot's $5,630 per customer × 1,610 = $9.1M is an upper bound because the pilot was residential-heavy; the 71 locked-in large accounts, with 48% of value in the top 5%, cut it further). Notes that $24M needs essentially all invoice value on card. Treats Tradesly as a residential-mix result, not proof, and engineering capacity, licensing and competition as secondary. Keeps what the evidence supports: card payments work for residential invoices, and the strongest pull is commercial cash flow (17 of 22), so it reshapes the bet toward faster payout or financing on commercial invoices, which earns on the volume the processing model can't. Proposes a cheap test with a threshold that would prove the challenge wrong (for example, commercial card share well above 20% with pay-by-link in a six-week cohort, or uptake of paid instant payout), and is written for a room that includes the CEO: firm on the numbers, not personal.

Critical failures (cap the score)
  • Accepts the $24M projection or the 0.7% blended net take rate as sound
  • Invents market sizes, competitor figures or customer statistics that are not in the pack
Case

v1.5 · synthetic · B2B SaaS, fintech, strategy, staff-level

Usable with a quick edit11 of 11 checks passed

The memo is complete and well-supported; it could be used with only minor wording edits.

Output

To: Priya Raman; CEO; CFO; Head of Sales From: Staff PM Subject: Fieldline Pay: challenge the revenue thesis before committing three squads

Recommendation

Do not commit three squads for three quarters or pause the scheduling rewrite. Authorize a six-week commercial test using the existing pilot product and partner.

Embedded payments may be worth building. The evidence does not support presenting $24M as an expected month-24 revenue run-rate. My scenario range is approximately $2–$10M annualized, with $10M an optimistic case—not a base case.

1. The unsupported assumption: adopted invoice dollars yield 0.7%

The plan’s most consequential assumption is that the invoice value belonging to adopting customers becomes payment volume earning a blended 0.7% net take.

That is card economics applied to all invoices. It conflates three things: adopting Fieldline Pay, routing payments through it, and choosing a profitable payment method.

The available evidence contradicts it:

  • Only 19% of recorded payment value is card. ACH is 44%; checks and cash are 37%. These observations cover 30% of customers, so representativeness remains unproven.
  • ACH does not yield 0.7%. We earn $1.60 per payment. At the average commercial invoice of $3,800, that is 0.042%; at the residential average of $410, 0.390%.
  • Commercial work is 58% of invoice value. Even in the pilot, commercial card share reached only 6%, versus 71% for residential. A 2.9% card fee on a $3,800 invoice is about $110, before the fixed charge. Sales and interviews both indicate resistance to that fee.
  • The pilot overweights our best payments segment: 64% residential versus 42% across Fieldline. Tradesly’s roughly 80% residential customer mix makes its revenue share a poor benchmark for ours.

Volume concentration compounds the error. The largest 115 customers supply 48% of invoice value; 71 have processor contracts, mostly through 2028. We do not know those 71 customers’ exact volume, but cannot assume their invoices are available.

Consequently, 70% customer adoption does not establish 70% volume capture. The $2.1M average also obscures this concentration: the median is $640k, and customers outside the largest 5% average only $1.15M.

2. Re-estimate: approximately $2–$10M annualized

Use:

Revenue = eligible invoice value × adoption × routed share × payment-method yield.

The following scenarios assume adopters route all eligible invoices through Pay. That is generous. They use today’s customer base, not unsupported customer growth, and are scenarios rather than a statistical confidence interval.

Conservative scenario: approximately $2.3M

Assumptions:

  • Exclude the entire largest-5% cohort pending evidence of access. This is a conservative scenario, not a claim that all 115 are contractually blocked.
  • Remaining invoice value: $4.83B − $2.32B = $2.51B.
  • Adoption: pilot’s 61%.
  • Retain today’s 19% card and 44% ACH value shares; checks and cash earn nothing.
  • Value ACH at the commercial invoice size. This produces the lower ACH yield.

Blended yield:

19% × 0.7% + 44% × ($1.60 ÷ $3,800) = 0.152%.

Revenue:

$2.51B × 61% × 0.152% ≈ $2.3M.

This assumes the remaining cohort has the overall payment mix; the pack does not provide its actual mix.

Optimistic scenario: approximately $10.3M

Assumptions:

  • All $4.83B is accessible, despite processor contracts.
  • 70% adoption also captures 70% of invoice value.
  • Pilot card conversion transfers to Fieldline’s actual job mix.
  • Every non-card dollar becomes ACH: no checks, cash or external payments.

Card value share:

42% × 71% + 58% × 6% = 33.3%.

Blended yield, using job-specific invoice sizes:

33.3% × 0.7% + [42% × 29% × ($1.60 ÷ $410)] + [58% × 94% × ($1.60 ÷ $3,800)] = 0.304%.

Revenue:

$4.83B × 70% × 0.304% ≈ $10.3M.

Even this unusually favorable case is 57% below the CFO’s $23.7M.

A separate cross-check is the pilot’s actual monetization: 1,610 adopters × $5,630 = $9.1M. That is not a forecast: the pilot is residential-heavy, and we lack its invoice-volume denominator. But it does not validate $24M.

The range could fall below $2M if adoption or routing disappoints. We lack the data to assign probabilities.

3. What would prove this challenge wrong?

To generate $24M from $4.83B, Fieldline needs portfolio-wide annual net yield of approximately 0.50%. At 70% volume capture, captured payments must yield approximately 0.71%.

We would need evidence of either:

  1. Much more accessible, routed, high-margin card volume than today’s mix and pilot suggest; or
  2. A demonstrably profitable additional fee or product customers will actually pay for.

Adoption percentages, enthusiasm and competitor revenue shares would not suffice.

Cheapest six-week test

Use the existing pay-by-link product and partner; no new platform build.

Recruit approximately 50 non-pilot customers, stratified by invoice volume and residential/commercial mix, including a deliberate sample of large accounts with processor contracts. Weight results by invoice dollars, not customer count.

For each, establish contractual access and offer Pay at intended production pricing. Measure:

  • Activation and actual routed invoice value;
  • Card/ACH/check mix, invoice sizes and realized net revenue;
  • Payer refusals, especially commercial surcharge refusals;
  • Large-account willingness and ability to switch.

Commercial invoices take 52 days to pay on average. Include existing invoices due during the test, not just newly issued invoices, and report unpaid invoices separately rather than assuming conversion.

The test passes only if observed economics and credible volume-access commitments support a transparent path to $24M. A small test cannot prove 18-month adoption, but it can cheaply falsify the unit economics or access thesis.

4. Change the bet

Keep the scheduling rewrite staffed. It addresses the top-cited churn reason in a segment that lost $1.9M ARR last year; postponing it has a concrete opportunity cost ahead of our raise.

If the test supports attractive economics, fund a narrower, stage-gated residential payments rollout—not the current commitment.

Separately, test commercial customers’ willingness to pay for faster collection. Seventeen of 22 interviewees identified commercial payment delays as their biggest cash problem; nine would pay to get paid faster. Payments acceptance does not itself solve net-60 terms. Explore collections workflows or partner-funded acceleration before underwriting a financing business.

Finally, reset the board’s expectation now: $24M is an unvalidated upside hypothesis, not an evidence-backed plan.

Check by check

Got right · 11
  • Uses the supplied evidence correctlyEvery factual figure used comes from the supplied pack or follows directly from its arithmetic; no current-state fact is invented.
  • Addresses the actual decisionIt commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
  • Respects explicit constraintsIt is a memo addressed to the named exec readers, covers all requested sections, and is within the 1,200-word limit.
  • Identifies material uncertaintyIt names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
  • Avoids unsupported claimsInterpretations and forecasts are clearly framed as assumptions or scenarios, not established fact.
  • Produces the required deliverableThe deliverable is a complete pre-read memo usable by the exec team with light edits.
  • Finds the load-bearing assumptionIt makes the 0.7% blended take applied to all invoice value the central challenge, backed by payment-mix and unit-economics figures.
  • Re-estimates the revenue correctlyIt shows arithmetic landing in $2.3M to $10.3M with a $9.1M pilot cross-check and explains the pilot's residential-heavy bias.
  • Says what would prove it wrongIt names a six-week test with cohort, metrics and economics threshold around 0.71% captured yield needed to support $24M.
  • Reshapes the bet from the evidenceIt proposes keeping the scheduling rewrite, a stage-gated residential rollout, and testing commercial faster-payment demand, grounded in the evidence.
  • Tigers, not paper tigersIt triages take rate and volume access above secondary risks and dismisses the Tradesly analogy with the residential-mix reason.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly39%
  • passAddresses the actual decision100%
  • passRespects explicit constraints58%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims53%
  • passProduces the required deliverable81%
  • passFinds the load-bearing assumption100%
  • passRe-estimates the revenue correctly62%
  • passSays what would prove it wrong47%
  • passReshapes the bet from the evidence100%
  • passTigers, not paper tigers98%
Run
Run
#1
API response time
60 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.095.82None
2GPT-6 AstrawithChatGPT95.595.82None
3GPT-6 LunawithAPI93.291.72None
4Sonnet 5.5withAPI93.275.02None
5Gemini 3.5 Flash-LitewithGemini81.875.021 capped
6Opus 5.5withClaude90.991.721 capped

About the task

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Separates the risks that could kill it from the ones that only look scary
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill · Staff level: a company bet with a long evidence pack

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.