Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  2. Identifies material uncertainty100% pass
    It names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  3. Finds the load-bearing assumption100% pass
    It makes the 0.7% blended take applied to all invoice value the central challenge, backed by payment-mix and unit-economics figures.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet

Where it slips

  1. A cheap test that can actually read out67% pass
    The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.
    GPT-6.1 Sol · API · An AI SDR for small agencies
  2. Avoids unsupported claims77% pass
    The memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
    Opus 5.5 · Claude · The CEO's embedded-payments bet
  3. Uses the supplied evidence correctly79% pass
    The memo presents the pilot blended net take of ~0.38% as an achieved fact, but it is not supported by the pack and is an assumption-driven extrapolation.
    Opus 5.5 · Claude · The CEO's embedded-payments bet

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Fieldline. Our CEO wants to commit three squads for three quarters to Fieldline Pay, embedded payments for our customers' invoices, and has told the board it could add $24M of annual revenue within two years. Priya Raman, our CPO, has asked you to write the strongest challenge to the plan as a pre-read for next week's exec offsite. The CEO, CFO and Head of Sales will all read it. Write a memo of no more than 1,200 words that: 1. Names the single assumption the plan most depends on that the evidence does not support, and shows why, using the numbers in the pack. 2. Re-estimates the revenue from the supplied data, showing your working, as a range. 3. Says what we would need to see to be proved wrong, and the cheapest test that would show it within six weeks. 4. Says what, if anything, we should do instead or how the bet should change. The pack below is everything we have. Some of it matters more than the rest.

About FieldlineField-service software for trades businesses (plumbing, HVAC, electrical): scheduling, dispatch, quotes and invoicing. 2,300 customers, $41.0M ARR, average $17,800 per customer. Series C; the next raise is planned in about 14 months.
CEO's memo to the exec team (excerpt)“Every invoice our customers send is money we don't touch. Our customers invoiced $4.83B last year through Fieldline. If we process those payments, we become part of how they get paid, not just how they schedule. The model is simple: 70% of customers adopt within 18 months, the average customer invoices $2.1M a year, and we keep a 0.7% blended net take. That is $24M of new annual revenue by month 24, more than half our current ARR, and it makes the next raise a very different conversation. Tradesly has shown it works: payments are now 22% of their revenue. I've told the board I believe payments can be 35% of our revenue by 2028. I want three squads on this from next quarter, which means pausing the scheduling rewrite.”
CFO's revenue modelCustomers: 2,300. Adoption by month 18: 70% (1,610 customers). Annual invoiced value per customer: $2.1M (total invoiced $4.83B ÷ 2,300). Blended net take rate: 0.7%. Month-24 revenue run-rate: 1,610 × $2.1M × 0.7% = $23.7M. CFO's note: “Adoption and take rate are the CEO's assumptions. I haven't stress-tested them.”
Invoicing data (last 12 months, all customers)Total invoiced value: $4.83B. Mean per customer: $2.1M. Median per customer: $640k. The largest 5% of customers (115) account for 48% of invoiced value ($2.32B). By job type, residential jobs are 42% of invoiced value and commercial jobs (property managers, facilities contracts) are 58%.
How invoices are paid todayFrom the 690 customers (30%) who record the payment method in Fieldline, by share of invoice value: card 19%, ACH/bank transfer 44%, check 31%, cash 6%. Card share is 38% of residential invoice value and 5% of commercial. Average commercial invoice: $3,800; average residential invoice: $410. Commercial clients pay on net-45 or net-60 terms; the average commercial invoice is paid 52 days after it is sent.
Payments partner term sheet (unit economics)Card: the customer is charged 2.9% + $0.30 per payment; our all-in cost (interchange, network, partner fee) is about 2.2%, so we net about 0.7% of card value. ACH: the customer is charged a flat $2.00 per payment; our cost is $0.40. Checks and cash earn nothing unless the payer switches to card or ACH. The partner handles licensing, KYC and risk; they have approved our application.
Pilot (4 months)62 customers invited, 38 adopted (61%). Pilot customers' invoice value is 64% residential (the customer base is 42%). With pay-by-link on every invoice, card share of invoices paid through Fieldline Pay rose to 71% for residential and 6% for commercial. Net payments revenue, annualised: $214k across the 38 customers ($5,630 per customer per year).
Sales notes on the largest accountsOf the 115 largest customers, 71 have multi-year contracts with an existing payment processor, most running to 2028. Head of Sales, in Slack: “None of the big ones will move processors before their contracts end, and their property-manager clients will not pay 2.9% on a $3,800 invoice. They'll pay by ACH or check like they always have.”
Customer interviews (22 customers, last quarter)17 of 22 named getting paid on commercial jobs as their biggest cash problem (“I'm floating $200k of payroll while property managers sit on invoices for two months”). 9 said they would pay a fee to be paid faster. 6 said they won't offer card payment because clients fight the surcharge. Of the 14 customers with mostly commercial work, 11 said their clients require ACH or check.
CompetitorTradesly launched embedded payments in 2025 and says payments are now 22% of its revenue. Its published case studies are almost all home-services businesses; it describes its customers as about 80% residential.
EngineeringThe plan needs three squads for three quarters. The scheduling rewrite it would pause is the top-cited reason for churn among customers with more than 50 technicians (lost ARR from that segment last year: $1.9M).
What a strong answer does

Names the load-bearing assumption as the 0.7% blended net take: that is the card margin applied to all invoiced value, but only about 19% of value is paid by card (5% of commercial, which is 58% of value), ACH nets about 0.04% on a $3,800 invoice, and checks and cash earn nothing. Re-estimates with the supplied data: even at the pilot's improved mix (about a third of value on card), the blended take is roughly 0.2–0.25%, so the run-rate is about $5–9M, not $24M (the pilot's $5,630 per customer × 1,610 = $9.1M is an upper bound because the pilot was residential-heavy; the 71 locked-in large accounts, with 48% of value in the top 5%, cut it further). Notes that $24M needs essentially all invoice value on card. Treats Tradesly as a residential-mix result, not proof, and engineering capacity, licensing and competition as secondary. Keeps what the evidence supports: card payments work for residential invoices, and the strongest pull is commercial cash flow (17 of 22), so it reshapes the bet toward faster payout or financing on commercial invoices, which earns on the volume the processing model can't. Proposes a cheap test with a threshold that would prove the challenge wrong (for example, commercial card share well above 20% with pay-by-link in a six-week cohort, or uptake of paid instant payout), and is written for a room that includes the CEO: firm on the numbers, not personal.

Critical failures (cap the score)
  • Accepts the $24M projection or the 0.7% blended net take rate as sound
  • Invents market sizes, competitor figures or customer statistics that are not in the pack
Case

v1.5 · synthetic · B2B SaaS, fintech, strategy, staff-level

Usable with a quick edit11 of 11 checks passed

The memo is well-reasoned, evidence-backed, and directly usable for the offsite; only light editorial polish would be needed, with no factual or logical gaps to fill.

Output

To: Priya Raman and the executive team Subject: Fieldline Pay: validate the revenue engine before committing three squads

Recommendation: do not approve three squads for three quarters. Authorise a six-week commercial validation using the existing pilot and protect the scheduling rewrite. Payments may be a useful extension, but the supplied evidence supports an illustrative annual revenue range of $3.3M–$7.2M1, with substantial uncertainty. It does not support underwriting $24M or making that figure central to our next raise.

The assumption the plan depends on

The unsupported assumption is that customer adoption translates into monetising those customers’ entire invoice value at a 0.7% net yield.

The model multiplies customer count by average invoicing and card economics. Each component hides a material constraint:

  • Adoption is not payment capture. Enabling Fieldline Pay does not mean every payer switches. In the reporting sample, 81% of invoice value is currently paid without cards.
  • 0.7% is card economics, not blended economics. ACH earns $1.60 per payment. On a $3,800 commercial invoice, that is 0.042%, around one-seventeenth of the assumed yield. Checks and cash earn nothing.
  • Customer count is a poor proxy for accessible volume. The largest 5% generate 48% of invoicing. Of those 115 customers, 71 have processor contracts running mostly to 2028. Seventy per cent adoption among smaller customers could leave much of the revenue pool untouched.

Our mix compounds the problem. Commercial work represents 58% of invoice value. Commercial card share barely changed in the pilot, from 5% to 6%, despite pay-by-link on every invoice. Residential card share increased to 71%, but the pilot was 64% residential against 42% across Fieldline.

The customer problem is also different from the proposed solution. Seventeen of 22 interviewees identified slow commercial payment as their biggest cash problem. Offering another payment method does not establish that property managers will abandon net-45 or net-60 terms.

Tradesly shows payments can work for a business with approximately 80% residential customers. Its 22% revenue contribution does not validate our volume, adoption or margin assumptions.

What the supplied numbers support

Use:

Annual revenue = eligible invoice value × adopting share of that value × realised net yield.

The following are scenarios, not statistical bounds. Neither the exact invoicing of the 71 contracted accounts nor representative payment capture is supplied.

InputLower scenarioUpper scenario
Eligible annual invoice value$2.51B: exclude all largest 115 accounts$3.40B: exclude the 71 contracted accounts, assuming equal volume within the top 115
Adoption by invoice value61%, matching pilot customer adoption70%, matching the plan
Card share by valueCurrent reported 19%33.3%: 42% × 71% residential + 58% × 6% commercial
ACH treatmentCurrent reported 44% of valueEvery remaining payment becomes ACH
Modelled net yield0.216%0.304%
Annual revenue$3.3M$7.2M

Working:

Lower yield:

`19% × 0.7% + 44% × $1.60 × (42% ÷ $410 + 58% ÷ $3,800) = 0.216%`

`$2.51B × 61% × 0.216% = $3.3M`

Upper yield:

`33.3% × 0.7% + $1.60 × (42% × 29% ÷ $410 + 58% × 94% ÷ $3,800) = 0.304%`

`$3.40B × 70% × 0.304% = $7.2M`

These calculations assume eligible customers retain the overall residential/commercial mix and that adoption rates by customer also apply by value. The lower scenario assumes 44% ACH within both job types. Both assume full capture of the payment streams modelled. The upper scenario additionally interprets the pilot’s card shares as shares of value; if they are transaction counts, that extrapolation is unsupported. It also assumes complete conversion of checks and cash to ACH.

Consequently, $3.3M is not a floor. The reporting sample covers only 30% of customers, and actual volume concentration or capture could reduce revenue further.

A separate check points in the same direction: scaling the pilot’s $5,630 annual revenue per adopter to 61–70% of 2,300 customers produces $7.9M–$9.1M, before correcting its residential bias or account availability.

On the upper scenario’s eligible volume and 70% adoption, $23.7M would require approximately 1.0% net revenue on every adopted invoice dollar, above even the quoted card margin. Growth could alter this, but the pack supplies no growth assumptions to bridge the gap.

These are annualised net payments revenues under the partner terms, not profit or contracted software ARR.

What would prove this challenge wrong

We would need evidence that accessible customers generate materially more net revenue per total invoiced dollar than these scenarios imply, at sufficient committed volume to support $24M. That could come from much greater accessible volume, different economics or a demonstrably valuable paid service. Enthusiasm and payment-feature activation would not suffice.

The cheapest six-week test is to extend the existing pilot, with analyst, sales and limited engineering support:

  1. Establish the accessible denominator. Match the 71 processor contracts to actual invoice value, expiry dates and any permitted carve-outs. Replace the equal-volume assumption with account-level evidence.
  2. Recruit a stratified cohort. Include commercial-heavy businesses and large uncontracted accounts, alongside residential firms. Count refusals in adoption results and weight findings to eligible invoice value.
  3. Measure actual money movement. Track total invoicing, activated volume, captured volume, payment methods, realised fees and net revenue. Report residential and commercial results separately. Use existing receivables approaching their due dates to observe commercial payments within six weeks.

Before recruitment, agree that expanding investment requires a volume-weighted model supported by observed payments and credible customer commitments. To justify the original commitment, that model must show a defensible route to $24M by month 24. Six weeks cannot establish mature adoption or commercial collection cycles; an inconclusive result should not trigger automatic approval.

How the bet should change

If validation confirms residential economics, pursue a narrower residential payments offering with staged staffing and revenue gates. Separately test whether commercial customers will make concrete commitments to a faster-payment proposition. Nine expressions of willingness to pay justify discovery, not a financing product.

Keep the scheduling rewrite funded. The affected segment lost $1.9M ARR last year, with scheduling the top-cited churn reason. That is not all recoverable, but pausing the work has a documented cost against a speculative upside.

We should update the board with a scenario range and validation milestones. With the raise approximately 14 months away, protecting credibility and existing revenue matters more than preserving an unsupported headline.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisQuick edit

    What we’d changeFrame the range up front as illustrative full-capture scenarios, with actual capture still unmeasured, so it isn't read as a forecast.

Check by check

Got right · 11
  • Uses the supplied evidence correctlyEvery factual statement about the current situation is drawn directly from the supplied context or follows from it by arithmetic; nothing about the current situation is invented.
  • Addresses the actual decisionThe memo commits early to 'do not approve three squads', frames it for the executive team, and states exactly what evidence would change that (a defensible route to $24M from a volume-weighted model).
  • Respects explicit constraintsThe output is a memo under 1,200 words, addresses all four required points, and respects the reader and the decision context; it does not propose anything that would violate an explicit constraint.
  • Identifies material uncertaintyIt names the key unknowns (contractual lock-in, actual captured volume, commercial card uptake), bounds them with scenario analysis, and describes a test that would resolve them and change the call.
  • Avoids unsupported claimsScenarios are clearly labeled as such, assumptions are stated explicitly, and no forecast or cause is presented as established fact without evidence.
  • Produces the required deliverableA complete memo in the right format, word count within limit, addressed to the CPO and exec team, with sections that answer each prompt point; the reader can act on it directly.
  • Finds the load-bearing assumptionThe central challenge is the mismatch between the 0.7% card-based take rate and the actual payment mix, substantiated with ACH/card percentages and unit economics from the pack.
  • Re-estimates the revenue correctlyIt recalculates revenue as $3.3M–$7.2M using the payment mix and unit economics, shows working, notes the pilot-based upper bound of ~$9M, and explains why even that is an overestimate.
  • Says what would prove it wrongIt proposes a concrete six-week cohort test with specific measurements and a clear threshold: a volume-weighted model that shows a defensible path to $24M before expanding investment.
  • Reshapes the bet from the evidenceIt recommends a narrower residential payments offering, a separate test for commercial faster-payment demand, and keeping the scheduling rewrite funded, directly grounded in the residential-mix evidence and commercial cash-flow pain.
  • Tigers, not paper tigersIt dismisses the Tradesly comparison as irrelevant for a different mix, elevates the take-rate mismatch as the central risk, and surfaces the hidden cost of pausing the scheduling rewrite ($1.9M lost ARR) that the CEO's plan avoided.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly9%
  • passAddresses the actual decision99%
  • passRespects explicit constraints55%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims45%
  • passProduces the required deliverable75%
  • passFinds the load-bearing assumption100%
  • passRe-estimates the revenue correctly31%
  • partialSays what would prove it wrong49%
  • passReshapes the bet from the evidence96%
  • passTigers, not paper tigers93%
Run
Run
#1
Time to output
80 s
Submitted
25 Sept 2026

Needs more than a quick edit9 of 11 checks passed

The memo is factually strong and gives a clear go/no-go call, but its 'instead' is generic and misses the commercial cash-flow/faster-payout opportunity from the interviews. A PM would need to add that alternative reasoning rather than just edit wording.

Output

To: Priya Raman Subject: Fieldline Pay: challenge the $24M case before committing three squads

The strongest case against the current plan is not that embedded payments cannot work. It is that the $24M forecast treats invoice volume as if it were payment volume earning a card-like margin.

The critical unsupported assumption: 70% of customers will adopt and route enough of their invoices through Fieldline Pay to earn roughly 0.7% on that value. The 0.7% net take is specifically the card margin. ACH earns $1.60 per payment; checks and cash earn nothing. The CFO’s arithmetic is correct only if essentially every dollar from every adopting customer earns the card margin.

That is a poor fit with our mix. Commercial work is 58% of invoiced value; commercial clients typically require ACH or check, pay after 52 days, and may not accept a card surcharge on a $3,800 invoice. Eleven of 14 interviewed mostly-commercial customers said their clients require ACH or check. By contrast, the pilot’s customers were 64% residential, against 42% for Fieldline overall. And 71 of the 115 largest customers have processor contracts running mostly to 2028; the largest 5% of customers account for 48% of invoiced value. The pilot’s 61% adoption is encouraging, but it does not establish broad adoption or broad monetization. Tradesly’s 22% revenue share is also a weak comparator: its customer base is about 80% residential.

A more defensible annualized revenue range

The following estimates are run-rate scenarios, not a forecast. They use the supplied 61% pilot adoption rate and 70% plan assumption as bounds, and make explicit that the take depends on payment type.

Card revenue: Current recorded payment methods show 19% of invoice value paid by card. In the pilot, the reported card shares imply a 33.3% value-weighted card share if applied to our mix: 42% residential × 71% + 58% commercial × 6% = 33.3%.

At a 0.7% net card take:

  • Lower case: $4.83B × 61% adoption × 19% card share × 0.7% = $3.9M
  • Upper case: $4.83B × 70% adoption × 33.3% card share × 0.7% = $7.9M

The upper case assumes pilot card behavior transfers to the full customer mix, despite the pilot’s residential skew.

ACH revenue: The recorded-method sample reports ACH at 44% of invoice value, but value share does not tell us the number of $2 ACH payments—and the pilot does not report ACH usage. As an illustrative bound, if that 44% share transferred to adopters, ACH revenue would be about $0.5M–$2.1M after the $0.40 cost: the lower end assumes ACH dollars are mostly $3,800 commercial invoices; the upper end allows the pilot’s card mix and available non-card residential volume. This is uncertain upside, not a reliable forecast.

Together, that supports roughly $4.5M–$10M of annualized revenue, before accounting for contract lock-in or slower adoption among large accounts.1 The range is still conditional on transferring today’s payment mix and pilot behavior to adopters. As a cross-check, the pilot’s $214K annualized revenue across 38 customers is $5,630 per customer; scaling that to 61%–70% adoption gives $7.9M–$9.1M. That sits within the range, but is not a clean base case because the pilot skewed residential.

The $24M case requires $14,700 per adopting customer per year ($2.1M × 0.7%). That is about 2.6× the pilot’s $5,630, while assuming adoption and monetization across a materially more commercial customer base.

What would change my mind—and the cheapest six-week test

I would support a larger bet if a representative cohort showed both credible adoption and enough net revenue per adopter to approach the $14,700 annualized level the board case requires—especially among commercial customers and larger accounts not locked into existing processors. A high signup rate alone would not prove the economics.

Run a six-week, no-build test using the existing pay-by-link and partner setup. Recruit a small, stratified cohort across residential and commercial customers, including larger unlocked accounts. Track offers accepted, invoice value actually routed, card versus ACH mix, net revenue after costs, and payment timing. Report results by segment; do not annualize a handful of unusually large or residential-heavy customers as if they were representative. Include a holdout or compare with each customer’s recent invoices to distinguish new payment volume from existing behavior.

Recommendation

Do not commit three squads for three quarters or pause the scheduling rewrite on the current evidence. The rewrite addresses the top-cited churn reason among customers with more than 50 technicians, a segment that lost $1.9M ARR last year. Run the six-week test with a small team and existing capabilities, then make a stage-gated decision. If commercial routing and net revenue per adopter are strong, fund the next increment. If not, preserve the scheduling investment and narrow Pay to the segments where the data shows customers can—and will—use it.

What a PM had to fix

  1. 1OtherFix itTargeted repair

    What we’d changeThese assumptions produce an illustrative $4.5M–$10M envelope, not a validated revenue range. They assume customer adoption translates proportionally into invoice volume and the modelled payments route through Pay. Rebuild by segment: accessible invoice value × volume-weighted adoption × routing share × payment-method economics, excluding contracted volume.

Check by check

Got wrong · 1
  • Reshapes the bet from the evidenceThe recommendation is a generic stage-gate/narrow-to-segments move and does not propose the specific alternative the evidence supports, such as residential card focus or commercial cash-flow/faster payout.
Mixed · 1
  • Uses the supplied evidence correctlyAll facts and arithmetic trace to the supplied pack, with no invented current-state figures found.The two graders disagreed on this one.
Got right · 9
  • Addresses the actual decisionThe memo commits clearly to not committing now, proposes a stage-gated test, and states what would change the call.
  • Respects explicit constraintsIt is a memo addressed to Priya, within the requested length, and covers the four required elements.
  • Identifies material uncertaintyIt names payment-mix, adoption, contract lock-in, and pilot skew as unknowns and says how a six-week test would resolve them.
  • Avoids unsupported claimsEstimates are labelled as scenarios/bounds, and generalizations are tied to interview and sales evidence rather than asserted as fact.
  • Produces the required deliverableThe memo is a complete, usable pre-read for the exec audience with a clear recommendation and supporting analysis.
  • Finds the load-bearing assumptionIt makes the take-rate/payment-mix mismatch the central challenge, using the card, ACH, and check/cash unit economics from the pack.
  • Re-estimates the revenue correctlyIt recalculates revenue from the payment mix and pilot, lands in the $4.5-10M range, and explains why the pilot is an upper bound.
  • Says what would prove it wrongIt specifies a six-week stratified cohort test with net revenue per adopter approaching $14,700 as the threshold to change its mind.
  • Tigers, not paper tigersIt triages the take-rate mismatch as the sink, dismisses Tradesly as a weak comparator, and surfaces the scheduling-rewrite churn risk.

Grades and run details

Decision model 86 · LLM judge 10 of 12 checks
Decision model checks
  • failUses the supplied evidence correctly3%
  • passAddresses the actual decision95%
  • passRespects explicit constraints57%
  • passIdentifies material uncertainty98%
  • passAvoids unsupported claims31%
  • passProduces the required deliverable57%
  • passFinds the load-bearing assumption97%
  • passRe-estimates the revenue correctly82%
  • passSays what would prove it wrong79%
  • partialReshapes the bet from the evidence44%
  • passTigers, not paper tigers89%
Run
Run
#1
API response time
73 s
Submitted
29 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.095.82None
2GPT-6 AstrawithChatGPT95.595.82None
3GPT-6 LunawithAPI93.291.72None
4Sonnet 5.5withAPI93.275.02None
5Gemini 3.5 Flash-LitewithGemini81.875.021 capped
6Opus 5.5withClaude90.991.721 capped

About the task

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Separates the risks that could kill it from the ones that only look scary
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill · Staff level: a company bet with a long evidence pack

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.