Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  2. Identifies material uncertainty100% pass
    It names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  3. Finds the load-bearing assumption100% pass
    It makes the 0.7% blended take applied to all invoice value the central challenge, backed by payment-mix and unit-economics figures.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet

Where it slips

  1. A cheap test that can actually read out67% pass
    The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.
    GPT-6.1 Sol · API · An AI SDR for small agencies
  2. Avoids unsupported claims77% pass
    The memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
    Opus 5.5 · Claude · The CEO's embedded-payments bet
  3. Uses the supplied evidence correctly79% pass
    The memo presents the pilot blended net take of ~0.38% as an achieved fact, but it is not supported by the pack and is an assumption-driven extrapolation.
    Opus 5.5 · Claude · The CEO's embedded-payments bet

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We plan to sell an AI sales-development rep, an agent that finds prospects and sends personalised outbound email, to agencies with 5–20 staff. Before we commit a year to it, find the strongest reason this fails. Use the interview material below, including the three transcripts. Write it up for the founding team: the one reason, the evidence for it (quote the interviews), where the evidence cuts the other way, what we would need to see to be proved wrong, and the cheapest test that would show it. Keep it under 600 words.

ScenarioWe are a three-person founding team with £800k of pre-seed funding and about 20 months of runway. Planned price: $600 a month per agency. Two competitors have each raised over $10m in the last year, selling AI SDRs mostly to software companies with 50+ staff.
Founders' hypothesis (from our planning doc)“Small agencies live feast or famine. When a big client leaves they have no pipeline, because the founder is too busy delivering to sell. An AI SDR keeps the pipeline full in the background, so the famine never comes. At $600 a month, one extra client a year pays for it many times over.”
Interview summary (14 agencies, 5–20 staff, last six weeks)11 of 14 get most of their new revenue from referrals and repeat clients (86% on average across those 11). 9 had tried outbound in the last two years; 7 of those stopped within six months, and none of the 7 closed a deal they could attribute to it. The 2 who kept going (A04, A07) have 15–20 staff and a dedicated person for business development. 3 (A08, A12, A14) said they regularly turn work away. 6 described feast-or-famine swings in new business. Average deal size across the 14: $18k; typical sales cycle 6–10 weeks.
Interview logAgency (staff) · services · share of new revenue from referrals and repeat clients · outbound tried? · outcome · representative quote A01 (8) · brand and web studio · 90% · yes, cold-email agency · stopped after 3 months · "We got meetings with people who'd never buy from us." A02 (14) · performance marketing · 85% · yes, LinkedIn automation · stopped after 4 months, account restricted · "LinkedIn shut our founder's account down. That's our best referral network." A03 (6) · PR boutique · 95% · no · — · "Every client we have came from someone vouching for us." A04 (18) · Shopify development · 40% (plus 30% partners, 30% outbound) · yes, in-house BD lead · kept, 30% of new revenue · "It works because of the follow-up, not the first email." A05 (11) · content · 85% · yes, lead-gen agency · stopped after 2 months, no deals · "Content is a trust purchase." A06 (7) · video production · 85% · no · — · "Our clients find us through the videos. Someone shares one and we get a call." A07 (16) · SEO · 45% · yes, outsourced lead generation · kept · "It pays for itself, just. Most of the meetings are a waste of time, but one in ten turns into a retainer." A08 (5) · packaging design · 95% · no · — · "If you sent me ten more leads in October I'd have to say no to nine of them." A09 (12) · paid social · 75% · yes, cold email · stopped after 5 months · "The replies were mostly people asking us to take them off the list." A10 (20) · B2B marketing · 70% · yes, contract SDR · stopped after 6 months · "It cost us about £400 for every meeting, and the meetings didn't close." A11 (9) · WordPress web agency · 90% · yes, cold email · stopped after 3 months · "Our domain ended up on a spam list. Took weeks to fix." A12 (10) · branding · 85% · no · — · "We're booked out till March. I don't need more leads, I need another designer." A13 (15) · HubSpot partner · 48% (plus 40% from HubSpot's partner directory) · yes, cold email · stopped after 4 months · "HubSpot sends us more than we can handle in a good quarter." A14 (6) · UX research · 90% · no · — · "We're small on purpose. We say no to about a third of enquiries."
interview-A05-content-agency.md31 lines · Download
# Interview A05: content agency, 11 staff
Participant: founder and managing director. Interviewer: our co-founder. 34 minutes, lightly edited.

**Interviewer:** Where did your last five clients come from?

**Participant:** Let me think. Two were old clients coming back: one had changed jobs and brought us into her new company. Two were introductions, one from a web agency we partner with and one from a client's CFO who'd seen our work. The fifth found us through a talk I gave at a SaaS marketing meetup. So none of them from anything you'd call outbound.

…
interview-A04-shopify-agency.md23 lines · Download
# Interview A04: Shopify development agency, 18 staff
Participant: head of business development. Interviewer: our co-founder. 29 minutes, lightly edited.

**Interviewer:** Tell me how new business works for you.

**Participant:** We're a bit unusual for an agency our size. I'm a full-time business-development person, and outbound is about 30% of our new revenue. Referrals and repeat work are still the majority, maybe 40% referrals and the rest partners: Shopify's partner directory and a couple of app companies who send us projects.

…
interview-A08-packaging-studio.md23 lines · Download
# Interview A08: packaging design studio, 5 staff
Participant: founder and creative director. Interviewer: our co-founder. 22 minutes, lightly edited.

**Interviewer:** How do new clients find you?

**Participant:** Almost entirely word of mouth. Food and drink brands talk to each other. Someone launches a range, their packaging does well on the shelf, and the next founder asks who did it. I'd say 95% referral. We've never done any outbound.

…
What a strong answer does

Names the load-bearing assumption as 'small agencies' growth is limited by a shortage of new leads': the interviews say it is limited by trust and capacity. 11 of 14 win most work through referrals and repeat clients; 7 of 9 who tried outbound quit within six months with no attributable deals, for reasons tied to trust (A05's 'content is a trust purchase', A01's meetings with people who'd never buy) and reputation risk (A02's restricted LinkedIn account, A11's spam listing); 3 turn work away (A08: 'I'd have to say no to nine of them'). Quotes the interviews accurately and attributes them correctly. Engages the counter-evidence honestly: outbound works for A04 and A07, but both are 15–20 staff with a dedicated BD owner, already use tools like Apollo and Clay, and credit the follow-up rather than the first email, so the segment is small, tooled and not what the product automates. Treats feast or famine as a real pain but notes agencies fix it through past clients and referrals (A05's reactivation emails, and her willingness to pay for help keeping in touch with past clients), pointing at a different product. Says what would prove the challenge wrong (for example, a meaningful share of 10–20-staff agencies with a BD owner and budget, or a concierge outbound test producing qualified meetings that turn into proposals), and designs a cheap test that measures an early signal, since a 6–10 week sales cycle means a six-week test cannot observe won deals.

Critical failures (cap the score)
  • Invents interview quotes or attributes a quote to the wrong agency
  • Treats the founders' feast-or-famine hypothesis as confirmed by the interviews
Case

v1.8 · synthetic · B2B, go-to-market, customer research

Needs more than a quick edit7 of 11 checks passed

It is over the requested length and contains two unsupported current-state claims that an editor would need to remove or verify. Otherwise the reasoning is usable, but those are more than wording fixes.

Output

Pre-mortem: AI SDR for 5–20 person agencies

The one reason it fails

Small agencies do have a pipeline problem, but cold outbound is the wrong fix. Their work is sold on trust, so it comes through introductions. Cold email produces meetings that don't close, and the few agencies where outbound works have a person doing the follow-up. At $600 a month we would be selling a channel that this buyer has already tried and dropped, into a market where two funded competitors are moving down.

Evidence

  • Referrals dominate. 11 of 14 agencies get about 86% of new revenue from referrals and repeat clients. A05: "Content is a trust purchase." A03: "Every client we have came from someone vouching for us."
  • Outbound has a poor track record. 9 tried it, 7 quit within six months, and none of the 7 could attribute a closed deal to it. The reasons are structural: A01 "meetings with people who'd never buy from us"; A09 "mostly people asking us to take them off the list"; A11 "Our domain ended up on a spam list"; A02 "LinkedIn shut our founder's account down. That's our best referral network."
  • The cost is high. A10: "about £400 for every meeting, and the meetings didn't close."
  • Success depends on humans. A04: "It works because of the follow-up, not the first email… The AI could save me the research hours, sure. It can't do the call." And: "most agencies our size don't have someone like me… Most tried it once, got burned, and went back to referrals."
  • The famine is real, but warm outreach solved it. A05 recovered by emailing three years of past clients, and two came back within a month. Of our product, A05 said: "Cold? Probably not… Six hundred for cold outbound, no."
  • Some agencies don't need leads. A08, A12 and A14 turn work away. A08: "It would make my problem worse."

Where the evidence cuts the other way

  • 6 of 14 report feast-or-famine, so the pain is real.
  • A07 kept outbound: "It pays for itself, just… one in ten turns into a retainer." A04 would pay $600 if the tool replaced Apollo and Clay.
  • None of the failures used a modern AI agent. Better targeting and lower cost per meeting might change the results, though the trust and deliverability complaints suggest it won't fix most of them.
  • The two who kept going are not cleanly separable by size. Of the four agencies with 15–20 staff, two stopped (A10, A13).
  • A05 would pay about £200 a month for a warm-relationship tool. That points to a different product at a third of our price, not proof that ours works.
  • 14 interviews is a small sample, and willingness to pay was hypothetical.

What would prove us wrong

In a paid pilot with agencies that have no BD hire, the product produces qualified pipeline at a cost they'll keep paying for. Roughly: at least half the pilot agencies get a qualified opportunity or proposal within 90 days, and most say they'd continue at $600. The maths also needs checking. $600 a month is $7,200 a year, but the buyer's gain is the margin on an $18k deal, not the revenue. "Many times over" only holds if outbound reliably closes more than one deal a year.

Cheapest test

Run a concierge pilot before building anything more:

  1. Recruit 10 agencies with 5–14 staff and no BD person. Exclude the referral-saturated ones like A08, A12 and A14.
  2. Charge $600 a month, or at least a serious deposit. Free pilots tell us nothing about willingness to pay.
  3. Do the work ourselves with off-the-shelf tools (Apollo, Clay, an email sender) on separate sending domains. The agency approves every email, since A04 and A05 both demanded that.
  4. Set the pass line before starting: at least 5 of 10 agencies get a qualified proposal within roughly 14 weeks (the 6–10 week sales cycle plus setup), and at least 6 of 10 agree to continue.
  5. Track reply rate, unsubscribe and spam complaints, cost per qualified meeting, and meeting-to-proposal rate.

This costs a few weeks of founder time and about £1–2k in tools, roughly 4 months of the 20-month runway. In parallel, offer 5 of the same agencies a £200 warm-outreach version, to see whether the real product is relationship follow-up.

If the pilot misses the line, don't spend the year on cold outbound for this segment.

Check by check

Got wrong · 2
  • Respects explicit constraintsIt exceeds the 600-word limit, running roughly 700 words, even though it includes the requested sections.
  • Avoids unsupported claimsIt presents the competitor 'moving down' claim and the 'none used a modern AI agent' claim as fact rather than as labelled inference.
Mixed · 2
  • Uses the supplied evidence correctlyIt includes at least two current-state claims, that competitors are moving down and that none of the failures used a modern AI agent, that are not supported by the supplied context.The two graders disagreed on this one.
  • Produces the required deliverableIt has the right form and reader, but it is not within the requested length.The two graders disagreed on this one.
Got right · 7
  • Addresses the actual decisionIt commits early to the answer that cold outbound is the wrong fix for this segment and says what pilot result would change that.
  • Identifies material uncertaintyIt names the small sample, hypothetical willingness to pay, and the untested modern-AI possibility, and proposes thresholds and observable signals to resolve them.
  • Finds the load-bearing assumptionIt targets the assumption that growth is lead-constrained, using the referral, abandoned-outbound and turning-work-away evidence as the central challenge.
  • Uses the interviews faithfullyAll quotes or close paraphrases match the transcripts and log, and each is attributed to the correct agency.
  • Engages the counter-evidenceIt names A04 and A07, identifies their size/BD-owner/tooling and follow-up dependence, and explains what that means for the addressable segment.
  • A cheap test that can actually read outIt proposes a cheap concierge pilot with explicit thresholds and measures qualified proposals and meetings rather than won deals, matching the 6–10 week sales cycle.
  • Tigers, not paper tigersIt separates the central trust and capacity risks from fixable targeting/cost issues and surfaces the warm-outreach product alternative and margin-math problem.

Claims the judge couldn’t find in the brief

  • Two funded competitors are moving down into this market.
  • None of the failed outbound attempts used a modern AI agent.

Grades and run details

Decision model 91 · LLM judge 7 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly37%
  • passAddresses the actual decision95%
  • partialRespects explicit constraints30%
  • passIdentifies material uncertainty100%
  • partialAvoids unsupported claims25%
  • passProduces the required deliverable79%
  • passFinds the load-bearing assumption86%
  • passUses the interviews faithfully83%
  • passEngages the counter-evidence75%
  • passA cheap test that can actually read out99%
  • passTigers, not paper tigers98%
Run
Run
#1
API response time
30 s
Submitted
29 Sept 2026
Critical failure: Treats the founders' feast-or-famine hypothesis as confirmed by the interviews. Score capped at 40.

Needs more than a quick edit6 of 11 checks passed

The output includes an unsupported revenue claim about A07 and proposes a test for a different product, so a PM would need to recheck facts and redesign the test.

Output

Why an AI SDR for Small Agencies Will Fail: Pre-Mortem Analysis

For: Founding Team (£800k runway, 20 months remaining) Product: AI Sales-Development Rep ($600/month) targeting agencies with 5–20 staff

---

1. The One Reason It Fails

Small agencies do not suffer from a lack of prospecting volume; they suffer from a lack of trust and execution capacity. Cold outbound is fundamentally the wrong tool for their business model. Because creative and professional services are "trust purchases," agencies rely overwhelmingly on referrals and repeat clients. Automated cold emails damage their reputation rather than building a reliable pipeline.

---

2. The Evidence (Interview Quotes)

Trust & Referrals Outweigh Cold Acquisition

  • “Content is a trust purchase: you're letting someone write in your company's voice… When the CFO introduced us, the first call was about scope, not about whether we were any good.” — A05 (Content Agency)
  • “Every client we have came from someone vouching for us.” — A03 (PR Boutique)

Cold Outbound Fails Small Agencies

  • “The replies were mostly people asking us to take them off the list.” — A09 (Paid Social Agency)
  • “Our domain ended up on a spam list. Took weeks to fix.” — A11 (WordPress Agency)
  • “It cost us about £400 for every meeting, and the meetings didn't close.” — A10 (B2B Marketing Agency)

The Real Bottleneck is Capacity and Follow-Up, Not Leads

  • “If you sent me ten more leads in October I'd have to say no to nine of them.” — A08 (Packaging Design Studio)
  • “The AI could save me the research hours, sure. It can't do the call... and the follow-up is the job.” — A04 (Shopify Development Agency)

---

3. Where the Evidence Cuts the Other Way

  • The Pain is Real: Six of the 14 interviewed agencies explicitly described painful "feast-or-famine" revenue swings, validating the founders' premise that pipeline stability is a major issue.2
  • Outbound Can Work: Two agencies (A04 and A07) successfully use outbound to drive 30% to over 50% of new revenue4, proving that agency outbound can function under specific circumstances.3

---

4. What We Would Need to See to Be Proved Wrong

To believe this product can succeed at $600/month for 5–20 staff agencies, we would need to see empirical proof that small agencies without a dedicated in-house business development lead can: 1. Generate positive ROI from cold AI outreach without harming their domain reputation. 2. Convert cold prospects into high-trust creative/professional services contracts without a human-led multi-channel follow-up loop.

---

5. The Cheapest Test

Run a targeted messaging experiment to 50 agencies with 5–20 staff. Pivot the pitch away from a cold-outbound AI SDR toward what A05 actually requested: an automated relationship assistant that tracks past clients and referral networks, drafting warm check-in notes.

If agencies will not engage with or pre-order a warm-relationship retention tool at $200–$600/month, they will certainly reject a cold AI prospecting tool.1

What a PM had to fix

  1. 1Test or gate too weakTighten the testSubstantial rework

    What we’d changeTest the reason the memo gives: whether agencies without a BD person can turn AI-sent cold email into deals. Demand for a different product doesn't settle that.

  2. 2Hypothesis stated as factReframe it as a hypothesisTargeted repair

    Source: Founders' hypothesis (from our planning doc)Small agencies live feast or famine.

    What we’d changeSix of 14 is some agencies, not validation of the founders' premise. Say that.

  3. 3Contradiction missedSurface the contradictionTargeted repair

    What we’d changeSay what sets A04 and A07 apart: a dedicated BD person doing the follow-up. That's the evidence for the memo's own reason, and it's left unused.

  4. 4Invented evidenceVerify or remove the claimQuick edit

    What we’d changeRemove 'over 50%': A07's share of new revenue from outbound isn't given anywhere.

Check by check

Got wrong · 4
  • Avoids unsupported claimsThe unsupported claim about A07's outbound revenue share is presented without qualification, and the test assumes demand for a warm relationship tool rather than testing the AI SDR assumption.
  • Produces the required deliverableThe proposed test does not test the viability of the AI SDR; it tests demand for a different product, so the deliverable cannot be acted on without reworking the test.
  • Engages the counter-evidenceThe output mentions A04 and A07 but does not say what sets them apart (15-20 staff, dedicated BD owner, follow-up doing the work) or what that means for addressable segment, as required.
  • A cheap test that can actually read outThe proposed test shifts to a warm relationship tool, which does not test the assumption behind the AI SDR or measure an early signal for cold outbound; it doesn't fit the brief's request.
Mixed · 1
  • Uses the supplied evidence correctlyThe claim that A07 drives 'over 50% of new revenue' from outbound is not supported by the supplied context; no such figure appears for A07.The two graders disagreed on this one.
Got right · 6
  • Addresses the actual decisionThe output commits to a clear answer (the product fails) and states what would change the call (empirical proof of positive ROI and conversion without human follow-up).
  • Respects explicit constraintsThe output is under 600 words, addresses the founding team, and includes the requested sections.
  • Identifies material uncertaintyIt names conditions that would prove the idea wrong and outlines what evidence would be needed.
  • Finds the load-bearing assumptionThe output identifies that small agencies' growth is not limited by a shortage of leads, but by trust and capacity, which is the load-bearing assumption.
  • Uses the interviews faithfullyAll quotes are accurate and correctly attributed to the right agencies, with no invented quotes.
  • Tigers, not paper tigersThe output ranks capacity and trust over mere lead volume, dismisses the feast-or-famine pain as solvable through past clients, and surfaces the real product need (warm outreach) that the founders' proposal avoids.

Claims the judge couldn’t find in the brief

  • Two agencies (A04 and A07) successfully use outbound to drive 30% to over 50% of new revenue.

Grades and run details

Decision model 77 · LLM judge 6 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly37%
  • passAddresses the actual decision84%
  • partialRespects explicit constraints24%
  • passIdentifies material uncertainty35%
  • partialAvoids unsupported claims16%
  • partialProduces the required deliverable46%
  • passFinds the load-bearing assumption94%
  • passUses the interviews faithfully78%
  • partialEngages the counter-evidence90%
  • partialA cheap test that can actually read out88%
  • passTigers, not paper tigers66%
Run
Run
#1
Time to output
7 s
Submitted
25 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.095.82None
2GPT-6 AstrawithChatGPT95.595.82None
3GPT-6 LunawithAPI93.291.72None
4Sonnet 5.5withAPI93.275.02None
5Gemini 3.5 Flash-LitewithGemini81.875.021 capped
6Opus 5.5withClaude90.991.721 capped

About the task

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Separates the risks that could kill it from the ones that only look scary
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill · Staff level: a company bet with a long evidence pack

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.