Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 12 graded outputs by 6 models. 75% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  2. Identifies material uncertainty100% pass
    It names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  3. Finds the load-bearing assumption100% pass
    It makes the 0.7% blended take applied to all invoice value the central challenge, backed by payment-mix and unit-economics figures.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet

Where it slips

  1. A cheap test that can actually read out67% pass
    The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.
    GPT-6.1 Sol · API · An AI SDR for small agencies
  2. Avoids unsupported claims77% pass
    The memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
    Opus 5.5 · Claude · The CEO's embedded-payments bet
  3. Uses the supplied evidence correctly79% pass
    The memo presents the pilot blended net take of ~0.38% as an achieved fact, but it is not supported by the pack and is an assumption-driven extrapolation.
    Opus 5.5 · Claude · The CEO's embedded-payments bet

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We plan to sell an AI sales-development rep, an agent that finds prospects and sends personalised outbound email, to agencies with 5–20 staff. Before we commit a year to it, find the strongest reason this fails. Use the interview material below, including the three transcripts. Write it up for the founding team: the one reason, the evidence for it (quote the interviews), where the evidence cuts the other way, what we would need to see to be proved wrong, and the cheapest test that would show it. Keep it under 600 words.

ScenarioWe are a three-person founding team with £800k of pre-seed funding and about 20 months of runway. Planned price: $600 a month per agency. Two competitors have each raised over $10m in the last year, selling AI SDRs mostly to software companies with 50+ staff.
Founders' hypothesis (from our planning doc)“Small agencies live feast or famine. When a big client leaves they have no pipeline, because the founder is too busy delivering to sell. An AI SDR keeps the pipeline full in the background, so the famine never comes. At $600 a month, one extra client a year pays for it many times over.”
Interview summary (14 agencies, 5–20 staff, last six weeks)11 of 14 get most of their new revenue from referrals and repeat clients (86% on average across those 11). 9 had tried outbound in the last two years; 7 of those stopped within six months, and none of the 7 closed a deal they could attribute to it. The 2 who kept going (A04, A07) have 15–20 staff and a dedicated person for business development. 3 (A08, A12, A14) said they regularly turn work away. 6 described feast-or-famine swings in new business. Average deal size across the 14: $18k; typical sales cycle 6–10 weeks.
Interview logAgency (staff) · services · share of new revenue from referrals and repeat clients · outbound tried? · outcome · representative quote A01 (8) · brand and web studio · 90% · yes, cold-email agency · stopped after 3 months · "We got meetings with people who'd never buy from us." A02 (14) · performance marketing · 85% · yes, LinkedIn automation · stopped after 4 months, account restricted · "LinkedIn shut our founder's account down. That's our best referral network." A03 (6) · PR boutique · 95% · no · — · "Every client we have came from someone vouching for us." A04 (18) · Shopify development · 40% (plus 30% partners, 30% outbound) · yes, in-house BD lead · kept, 30% of new revenue · "It works because of the follow-up, not the first email." A05 (11) · content · 85% · yes, lead-gen agency · stopped after 2 months, no deals · "Content is a trust purchase." A06 (7) · video production · 85% · no · — · "Our clients find us through the videos. Someone shares one and we get a call." A07 (16) · SEO · 45% · yes, outsourced lead generation · kept · "It pays for itself, just. Most of the meetings are a waste of time, but one in ten turns into a retainer." A08 (5) · packaging design · 95% · no · — · "If you sent me ten more leads in October I'd have to say no to nine of them." A09 (12) · paid social · 75% · yes, cold email · stopped after 5 months · "The replies were mostly people asking us to take them off the list." A10 (20) · B2B marketing · 70% · yes, contract SDR · stopped after 6 months · "It cost us about £400 for every meeting, and the meetings didn't close." A11 (9) · WordPress web agency · 90% · yes, cold email · stopped after 3 months · "Our domain ended up on a spam list. Took weeks to fix." A12 (10) · branding · 85% · no · — · "We're booked out till March. I don't need more leads, I need another designer." A13 (15) · HubSpot partner · 48% (plus 40% from HubSpot's partner directory) · yes, cold email · stopped after 4 months · "HubSpot sends us more than we can handle in a good quarter." A14 (6) · UX research · 90% · no · — · "We're small on purpose. We say no to about a third of enquiries."
interview-A05-content-agency.md31 lines · Download
# Interview A05: content agency, 11 staff
Participant: founder and managing director. Interviewer: our co-founder. 34 minutes, lightly edited.

**Interviewer:** Where did your last five clients come from?

**Participant:** Let me think. Two were old clients coming back: one had changed jobs and brought us into her new company. Two were introductions, one from a web agency we partner with and one from a client's CFO who'd seen our work. The fifth found us through a talk I gave at a SaaS marketing meetup. So none of them from anything you'd call outbound.

…
interview-A04-shopify-agency.md23 lines · Download
# Interview A04: Shopify development agency, 18 staff
Participant: head of business development. Interviewer: our co-founder. 29 minutes, lightly edited.

**Interviewer:** Tell me how new business works for you.

**Participant:** We're a bit unusual for an agency our size. I'm a full-time business-development person, and outbound is about 30% of our new revenue. Referrals and repeat work are still the majority, maybe 40% referrals and the rest partners: Shopify's partner directory and a couple of app companies who send us projects.

…
interview-A08-packaging-studio.md23 lines · Download
# Interview A08: packaging design studio, 5 staff
Participant: founder and creative director. Interviewer: our co-founder. 22 minutes, lightly edited.

**Interviewer:** How do new clients find you?

**Participant:** Almost entirely word of mouth. Food and drink brands talk to each other. Someone launches a range, their packaging does well on the shelf, and the next founder asks who did it. I'd say 95% referral. We've never done any outbound.

…
What a strong answer does

Names the load-bearing assumption as 'small agencies' growth is limited by a shortage of new leads': the interviews say it is limited by trust and capacity. 11 of 14 win most work through referrals and repeat clients; 7 of 9 who tried outbound quit within six months with no attributable deals, for reasons tied to trust (A05's 'content is a trust purchase', A01's meetings with people who'd never buy) and reputation risk (A02's restricted LinkedIn account, A11's spam listing); 3 turn work away (A08: 'I'd have to say no to nine of them'). Quotes the interviews accurately and attributes them correctly. Engages the counter-evidence honestly: outbound works for A04 and A07, but both are 15–20 staff with a dedicated BD owner, already use tools like Apollo and Clay, and credit the follow-up rather than the first email, so the segment is small, tooled and not what the product automates. Treats feast or famine as a real pain but notes agencies fix it through past clients and referrals (A05's reactivation emails, and her willingness to pay for help keeping in touch with past clients), pointing at a different product. Says what would prove the challenge wrong (for example, a meaningful share of 10–20-staff agencies with a BD owner and budget, or a concierge outbound test producing qualified meetings that turn into proposals), and designs a cheap test that measures an early signal, since a 6–10 week sales cycle means a six-week test cannot observe won deals.

Critical failures (cap the score)
  • Invents interview quotes or attributes a quote to the wrong agency
  • Treats the founders' feast-or-famine hypothesis as confirmed by the interviews
Case

v1.8 · synthetic · B2B, go-to-market, customer research

Usable with a quick edit10 of 11 checks passed

The memo is structurally complete and well‑reasoned; a quick edit to shift the test gate to an early signal (e.g., qualified proposals) would make it fully usable without rechecking evidence or redoing the reasoning.

Output

To: Founding team Subject: We risk automating a channel our customers cannot turn into revenue

Do not commit a year to this product yet. The strongest reason it fails is that finding prospects and sending personalised emails does not remove the constraint on small agencies’ sales: the human work of establishing trust and converting interest. Our promise assumes an agent can keep revenue flowing while the founder stays busy delivering. The interviews suggest we would generate more work for that same founder.

Eleven of 14 agencies get most new revenue from referrals and repeat clients, averaging 86% across those eleven. Of nine that tried outbound, seven stopped within six months without an attributable deal. This is more than unfamiliarity with the channel: most have bought or attempted it and abandoned it.

A05 shows the mechanism. Its lead-generation agency booked six or seven meetings a month, but prospects lacked budgets or wanted evidence of relevant experience. With a referred prospect, “the first call was about scope, not about whether we were any good”. Personalisation may earn attention; these interviews do not establish that it supplies the credibility needed to buy.

The successful case makes the constraint clearer. A04 generates 30% of new revenue through outbound, but has a dedicated BD lead who follows emails with site teardowns, videos and calls: “They don’t have the time to follow up, and the follow-up is the job.” Automating the first step leaves that job with the founder.

Some agencies do not need additional demand at all. A08 says, “Getting them isn’t. Handling them is,” and needs a senior designer. Three agencies regularly turn work away. Feast or famine is not a universal unmet need.

What cuts against this conclusion

Six agencies do report feast-or-famine swings. A04 and A07 prove outbound can work, although both have dedicated business-development capacity. A04 would pay $600 if we replaced research tools, saved time, integrated with HubSpot and required email approval. That supports a narrower research assistant proposition.

Past failures also do not prove our product would fail: targeting and execution may improve, and A05’s two-month trial barely covered the typical 6–10-week sales cycle. Fourteen interviews cannot establish market prevalence.

But “one extra client pays for it” skips the disputed step: winning that client. The $7,200 annual subscription is 40% of the $18,000 average deal’s revenue, before delivery costs and selling time.

What would prove us wrong, and the cheapest test

Run a paid concierge pilot before building. Recruit five agencies with spare delivery capacity, a stated growth goal and no dedicated BD person. Charge the planned $600 monthly price. Using existing tools, manually deliver prospect research and personalised drafts, with approval before sending. Founders must handle follow-up; we must not secretly supply the missing salesperson.

Run for 12 weeks to cover the stated sales cycle.1 Agree a founder-time ceiling upfront, record all selling time and track qualified opportunities through to attributable signed deals, contribution margin and renewal.

A proposed gate: at least three of five win business whose expected contribution covers subscription and selling costs, remain within the time ceiling and renew at full price. Meetings alone do not pass.

If only agencies with BD staff succeed, narrow the customer and promise. A05’s request for past-client reminders suggests another test, but “I’d pay for that tomorrow” is not payment. Do not turn one appealing interview into the next year-long commitment.

What a PM had to fix

  1. 1Test or gate too weakTighten the testQuick edit

    What we’d changeSigning and renewing inside 12 weeks is a bar that quietly favours the memo's own conclusion. Extend the window or add a leading indicator such as qualified proposals, and include the risk to agencies' LinkedIn channel.

Check by check

Got wrong · 1
  • A cheap test that can actually read outThe proposed test gate relies on signed deals (won business), but a 6–10 week sales cycle means the 12-week test likely cannot observe this late signal; the criteria require measuring an early signal like qualified meetings or proposals.
Got right · 10
  • Uses the supplied evidence correctlyAll factual claims about the current situation are drawn accurately from the brief and supplied context.
  • Addresses the actual decisionThe output gives a clear 'do not commit a year' answer early, framed for the founding team, and specifies what would change it.
  • Respects explicit constraintsThe memo is under 600 words, addresses the founding team, and includes the required one reason, evidence, counter-evidence, proof-wrong condition, and cheapest test.
  • Identifies material uncertaintyIt acknowledges that past failures do not prove the product will fail, the small sample limits prevalence estimates, and states what results would change the call.
  • Avoids unsupported claimsInterpretations are clearly labelled as hypotheses or suggestions, not presented as established facts.
  • Produces the required deliverableThe memo is a complete, usable document in the requested form, length, and reader, requiring only light editing.
  • Finds the load-bearing assumptionThe critique centers on the assumption that small agencies' growth is limited by a shortage of new leads, showing trust and capacity as the real constraints.
  • Uses the interviews faithfullyQuotes and paraphrases from A05, A04, A08 are accurate and correctly attributed, with no invented material.
  • Engages the counter-evidenceA04 and A07 are named, their dedicated BD capacity and follow-up are highlighted, and the implication of a narrow addressable segment is stated.
  • Tigers, not paper tigersIt dismisses the worry that a small sample invalidates the finding by noting that 14 interviews cannot establish prevalence, and surfaces the unspoken issue that outbound adds follow‑up work the founder cannot do.

Grades and run details

Decision model 95 · LLM judge 11 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly59%
  • passAddresses the actual decision99%
  • passRespects explicit constraints40%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims21%
  • passProduces the required deliverable86%
  • passFinds the load-bearing assumption99%
  • passUses the interviews faithfully81%
  • passEngages the counter-evidence81%
  • partialA cheap test that can actually read out42%
  • passTigers, not paper tigers98%
Run
Run
#1
Time to output
72 s
Submitted
25 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 87% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.095.82None
2GPT-6 AstrawithChatGPT95.595.82None
3GPT-6 LunawithAPI93.291.72None
4Sonnet 5.5withAPI93.275.02None
5Gemini 3.5 Flash-LitewithGemini81.875.021 capped
6Opus 5.5withClaude90.991.721 capped

About the task

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Separates the risks that could kill it from the ones that only look scary
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill · Staff level: a company bet with a long evidence pack

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.