Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.2 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 64% were usable with at most a quick edit.

Reliably right

  1. Addresses the actual decision100% pass
    It commits early to not committing three squads and authorizing a six-week test, and specifies what result would change that.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  2. Identifies material uncertainty100% pass
    It names the unknowns (representativeness of payment data, large-account volume access, actual mix) and resolves them with a bounded test.
    GPT-6.1 Sol · API · The CEO's embedded-payments bet
  3. Uses the interviews faithfully100% pass
    All quotes are accurate and correctly attributed to A05, A08, A04, A07, and A12.
    GPT-6 Luna · API · An AI SDR for small agencies

Where it slips

  1. A cheap test that can actually read out64% pass
    The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.
    GPT-6.1 Sol · API · An AI SDR for small agencies
  2. Re-estimates the revenue correctly71% pass
    It shows working but lands on a central $2.3M and $1M–$6M range, not roughly the $5–10M range required by the grading rubric, and its $9.1M pilot upper bound is not used as the main re-estimate.
    Sonnet 5.5 · API · The CEO's embedded-payments bet
  3. Avoids unsupported claims71% pass
    The memo claims the pilot achieved ~0.38% blended take as a fact, which is not in the evidence and is not derived from it arithmetically; it also treats the pro-rata $1.43B locked value as a hard constraint without flagging the assumption.
    Opus 5.5 · Claude · The CEO's embedded-payments bet

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

We plan to sell an AI sales-development rep, an agent that finds prospects and sends personalised outbound email, to agencies with 5–20 staff. Before we commit a year to it, find the strongest reason this fails. Use the interview material below, including the three transcripts. Write it up for the founding team: the one reason, the evidence for it (quote the interviews), where the evidence cuts the other way, what we would need to see to be proved wrong, and the cheapest test that would show it. Keep it under 600 words.

What the model was given7 items: Scenario, Founders' hypothesis (from our planning doc), Interview summary (14 agencies, 5–20 staff, last six weeks), Interview log, interview-A05-content-agency.md, interview-A04-shopify-agency.md, interview-A08-packaging-studio.md
ScenarioWe are a three-person founding team with £800k of pre-seed funding and about 20 months of runway. Planned price: $600 a month per agency. Two competitors have each raised over $10m in the last year, selling AI SDRs mostly to software companies with 50+ staff.
Founders' hypothesis (from our planning doc)“Small agencies live feast or famine. When a big client leaves they have no pipeline, because the founder is too busy delivering to sell. An AI SDR keeps the pipeline full in the background, so the famine never comes. At $600 a month, one extra client a year pays for it many times over.”
Interview summary (14 agencies, 5–20 staff, last six weeks)11 of 14 get most of their new revenue from referrals and repeat clients (86% on average across those 11). 9 had tried outbound in the last two years; 7 of those stopped within six months, and none of the 7 closed a deal they could attribute to it. The 2 who kept going (A04, A07) have 15–20 staff and a dedicated person for business development. 3 (A08, A12, A14) said they regularly turn work away. 6 described feast-or-famine swings in new business. Average deal size across the 14: $18k; typical sales cycle 6–10 weeks.
Interview logAgency (staff) · services · share of new revenue from referrals and repeat clients · outbound tried? · outcome · representative quote A01 (8) · brand and web studio · 90% · yes, cold-email agency · stopped after 3 months · "We got meetings with people who'd never buy from us." A02 (14) · performance marketing · 85% · yes, LinkedIn automation · stopped after 4 months, account restricted · "LinkedIn shut our founder's account down. That's our best referral network." A03 (6) · PR boutique · 95% · no · — · "Every client we have came from someone vouching for us." A04 (18) · Shopify development · 40% (plus 30% partners, 30% outbound) · yes, in-house BD lead · kept, 30% of new revenue · "It works because of the follow-up, not the first email." A05 (11) · content · 85% · yes, lead-gen agency · stopped after 2 months, no deals · "Content is a trust purchase." A06 (7) · video production · 85% · no · — · "Our clients find us through the videos. Someone shares one and we get a call." A07 (16) · SEO · 45% · yes, outsourced lead generation · kept · "It pays for itself, just. Most of the meetings are a waste of time, but one in ten turns into a retainer." A08 (5) · packaging design · 95% · no · — · "If you sent me ten more leads in October I'd have to say no to nine of them." A09 (12) · paid social · 75% · yes, cold email · stopped after 5 months · "The replies were mostly people asking us to take them off the list." A10 (20) · B2B marketing · 70% · yes, contract SDR · stopped after 6 months · "It cost us about £400 for every meeting, and the meetings didn't close." A11 (9) · WordPress web agency · 90% · yes, cold email · stopped after 3 months · "Our domain ended up on a spam list. Took weeks to fix." A12 (10) · branding · 85% · no · — · "We're booked out till March. I don't need more leads, I need another designer." A13 (15) · HubSpot partner · 48% (plus 40% from HubSpot's partner directory) · yes, cold email · stopped after 4 months · "HubSpot sends us more than we can handle in a good quarter." A14 (6) · UX research · 90% · no · — · "We're small on purpose. We say no to about a third of enquiries."
interview-A05-content-agency.md31 lines · Download
# Interview A05: content agency, 11 staff
Participant: founder and managing director. Interviewer: our co-founder. 34 minutes, lightly edited.

**Interviewer:** Where did your last five clients come from?

**Participant:** Let me think. Two were old clients coming back: one had changed jobs and brought us into her new company. Two were introductions, one from a web agency we partner with and one from a client's CFO who'd seen our work. The fifth found us through a talk I gave at a SaaS marketing meetup. So none of them from anything you'd call outbound.

…
interview-A04-shopify-agency.md23 lines · Download
# Interview A04: Shopify development agency, 18 staff
Participant: head of business development. Interviewer: our co-founder. 29 minutes, lightly edited.

**Interviewer:** Tell me how new business works for you.

**Participant:** We're a bit unusual for an agency our size. I'm a full-time business-development person, and outbound is about 30% of our new revenue. Referrals and repeat work are still the majority, maybe 40% referrals and the rest partners: Shopify's partner directory and a couple of app companies who send us projects.

…
interview-A08-packaging-studio.md23 lines · Download
# Interview A08: packaging design studio, 5 staff
Participant: founder and creative director. Interviewer: our co-founder. 22 minutes, lightly edited.

**Interviewer:** How do new clients find you?

**Participant:** Almost entirely word of mouth. Food and drink brands talk to each other. Someone launches a range, their packaging does well on the shelf, and the next founder asks who did it. I'd say 95% referral. We've never done any outbound.

…
What a strong answer doesThe answer key the graders mark against

Names the load-bearing assumption as 'small agencies' growth is limited by a shortage of new leads': the interviews say it is limited by trust and capacity. 11 of 14 win most work through referrals and repeat clients; 7 of 9 who tried outbound quit within six months with no attributable deals, for reasons tied to trust (A05's 'content is a trust purchase', A01's meetings with people who'd never buy) and reputation risk (A02's restricted LinkedIn account, A11's spam listing); 3 turn work away (A08: 'I'd have to say no to nine of them'). Quotes the interviews accurately and attributes them correctly. Engages the counter-evidence honestly: outbound works for A04 and A07, but both are 15–20 staff with a dedicated BD owner, already use tools like Apollo and Clay, and credit the follow-up rather than the first email, so the segment is small, tooled and not what the product automates. Treats feast or famine as a real pain but notes agencies fix it through past clients and referrals (A05's reactivation emails, and her willingness to pay for help keeping in touch with past clients), pointing at a different product. Says what would prove the challenge wrong (for example, a meaningful share of 10–20-staff agencies with a BD owner and budget, or a concierge outbound test producing qualified meetings that turn into proposals), and designs a cheap test that measures an early signal, since a 6–10 week sales cycle means a six-week test cannot observe won deals.

Critical failures (cap the score at 40)
  • Invents interview quotes or attributes a quote to the wrong agency
  • Treats the founders' feast-or-famine hypothesis as confirmed by the interviews
Case

Staff-level · v1.10 · synthetic · B2B, go-to-market, customer research

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 11 checks passed

The core answer, evidence, and risk triage are usable as-is; the only needed edit is shifting the pilot gate to an earlier observable signal.

Output

Strongest reason this fails: you automate prospecting, but prospecting is not the binding constraint

The proposed AI SDR supplies cold introductions. Most agencies interviewed win through transferred trust; converting cold prospects requires substantial human follow-up that busy founders cannot supply. The product therefore risks delivering meetings rather than revenue—and being cancelled before it earns its keep.

This is a mismatch with the promised outcome, not proof that AI cannot write good emails. Cheaper, more personalised outreach does not automatically create trust, budget or founder time.

Evidence

  • Existing behaviour is a warning: 11 of 14 agencies get most new revenue from referrals and repeat clients—86% on average within that group. Nine tried outbound; seven stopped within six months, with no attributable deals.
  • A05 confirms the pain but rejects the proposed solution. Its founder described “feast or famine,” yet previous outbound generated six or seven meetings monthly without a sale: “Most of them didn’t have a budget for content.” Introductions worked because “they’d already been told we were good.” After losing three clients, personal notes to former clients brought two back within a month. On pricing: “Six hundred for cold outbound, no.”
  • Even the strongest success case exposes the missing work. A04 gets 30% of new revenue from outbound, but employs a full-time BD lead: “What wins deals is that I then send them a teardown of their site, a short video, and I’m on the phone with them within a week.” For typical founders, “They don’t have the time to follow up, and the follow-up is the job.”
  • Some agencies do not need more demand. A08: “Getting them isn’t [a problem]. Handling them is.” Its answer to the product was: “It would make my problem worse.” Three agencies regularly turn work away.

“One extra client pays for it” only holds if the tool causes an incremental, profitable win—not merely a meeting—and the agency can serve it.

Where the evidence cuts the other way

Six agencies report feast-or-famine swings. A04 would pay $600 if the tool replaced research tools and saved three hours daily, with HubSpot integration and email approval. A07 says outbound “pays for itself, just.”

These suggest a viable BD-assistance niche, particularly among larger agencies with a dedicated salesperson. They do not establish demand for autonomous pipeline generation across 5–20-person agencies.

The interviews are also a small sample, and failed historical outbound does not prove better targeting cannot work. A05’s two-month trial barely covered a typical sales cycle.

What would prove us wrong

Several agencies without dedicated BD staff pay $600, convert agent-sourced cold prospects into profitable clients with little founder follow-up, and renew. Success only among A04-like firms would support a narrower product, not the current thesis.

Cheapest decisive test

Run a 12–16-week paid concierge pilot with six agencies that genuinely need clients, primarily without BD staff. Charge $600 monthly; use existing prospecting and email tools rather than build software. Obtain sending approval and protect their primary domains.

Track qualified opportunities, founder follow-up hours, attributable signed deals, delivery capacity and paid renewals—not replies or booked meetings. Pre-register a continuation gate: at least three agencies close an attributable profitable deal and renew, within a founder-time limit they set beforehand.

Failure to recruit paying pilots is useful evidence itself. Do this before committing a year.

Grades and run details

Decision model 100 · LLM judge 11 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly79%
  • passAddresses the actual decision97%
  • passRespects explicit constraints41%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims46%
  • passProduces the required deliverable88%
  • passFinds the load-bearing assumption100%
  • passUses the interviews faithfully84%
  • passEngages the counter-evidence89%
  • passA cheap test that can actually read out68%
  • passTigers, not paper tigers97%
Run
Run
#1
API response time
17 s
Submitted
29 Sept 2026
Critical failure: Treats the founders' feast-or-famine hypothesis as confirmed by the interviews. Score capped at 40.

Usable with a quick edit11 of 11 checks passed

The memo is well-structured, evidence-backed, and actionable, requiring at most light editing for polish.

Output

Why the AI SDR for small agencies fails

Sep 25, 2026 · @Martin Slaney

The reason

Small agencies win work on trust, and cold outbound only closes when someone does the follow-up. An AI SDR automates the first email, the cheap part. It hands the follow-up, which is the actual job, to a founder our own hypothesis says is too busy delivering to sell. We would be adding meetings to the one resource that is already the constraint.

Evidence

  • Referrals dominate. 11 of 14 get most new revenue (86% on average) from referrals and repeat clients. A03: "Every client we have came from someone vouching for us."
  • Outbound was tried and dropped. 9 tried it; 7 stopped within six months and none of the 7 closed an attributable deal. They used five different channels (lead-gen agency, cold-email agency, contract SDR, LinkedIn automation, DIY email), so this isn't one bad vendor.
  • Meetings weren't the problem; closing was. A05 got six or seven meetings a month and closed nothing: "Content is a trust purchase." A01: "We got meetings with people who'd never buy from us." A10: "It cost us about £400 for every meeting, and the meetings didn't close."
  • Where it works, a human does the work. Both survivors (A04, A07) have 15–20 staff and a dedicated BD person. A04: "It works because of the follow-up, not the first email" and "most agencies our size don't have someone like me."
  • The downside hits their best channel. A02 lost the founder's LinkedIn, "our best referral network"; A11's domain was spam-listed. A04 and A05 both refuse to let anything send unread.
  • Some don't need leads at all. A08, A12 and A14 turn work away. A12: "I don't need more leads, I need another designer."

Where it cuts the other way

  • The famine is real.1 6 of 14 describe feast-or-famine swings; A05 "spent April panicking" after three clients left at once.
  • The maths are forgiving. At an $18k average deal, one closed client a year covers $7.2k of subscription.
  • One buyer exists. A04 would pay $600 today if it replaced Apollo, Clay and three hours a day of research, but only with HubSpot sync and approval of every email.
  • Past failures were generic. A05's buyers wanted proof of work in their niche; better enrichment might fix that. Untested.
  • Small sample. 14 interviews, one interviewer.
  • Warm outreach worked. A05's emails to past clients won two back in a month, and she'd pay "a couple of hundred a month" for a tool that did it. That supports the thesis against cold outbound, but points at a different product.

What would prove us wrong

Agencies without a dedicated BD person closing deals from AI-sent cold email, without damage to their reputation. Concretely, in a 12-week pilot with 5 such agencies:

  • at least 3 reach a scoped proposal and at least 1 closes;2
  • no domain, inbox or LinkedIn penalties;
  • owners spend under 2 hours a week on follow-up;
  • they keep paying when the pilot ends.

Cheapest test

Pre-sell a hand-run pilot before building anything. Offer the 9 agencies with no BD person that aren't turning work away (A01–A03, A05, A06, A09–A11, A13) a 12-week pilot at $600 a month, paid upfront. We run it by hand: Apollo and Clay for prospects, LLM-drafted emails, owner approves every send, separate sending domain.

  1. Gate 1, one week: fewer than 3 pay, stop. There is no demand at this price.
  2. Gate 2, 12 weeks (covers the 6–10 week sales cycle): score against the thresholds above and log owner follow-up hours.

In parallel, offer A05's version (past-client and referrer re-engagement, about $200 a month) as a second arm. If that sells and cold doesn't, we have our answer.

Cost: roughly £1–2k in tools and a quarter of one founder's time3, against a year and most of the £800k.

What a PM had to fix

  1. 1Hypothesis stated as factReframe it as a hypothesisTargeted repair

    Source: Founders' hypothesis (from our planning doc)Small agencies live feast or famine.

    What we’d changeSay 'Famine is real for some: 6 of 14' rather than confirming the founders' hypothesis for the whole segment.

  2. 2Test or gate too weakTighten the testQuick edit

    What we’d changeRaise the bar for proving the memo wrong: one close in five agencies doesn't show the product works.

  3. 3OtherFix itQuick edit

    What we’d changeThe cost looks light for running a hand-run pilot across nine agencies. Build it up line by line.

Grades and run details

Decision model 95 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly42%
  • passAddresses the actual decision95%
  • partialRespects explicit constraints40%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims16%
  • passProduces the required deliverable79%
  • passFinds the load-bearing assumption92%
  • passUses the interviews faithfully79%
  • passEngages the counter-evidence97%
  • passA cheap test that can actually read out92%
  • passTigers, not paper tigers98%
Run
Run
#1
Time to output
60 s
Submitted
25 Sept 2026

Usable with a quick edit11 of 11 checks passed

The memo is well-structured, evidence-based, and actionable; a product manager could use it with only light editing.

Output

Founding-team memo: the strongest reason this fails

Reason: The product may solve the wrong bottleneck. Many small agencies are not short of names to contact: they win through trust, referrals and repeat work, and some already lack capacity to take on more. Cold outbound can create low-quality meetings—or reputational risk—without solving the follow-up needed to convert them. At $600/month, an AI SDR could therefore add work rather than prevent famine.

Evidence for it: In 11 of 14 agencies, referrals and repeat clients account for most new revenue (86% on average). Of the nine that tried outbound, seven stopped within six months; none attributed a deal to it. A05’s experience is telling: a lead-gen firm booked “maybe six or seven a month,” but “most of them didn’t have a budget for content.” The founder said, “Content is a trust purchase,” and would pay for prompts and drafts to contact past clients—not “$600 for cold outbound.”

Capacity is also a real constraint: A08 said, “If you sent me ten more leads in October I’d have to say no to nine of them”; A12 needs another designer, not more leads. And even outbound that works appears to require substantial human selling. A04 said, “It works because of the follow-up, not the first email,” then described sending a site teardown, video and making a call. A07 said most meetings were “a waste of time.” The risk is not merely that the AI writes mediocre emails; it is that the agency cannot or does not convert what it sends.

Where the evidence cuts the other way: Six agencies reported feast-or-famine swings, and A05 described panicking after three clients ended together. A04 gets 30% of new revenue from outbound and would pay if the tool saved research time; A07 says its outsourced lead generation “pays for itself, just.” With an $18k average deal, one win could easily cover the subscription. These examples show outbound can work—but the two continuing users are larger agencies with dedicated BD, unlike the founder-led customer we plan to target. A05’s warm reactivation success also suggests the underlying need may be real, but the best channel may not be cold prospecting.

What would prove us wrong: Typical 5–20-person agencies without dedicated BD will pay $600/month for cold outbound, follow up promptly, and generate attributable profitable work—not just meetings—without harming their reputation or exceeding delivery capacity.

Cheapest test: Don’t build the agent. Recruit five founder-led agencies with capacity and a genuine pipeline gap, charge the planned $600/month, and manually run a 12-week, human-approved outbound pilot. Track qualified meetings, founder follow-up time, attributable wins and renewals. Given 6–10-week sales cycles, a strong signal would be at least two attributable wins and three agencies renewing at full price. If interest stops at free trials, meetings fail to convert, or founders cannot follow up, stop or test a warm-client reactivation product instead.

Grades and run details

Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
  • passUses the supplied evidence correctly65%
  • passAddresses the actual decision93%
  • passRespects explicit constraints60%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims49%
  • passProduces the required deliverable94%
  • passFinds the load-bearing assumption99%
  • passUses the interviews faithfully75%
  • passEngages the counter-evidence96%
  • passA cheap test that can actually read out73%
  • passTigers, not paper tigers96%
Run
Run
#1
API response time
27 s
Submitted
29 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 1

A cheap test that can actually read outMixedRightRight
GPT-6.1 Sol · API

The pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.

Opus 5.5 · Claude

The test is a hand-run pilot that measures proposals within 12 weeks (covering the sales cycle) and uses simple gates; it is cheap and provides an early signal before full build.

GPT-6 Luna · API

The test is a cheap, manual pilot with a concrete threshold, tracking early signals (qualified meetings, follow-up time) and acknowledging the 6–10 week sales cycle.

All got right 10

Uses the supplied evidence correctlyRightRightRight
GPT-6.1 Sol · API

The facts, figures, and quotes come from the supplied summary and transcripts without inventing current-situation claims.

Opus 5.5 · Claude

Every factual claim about the current situation is drawn correctly from the brief or supplied context, with accurate quotes and attributions.

GPT-6 Luna · API

All factual claims about the current situation are directly supported by the supplied context, with no inventions.

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

It commits early to one answer—prospecting is not the binding constraint—and states what evidence would change it.

Opus 5.5 · Claude

The output commits to a single reason for failure early on, addresses the founding team, and specifies the evidence and conditions that would change the call.

GPT-6 Luna · API

The output commits to a clear answer (the product solves the wrong bottleneck) and states what would prove it wrong.

Respects explicit constraintsRightRightRight
GPT-6.1 Sol · API

It is written for the founding team, under 600 words, and includes the requested one reason, evidence, counter-evidence, falsification condition, and test.

Opus 5.5 · Claude

The output is a memo under 600 words, covers all requested sections, and the proposals (test, thresholds) enforce the constraints in practice.

GPT-6 Luna · API

The memo is written for the founding team, includes all requested sections, and stays under 600 words.

Identifies material uncertaintyRightRightRight
GPT-6.1 Sol · API

It names the small sample, historical outbound not proving better targeting impossible, and the condition that would prove the thesis wrong.

Opus 5.5 · Claude

It names the small sample size, whether better enrichment could fix niche trust, and that the test result would resolve the uncertainty.

GPT-6 Luna · API

It names the key unknown (whether typical agencies without dedicated BD will pay and convert) and says how a test would resolve it.

Avoids unsupported claimsRightRightRight
GPT-6.1 Sol · API

Interpretations and forecasts are framed as risks or suggestions rather than established fact.

Opus 5.5 · Claude

Interpretations such as 'small agencies win work on trust' are clearly positioned as the reason and backed with evidence; no causal claims are presented as fact without support.

GPT-6 Luna · API

Interpretations and risks are clearly framed as hypotheses, not as established facts.

Produces the required deliverableRightRightRight
GPT-6.1 Sol · API

The requested memo is complete, usable, and within the requested length.

Opus 5.5 · Claude

The memo includes the reason, quoted evidence, counter-evidence, a falsification condition, a cheap test, and stays under 600 words; a founding team could act on it.

GPT-6 Luna · API

The deliverable is a complete, actionable memo for the founding team, within the word limit.

Finds the load-bearing assumptionRightRightRight
GPT-6.1 Sol · API

It makes lead-constrained growth the central challenge and supports it with referral, abandoned-outbound, and turning-work-away evidence.

Opus 5.5 · Claude

The critique directly challenges the assumption that these agencies are lead-constrained, arguing that they lack the capacity to convert leads and that outbound fails without follow-up.

GPT-6 Luna · API

The memo targets the load-bearing assumption that growth is limited by a shortage of new leads, using referral, abandoned-outbound and capacity evidence.

Uses the interviews faithfullyRightRightRight
GPT-6.1 Sol · API

Quotes are accurate and correctly attributed to A04, A05, A07, and A08.

Opus 5.5 · Claude

All quoted material is taken from the transcripts or log, attributed to the correct agency, and none is invented.

GPT-6 Luna · API

All quotes are accurate and correctly attributed to A05, A08, A04, A07, and A12.

Engages the counter-evidenceRightRightRight
GPT-6.1 Sol · API

It names A04 and A07, identifies the larger dedicated-BD agency pattern and follow-up dependence, and concludes the viable niche is narrower than the thesis.

Opus 5.5 · Claude

A04 and A07 are named, their shared characteristics (size, dedicated BD, follow-up effort) are identified, and the implication that the addressable segment is very small is drawn.

GPT-6 Luna · API

It names A04 and A07, notes their larger size and dedicated BD, and explains that follow-up, not the first email, drives their success.

Tigers, not paper tigersRightRightRight
GPT-6.1 Sol · API

It dismisses email quality and small-sample objections as less decisive while surfacing the unspoken founder-follow-up and delivery-capacity constraints.

Opus 5.5 · Claude

It dismisses the worry that past failures were just due to poor tooling as untested, and surfaces the unspoken issue that a warm-outreach re-engagement tool might be a better product.

GPT-6 Luna · API

It separates the real bottleneck (trust, capacity, follow-up) from manageable risks, dismisses the idea that outbound never works, and surfaces the unspoken issue that agencies lack conversion capacity.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 88% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI100.095.82None
2GPT-6 AstrawithChatGPT93.295.82None
3GPT-6 LunawithAPI97.791.72None
4Sonnet 5.5withAPI90.975.02None
5Gemini 3.5 Flash-LitewithGemini88.675.021 capped
6Opus 5.5withClaude93.291.721 capped
7Gemini 3.8 FlashwithAPI93.237.521 capped

About the task

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Separates the risks that could kill it from the ones that only look scary
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review

Variants

Vanilla prompt (core) · With Roast Me skill · Staff level: a company bet with a long evidence pack

This task measures the whole setup. Tools, instructions and skills in the harness do real work here, so read the harness as carefully as the model name.