Usable with a quick edit10 of 11 checks passed
The core answer, evidence, and risk triage are usable as-is; the only needed edit is shifting the pilot gate to an earlier observable signal.
Output
Strongest reason this fails: you automate prospecting, but prospecting is not the binding constraint
The proposed AI SDR supplies cold introductions. Most agencies interviewed win through transferred trust; converting cold prospects requires substantial human follow-up that busy founders cannot supply. The product therefore risks delivering meetings rather than revenue—and being cancelled before it earns its keep.
This is a mismatch with the promised outcome, not proof that AI cannot write good emails. Cheaper, more personalised outreach does not automatically create trust, budget or founder time.
Evidence
- Existing behaviour is a warning: 11 of 14 agencies get most new revenue from referrals and repeat clients—86% on average within that group. Nine tried outbound; seven stopped within six months, with no attributable deals.
- A05 confirms the pain but rejects the proposed solution. Its founder described “feast or famine,” yet previous outbound generated six or seven meetings monthly without a sale: “Most of them didn’t have a budget for content.” Introductions worked because “they’d already been told we were good.” After losing three clients, personal notes to former clients brought two back within a month. On pricing: “Six hundred for cold outbound, no.”
- Even the strongest success case exposes the missing work. A04 gets 30% of new revenue from outbound, but employs a full-time BD lead: “What wins deals is that I then send them a teardown of their site, a short video, and I’m on the phone with them within a week.” For typical founders, “They don’t have the time to follow up, and the follow-up is the job.”
- Some agencies do not need more demand. A08: “Getting them isn’t [a problem]. Handling them is.” Its answer to the product was: “It would make my problem worse.” Three agencies regularly turn work away.
“One extra client pays for it” only holds if the tool causes an incremental, profitable win—not merely a meeting—and the agency can serve it.
Where the evidence cuts the other way
Six agencies report feast-or-famine swings. A04 would pay $600 if the tool replaced research tools and saved three hours daily, with HubSpot integration and email approval. A07 says outbound “pays for itself, just.”
These suggest a viable BD-assistance niche, particularly among larger agencies with a dedicated salesperson. They do not establish demand for autonomous pipeline generation across 5–20-person agencies.
The interviews are also a small sample, and failed historical outbound does not prove better targeting cannot work. A05’s two-month trial barely covered a typical sales cycle.
What would prove us wrong
Several agencies without dedicated BD staff pay $600, convert agent-sourced cold prospects into profitable clients with little founder follow-up, and renew. Success only among A04-like firms would support a narrower product, not the current thesis.
Cheapest decisive test
Run a 12–16-week paid concierge pilot with six agencies that genuinely need clients, primarily without BD staff. Charge $600 monthly; use existing prospecting and email tools rather than build software. Obtain sending approval and protect their primary domains.
Track qualified opportunities, founder follow-up hours, attributable signed deals, delivery capacity and paid renewals—not replies or booked meetings. Pre-register a continuation gate: at least three agencies close an attributable profitable deal and renew, within a founder-time limit they set beforehand.
Failure to recruit paying pilots is useful evidence itself. Do this before committing a year.
Check by check
Mixed · 1
- A cheap test that can actually read outThe pre-registered gate requires signed deals and renewal rather than an early signal such as qualified meetings or proposals, so it risks being too late for a cheap read-out.The two graders disagreed on this one.
Got right · 10
- Uses the supplied evidence correctlyThe facts, figures, and quotes come from the supplied summary and transcripts without inventing current-situation claims.
- Addresses the actual decisionIt commits early to one answer—prospecting is not the binding constraint—and states what evidence would change it.
- Respects explicit constraintsIt is written for the founding team, under 600 words, and includes the requested one reason, evidence, counter-evidence, falsification condition, and test.
- Identifies material uncertaintyIt names the small sample, historical outbound not proving better targeting impossible, and the condition that would prove the thesis wrong.
- Avoids unsupported claimsInterpretations and forecasts are framed as risks or suggestions rather than established fact.
- Produces the required deliverableThe requested memo is complete, usable, and within the requested length.
- Finds the load-bearing assumptionIt makes lead-constrained growth the central challenge and supports it with referral, abandoned-outbound, and turning-work-away evidence.
- Uses the interviews faithfullyQuotes are accurate and correctly attributed to A04, A05, A07, and A08.
- Engages the counter-evidenceIt names A04 and A07, identifies the larger dedicated-BD agency pattern and follow-up dependence, and concludes the viable niche is narrower than the thesis.
- Tigers, not paper tigersIt dismisses email quality and small-sample objections as less decisive while surfacing the unspoken founder-follow-up and delivery-capacity constraints.
Grades and run details
Decision model 100 · LLM judge 11 of 12 checks
Decision model checks
- passUses the supplied evidence correctly68%
- passAddresses the actual decision94%
- passRespects explicit constraints45%
- passIdentifies material uncertainty99%
- passAvoids unsupported claims46%
- passProduces the required deliverable87%
- passFinds the load-bearing assumption99%
- passUses the interviews faithfully84%
- passEngages the counter-evidence89%
- passA cheap test that can actually read out64%
- passTigers, not paper tigers97%
Run
- Run
- #1
- API response time
- 17 s
- Submitted
- 29 Sept 2026