Needs more than a quick edit3 of 11 checks passed
A PM would need to re-check evidence and repair the test before acting. The core argument is useful but not reliable as written.
Output
The One Reason This Fails
Agency services are high-trust purchases that cold automated outreach cannot convert without intensive, bespoke human follow-up—which time-poor founders cannot provide. Automating email generation solves top-of-funnel activity, but the actual point of failure is trust and follow-up capacity.
---
The Evidence
1. Zero conversion for founder-led shops: 7 of 9 agencies that tried outbound stopped within six months having closed zero deals. 2. Cold outreach delivers the wrong buyers: Services require credibility upfront. * A05 (11 staff): “Content is a trust purchase... The people on those calls had no idea who we were and, honestly, most of them didn’t have a budget.” * A01 (8 staff): “We got meetings with people who’d never buy from us.” * A10 (20 staff): “Cost us about £400 for every meeting, and the meetings didn’t close.” 3. The constraint is manual follow-up, not initial emails: * A04 (18 staff, dedicated BD): “It works because of the follow-up, not the first email... What wins deals is that I then send them a teardown of their site, a short video... The owner does sales on the side. They don't have the time to follow up, and the follow-up is the job.” 4. Asymmetric downside risk: Small agencies live on reputation. * A11: “Our domain ended up on a spam list.” * A02: “LinkedIn shut our founder’s account down. That’s our best referral network.” * A05: “I’d worry about our name going out on emails I haven’t read.”
---
Where the Evidence Cuts the Other Way
- Feast-and-famine is real: 6 of 14 experience painful dry spells. A05 panicked when three clients left simultaneously.
- Outbound works at scale with dedicated staff: A04 generates 30% of revenue via outbound; A07 makes it pay (“one in ten turns into a retainer”).
- Willingness to pay exists if positioned right: A04 would pay $600/month if it replaced Apollo/Clay and saved research hours. A05 would pay $200/month for an automated tool to reactivate past clients and referral networks.
---
What We Would Need to See to Be Proved Wrong
- A solo founder (with no dedicated BD rep) closing an outbound deal with typical agency pricing ($18k) using standard outbound messaging, without spending >5 hours/week on bespoke pitch work.
- Willingness of agency owners to allow an autonomous agent to send cold outbound without manual pre-approval of every single touchpoint.
---
The Cheapest Test to Prove It
A 3-week "Wizard of Oz" Concierge Test with 3 Agencies (Cost: <$500): * Partner with 3 agencies suffering feast-or-famine (e.g., A05, A09, A10) who will commit to running a free, 30-day outbound campaign. * Manually emulate the AI SDR: scrape prospect lists, generate personalized emails, and handle automated sequencing from a secondary domain. * Kill criteria: If after booking 10–15 meetings across these agencies, zero advance to contract negotiations—or founders drop out because they cannot keep up with follow-ups and custom scoping—kill the product.
(Pivot signal: A05 offered to pay immediately for an AI assistant that drafts warm re-engagement notes to past clients and referral partners. That solves the same famine problem without the friction of cold outbound.)
Claims the judge couldn’t find in the brief
- 7 of 9 agencies that tried outbound stopped within six months having closed zero deals.
- 6 of 14 agencies experience painful dry spells.
- A04 generates 30% of revenue via outbound.
- A05, A09, and A10 are agencies suffering feast-or-famine.
Grades and run details
Decision model 95 · LLM judge 3 of 12 checks
Decision model checks
- passUses the supplied evidence correctly44%
- passAddresses the actual decision96%
- passRespects explicit constraints36%
- passIdentifies material uncertainty97%
- partialAvoids unsupported claims23%
- passProduces the required deliverable71%
- passFinds the load-bearing assumption26%
- passUses the interviews faithfully74%
- passEngages the counter-evidence79%
- passA cheap test that can actually read out94%
- passTigers, not paper tigers94%
Run
- Run
- #1
- API response time
- 22 s
- Submitted
- 1 Oct 2026