Usable with a quick edit11 of 11 checks passed
The memo is well-structured, evidence-backed, and actionable, requiring at most light editing for polish.
Output
Why the AI SDR for small agencies fails
Sep 25, 2026 · @Martin Slaney
The reason
Small agencies win work on trust, and cold outbound only closes when someone does the follow-up. An AI SDR automates the first email, the cheap part. It hands the follow-up, which is the actual job, to a founder our own hypothesis says is too busy delivering to sell. We would be adding meetings to the one resource that is already the constraint.
Evidence
- Referrals dominate. 11 of 14 get most new revenue (86% on average) from referrals and repeat clients. A03: "Every client we have came from someone vouching for us."
- Outbound was tried and dropped. 9 tried it; 7 stopped within six months and none of the 7 closed an attributable deal. They used five different channels (lead-gen agency, cold-email agency, contract SDR, LinkedIn automation, DIY email), so this isn't one bad vendor.
- Meetings weren't the problem; closing was. A05 got six or seven meetings a month and closed nothing: "Content is a trust purchase." A01: "We got meetings with people who'd never buy from us." A10: "It cost us about £400 for every meeting, and the meetings didn't close."
- Where it works, a human does the work. Both survivors (A04, A07) have 15–20 staff and a dedicated BD person. A04: "It works because of the follow-up, not the first email" and "most agencies our size don't have someone like me."
- The downside hits their best channel. A02 lost the founder's LinkedIn, "our best referral network"; A11's domain was spam-listed. A04 and A05 both refuse to let anything send unread.
- Some don't need leads at all. A08, A12 and A14 turn work away. A12: "I don't need more leads, I need another designer."
Where it cuts the other way
- The famine is real.1 6 of 14 describe feast-or-famine swings; A05 "spent April panicking" after three clients left at once.
- The maths are forgiving. At an $18k average deal, one closed client a year covers $7.2k of subscription.
- One buyer exists. A04 would pay $600 today if it replaced Apollo, Clay and three hours a day of research, but only with HubSpot sync and approval of every email.
- Past failures were generic. A05's buyers wanted proof of work in their niche; better enrichment might fix that. Untested.
- Small sample. 14 interviews, one interviewer.
- Warm outreach worked. A05's emails to past clients won two back in a month, and she'd pay "a couple of hundred a month" for a tool that did it. That supports the thesis against cold outbound, but points at a different product.
What would prove us wrong
Agencies without a dedicated BD person closing deals from AI-sent cold email, without damage to their reputation. Concretely, in a 12-week pilot with 5 such agencies:
- at least 3 reach a scoped proposal and at least 1 closes;2
- no domain, inbox or LinkedIn penalties;
- owners spend under 2 hours a week on follow-up;
- they keep paying when the pilot ends.
Cheapest test
Pre-sell a hand-run pilot before building anything. Offer the 9 agencies with no BD person that aren't turning work away (A01–A03, A05, A06, A09–A11, A13) a 12-week pilot at $600 a month, paid upfront. We run it by hand: Apollo and Clay for prospects, LLM-drafted emails, owner approves every send, separate sending domain.
- Gate 1, one week: fewer than 3 pay, stop. There is no demand at this price.
- Gate 2, 12 weeks (covers the 6–10 week sales cycle): score against the thresholds above and log owner follow-up hours.
In parallel, offer A05's version (past-client and referrer re-engagement, about $200 a month) as a second arm. If that sells and cold doesn't, we have our answer.
Cost: roughly £1–2k in tools and a quarter of one founder's time3, against a year and most of the £800k.
What a PM had to fix
1Hypothesis stated as factReframe it as a hypothesisTargeted repair
Source: Founders' hypothesis (from our planning doc)
Small agencies live feast or famine.
What we’d changeSay 'Famine is real for some: 6 of 14' rather than confirming the founders' hypothesis for the whole segment.
2Test or gate too weakTighten the testQuick edit
What we’d changeRaise the bar for proving the memo wrong: one close in five agencies doesn't show the product works.
3OtherFix itQuick edit
What we’d changeThe cost looks light for running a hand-run pilot across nine agencies. Build it up line by line.
Check by check
Got right · 11
- Uses the supplied evidence correctlyEvery factual claim about the current situation is drawn correctly from the brief or supplied context, with accurate quotes and attributions.
- Addresses the actual decisionThe output commits to a single reason for failure early on, addresses the founding team, and specifies the evidence and conditions that would change the call.
- Respects explicit constraintsThe output is a memo under 600 words, covers all requested sections, and the proposals (test, thresholds) enforce the constraints in practice.
- Identifies material uncertaintyIt names the small sample size, whether better enrichment could fix niche trust, and that the test result would resolve the uncertainty.
- Avoids unsupported claimsInterpretations such as 'small agencies win work on trust' are clearly positioned as the reason and backed with evidence; no causal claims are presented as fact without support.
- Produces the required deliverableThe memo includes the reason, quoted evidence, counter-evidence, a falsification condition, a cheap test, and stays under 600 words; a founding team could act on it.
- Finds the load-bearing assumptionThe critique directly challenges the assumption that these agencies are lead-constrained, arguing that they lack the capacity to convert leads and that outbound fails without follow-up.
- Uses the interviews faithfullyAll quoted material is taken from the transcripts or log, attributed to the correct agency, and none is invented.
- Engages the counter-evidenceA04 and A07 are named, their shared characteristics (size, dedicated BD, follow-up effort) are identified, and the implication that the addressable segment is very small is drawn.
- A cheap test that can actually read outThe test is a hand-run pilot that measures proposals within 12 weeks (covering the sales cycle) and uses simple gates; it is cheap and provides an early signal before full build.
- Tigers, not paper tigersIt dismisses the worry that past failures were just due to poor tooling as untested, and surfaces the unspoken issue that a warm-outreach re-engagement tool might be a better product.
Grades and run details
Decision model 91 · LLM judge 12 of 12 checks
Decision model checks
- passUses the supplied evidence correctly36%
- passAddresses the actual decision94%
- partialRespects explicit constraints22%
- passIdentifies material uncertainty100%
- partialAvoids unsupported claims24%
- passProduces the required deliverable79%
- passFinds the load-bearing assumption89%
- passUses the interviews faithfully82%
- passEngages the counter-evidence96%
- passA cheap test that can actually read out89%
- passTigers, not paper tigers98%
Run
- Run
- #1
- Time to output
- 60 s
- Submitted
- 25 Sept 2026