Needs more than a quick edit6 of 11 checks passed
The output includes an unsupported revenue claim about A07 and proposes a test for a different product, so a PM would need to recheck facts and redesign the test.
Output
Why an AI SDR for Small Agencies Will Fail: Pre-Mortem Analysis
For: Founding Team (£800k runway, 20 months remaining) Product: AI Sales-Development Rep ($600/month) targeting agencies with 5–20 staff
---
1. The One Reason It Fails
Small agencies do not suffer from a lack of prospecting volume; they suffer from a lack of trust and execution capacity. Cold outbound is fundamentally the wrong tool for their business model. Because creative and professional services are "trust purchases," agencies rely overwhelmingly on referrals and repeat clients. Automated cold emails damage their reputation rather than building a reliable pipeline.
---
2. The Evidence (Interview Quotes)
Trust & Referrals Outweigh Cold Acquisition
- “Content is a trust purchase: you're letting someone write in your company's voice… When the CFO introduced us, the first call was about scope, not about whether we were any good.” — A05 (Content Agency)
- “Every client we have came from someone vouching for us.” — A03 (PR Boutique)
Cold Outbound Fails Small Agencies
- “The replies were mostly people asking us to take them off the list.” — A09 (Paid Social Agency)
- “Our domain ended up on a spam list. Took weeks to fix.” — A11 (WordPress Agency)
- “It cost us about £400 for every meeting, and the meetings didn't close.” — A10 (B2B Marketing Agency)
The Real Bottleneck is Capacity and Follow-Up, Not Leads
- “If you sent me ten more leads in October I'd have to say no to nine of them.” — A08 (Packaging Design Studio)
- “The AI could save me the research hours, sure. It can't do the call... and the follow-up is the job.” — A04 (Shopify Development Agency)
---
3. Where the Evidence Cuts the Other Way
- The Pain is Real: Six of the 14 interviewed agencies explicitly described painful "feast-or-famine" revenue swings, validating the founders' premise that pipeline stability is a major issue.2
- Outbound Can Work: Two agencies (A04 and A07) successfully use outbound to drive 30% to over 50% of new revenue4, proving that agency outbound can function under specific circumstances.3
---
4. What We Would Need to See to Be Proved Wrong
To believe this product can succeed at $600/month for 5–20 staff agencies, we would need to see empirical proof that small agencies without a dedicated in-house business development lead can: 1. Generate positive ROI from cold AI outreach without harming their domain reputation. 2. Convert cold prospects into high-trust creative/professional services contracts without a human-led multi-channel follow-up loop.
---
5. The Cheapest Test
Run a targeted messaging experiment to 50 agencies with 5–20 staff. Pivot the pitch away from a cold-outbound AI SDR toward what A05 actually requested: an automated relationship assistant that tracks past clients and referral networks, drafting warm check-in notes.
If agencies will not engage with or pre-order a warm-relationship retention tool at $200–$600/month, they will certainly reject a cold AI prospecting tool.1
What a PM had to fix
1Test or gate too weakTighten the testSubstantial rework
What we’d changeTest the reason the memo gives: whether agencies without a BD person can turn AI-sent cold email into deals. Demand for a different product doesn't settle that.
2Hypothesis stated as factReframe it as a hypothesisTargeted repair
Source: Founders' hypothesis (from our planning doc)
Small agencies live feast or famine.
What we’d changeSix of 14 is some agencies, not validation of the founders' premise. Say that.
3Contradiction missedSurface the contradictionTargeted repair
What we’d changeSay what sets A04 and A07 apart: a dedicated BD person doing the follow-up. That's the evidence for the memo's own reason, and it's left unused.
4Invented evidenceVerify or remove the claimQuick edit
What we’d changeRemove 'over 50%': A07's share of new revenue from outbound isn't given anywhere.
Check by check
Got wrong · 4
- Avoids unsupported claimsThe unsupported claim about A07's outbound revenue share is presented without qualification, and the test assumes demand for a warm relationship tool rather than testing the AI SDR assumption.
- Produces the required deliverableThe proposed test does not test the viability of the AI SDR; it tests demand for a different product, so the deliverable cannot be acted on without reworking the test.
- Engages the counter-evidenceThe output mentions A04 and A07 but does not say what sets them apart (15-20 staff, dedicated BD owner, follow-up doing the work) or what that means for addressable segment, as required.
- A cheap test that can actually read outThe proposed test shifts to a warm relationship tool, which does not test the assumption behind the AI SDR or measure an early signal for cold outbound; it doesn't fit the brief's request.
Mixed · 1
- Uses the supplied evidence correctlyThe claim that A07 drives 'over 50% of new revenue' from outbound is not supported by the supplied context; no such figure appears for A07.The two graders disagreed on this one.
Got right · 6
- Addresses the actual decisionThe output commits to a clear answer (the product fails) and states what would change the call (empirical proof of positive ROI and conversion without human follow-up).
- Respects explicit constraintsThe output is under 600 words, addresses the founding team, and includes the requested sections.
- Identifies material uncertaintyIt names conditions that would prove the idea wrong and outlines what evidence would be needed.
- Finds the load-bearing assumptionThe output identifies that small agencies' growth is not limited by a shortage of leads, but by trust and capacity, which is the load-bearing assumption.
- Uses the interviews faithfullyAll quotes are accurate and correctly attributed to the right agencies, with no invented quotes.
- Tigers, not paper tigersThe output ranks capacity and trust over mere lead volume, dismisses the feast-or-famine pain as solvable through past clients, and surfaces the real product need (warm outreach) that the founders' proposal avoids.
Claims the judge couldn’t find in the brief
- Two agencies (A04 and A07) successfully use outbound to drive 30% to over 50% of new revenue.
Grades and run details
Decision model 77 · LLM judge 6 of 12 checks
Decision model checks
- passUses the supplied evidence correctly37%
- passAddresses the actual decision84%
- partialRespects explicit constraints24%
- passIdentifies material uncertainty35%
- partialAvoids unsupported claims16%
- partialProduces the required deliverable46%
- passFinds the load-bearing assumption94%
- passUses the interviews faithfully78%
- partialEngages the counter-evidence90%
- partialA cheap test that can actually read out88%
- passTigers, not paper tigers66%
Run
- Run
- #1
- Time to output
- 7 s
- Submitted
- 25 Sept 2026