Needs more than a quick edit7 of 11 checks passed
It is over the requested length and contains two unsupported current-state claims that an editor would need to remove or verify. Otherwise the reasoning is usable, but those are more than wording fixes.
Output
Pre-mortem: AI SDR for 5–20 person agencies
The one reason it fails
Small agencies do have a pipeline problem, but cold outbound is the wrong fix. Their work is sold on trust, so it comes through introductions. Cold email produces meetings that don't close, and the few agencies where outbound works have a person doing the follow-up. At $600 a month we would be selling a channel that this buyer has already tried and dropped, into a market where two funded competitors are moving down.
Evidence
- Referrals dominate. 11 of 14 agencies get about 86% of new revenue from referrals and repeat clients. A05: "Content is a trust purchase." A03: "Every client we have came from someone vouching for us."
- Outbound has a poor track record. 9 tried it, 7 quit within six months, and none of the 7 could attribute a closed deal to it. The reasons are structural: A01 "meetings with people who'd never buy from us"; A09 "mostly people asking us to take them off the list"; A11 "Our domain ended up on a spam list"; A02 "LinkedIn shut our founder's account down. That's our best referral network."
- The cost is high. A10: "about £400 for every meeting, and the meetings didn't close."
- Success depends on humans. A04: "It works because of the follow-up, not the first email… The AI could save me the research hours, sure. It can't do the call." And: "most agencies our size don't have someone like me… Most tried it once, got burned, and went back to referrals."
- The famine is real, but warm outreach solved it. A05 recovered by emailing three years of past clients, and two came back within a month. Of our product, A05 said: "Cold? Probably not… Six hundred for cold outbound, no."
- Some agencies don't need leads. A08, A12 and A14 turn work away. A08: "It would make my problem worse."
Where the evidence cuts the other way
- 6 of 14 report feast-or-famine, so the pain is real.
- A07 kept outbound: "It pays for itself, just… one in ten turns into a retainer." A04 would pay $600 if the tool replaced Apollo and Clay.
- None of the failures used a modern AI agent. Better targeting and lower cost per meeting might change the results, though the trust and deliverability complaints suggest it won't fix most of them.
- The two who kept going are not cleanly separable by size. Of the four agencies with 15–20 staff, two stopped (A10, A13).
- A05 would pay about £200 a month for a warm-relationship tool. That points to a different product at a third of our price, not proof that ours works.
- 14 interviews is a small sample, and willingness to pay was hypothetical.
What would prove us wrong
In a paid pilot with agencies that have no BD hire, the product produces qualified pipeline at a cost they'll keep paying for. Roughly: at least half the pilot agencies get a qualified opportunity or proposal within 90 days, and most say they'd continue at $600. The maths also needs checking. $600 a month is $7,200 a year, but the buyer's gain is the margin on an $18k deal, not the revenue. "Many times over" only holds if outbound reliably closes more than one deal a year.
Cheapest test
Run a concierge pilot before building anything more:
- Recruit 10 agencies with 5–14 staff and no BD person. Exclude the referral-saturated ones like A08, A12 and A14.
- Charge $600 a month, or at least a serious deposit. Free pilots tell us nothing about willingness to pay.
- Do the work ourselves with off-the-shelf tools (Apollo, Clay, an email sender) on separate sending domains. The agency approves every email, since A04 and A05 both demanded that.
- Set the pass line before starting: at least 5 of 10 agencies get a qualified proposal within roughly 14 weeks (the 6–10 week sales cycle plus setup), and at least 6 of 10 agree to continue.
- Track reply rate, unsubscribe and spam complaints, cost per qualified meeting, and meeting-to-proposal rate.
This costs a few weeks of founder time and about £1–2k in tools, roughly 4 months of the 20-month runway. In parallel, offer 5 of the same agencies a £200 warm-outreach version, to see whether the real product is relationship follow-up.
If the pilot misses the line, don't spend the year on cold outbound for this segment.
Check by check
Got wrong · 2
- Respects explicit constraintsIt exceeds the 600-word limit, running roughly 700 words, even though it includes the requested sections.
- Avoids unsupported claimsIt presents the competitor 'moving down' claim and the 'none used a modern AI agent' claim as fact rather than as labelled inference.
Mixed · 2
- Uses the supplied evidence correctlyIt includes at least two current-state claims, that competitors are moving down and that none of the failures used a modern AI agent, that are not supported by the supplied context.The two graders disagreed on this one.
- Produces the required deliverableIt has the right form and reader, but it is not within the requested length.The two graders disagreed on this one.
Got right · 7
- Addresses the actual decisionIt commits early to the answer that cold outbound is the wrong fix for this segment and says what pilot result would change that.
- Identifies material uncertaintyIt names the small sample, hypothetical willingness to pay, and the untested modern-AI possibility, and proposes thresholds and observable signals to resolve them.
- Finds the load-bearing assumptionIt targets the assumption that growth is lead-constrained, using the referral, abandoned-outbound and turning-work-away evidence as the central challenge.
- Uses the interviews faithfullyAll quotes or close paraphrases match the transcripts and log, and each is attributed to the correct agency.
- Engages the counter-evidenceIt names A04 and A07, identifies their size/BD-owner/tooling and follow-up dependence, and explains what that means for the addressable segment.
- A cheap test that can actually read outIt proposes a cheap concierge pilot with explicit thresholds and measures qualified proposals and meetings rather than won deals, matching the 6–10 week sales cycle.
- Tigers, not paper tigersIt separates the central trust and capacity risks from fixable targeting/cost issues and surfaces the warm-outreach product alternative and margin-math problem.
Claims the judge couldn’t find in the brief
- Two funded competitors are moving down into this market.
- None of the failed outbound attempts used a modern AI agent.
Grades and run details
Decision model 91 · LLM judge 7 of 12 checks
Decision model checks
- passUses the supplied evidence correctly37%
- passAddresses the actual decision95%
- partialRespects explicit constraints30%
- passIdentifies material uncertainty100%
- partialAvoids unsupported claims25%
- passProduces the required deliverable79%
- passFinds the load-bearing assumption86%
- passUses the interviews faithfully83%
- passEngages the counter-evidence75%
- passA cheap test that can actually read out99%
- passTigers, not paper tigers98%
Run
- Run
- #1
- API response time
- 30 s
- Submitted
- 29 Sept 2026