Usable with a quick edit11 of 11 checks passed
The memo is well-structured, evidence-based, and actionable; a product manager could use it with only light editing.
Output
Founding-team memo: the strongest reason this fails
Reason: The product may solve the wrong bottleneck. Many small agencies are not short of names to contact: they win through trust, referrals and repeat work, and some already lack capacity to take on more. Cold outbound can create low-quality meetings—or reputational risk—without solving the follow-up needed to convert them. At $600/month, an AI SDR could therefore add work rather than prevent famine.
Evidence for it: In 11 of 14 agencies, referrals and repeat clients account for most new revenue (86% on average). Of the nine that tried outbound, seven stopped within six months; none attributed a deal to it. A05’s experience is telling: a lead-gen firm booked “maybe six or seven a month,” but “most of them didn’t have a budget for content.” The founder said, “Content is a trust purchase,” and would pay for prompts and drafts to contact past clients—not “$600 for cold outbound.”
Capacity is also a real constraint: A08 said, “If you sent me ten more leads in October I’d have to say no to nine of them”; A12 needs another designer, not more leads. And even outbound that works appears to require substantial human selling. A04 said, “It works because of the follow-up, not the first email,” then described sending a site teardown, video and making a call. A07 said most meetings were “a waste of time.” The risk is not merely that the AI writes mediocre emails; it is that the agency cannot or does not convert what it sends.
Where the evidence cuts the other way: Six agencies reported feast-or-famine swings, and A05 described panicking after three clients ended together. A04 gets 30% of new revenue from outbound and would pay if the tool saved research time; A07 says its outsourced lead generation “pays for itself, just.” With an $18k average deal, one win could easily cover the subscription. These examples show outbound can work—but the two continuing users are larger agencies with dedicated BD, unlike the founder-led customer we plan to target. A05’s warm reactivation success also suggests the underlying need may be real, but the best channel may not be cold prospecting.
What would prove us wrong: Typical 5–20-person agencies without dedicated BD will pay $600/month for cold outbound, follow up promptly, and generate attributable profitable work—not just meetings—without harming their reputation or exceeding delivery capacity.
Cheapest test: Don’t build the agent. Recruit five founder-led agencies with capacity and a genuine pipeline gap, charge the planned $600/month, and manually run a 12-week, human-approved outbound pilot. Track qualified meetings, founder follow-up time, attributable wins and renewals. Given 6–10-week sales cycles, a strong signal would be at least two attributable wins and three agencies renewing at full price. If interest stops at free trials, meetings fail to convert, or founders cannot follow up, stop or test a warm-client reactivation product instead.
Check by check
Got right · 11
- Uses the supplied evidence correctlyAll factual claims about the current situation are directly supported by the supplied context, with no inventions.
- Addresses the actual decisionThe output commits to a clear answer (the product solves the wrong bottleneck) and states what would prove it wrong.
- Respects explicit constraintsThe memo is written for the founding team, includes all requested sections, and stays under 600 words.
- Identifies material uncertaintyIt names the key unknown (whether typical agencies without dedicated BD will pay and convert) and says how a test would resolve it.
- Avoids unsupported claimsInterpretations and risks are clearly framed as hypotheses, not as established facts.
- Produces the required deliverableThe deliverable is a complete, actionable memo for the founding team, within the word limit.
- Finds the load-bearing assumptionThe memo targets the load-bearing assumption that growth is limited by a shortage of new leads, using referral, abandoned-outbound and capacity evidence.
- Uses the interviews faithfullyAll quotes are accurate and correctly attributed to A05, A08, A04, A07, and A12.
- Engages the counter-evidenceIt names A04 and A07, notes their larger size and dedicated BD, and explains that follow-up, not the first email, drives their success.
- A cheap test that can actually read outThe test is a cheap, manual pilot with a concrete threshold, tracking early signals (qualified meetings, follow-up time) and acknowledging the 6–10 week sales cycle.
- Tigers, not paper tigersIt separates the real bottleneck (trust, capacity, follow-up) from manageable risks, dismisses the idea that outbound never works, and surfaces the unspoken issue that agencies lack conversion capacity.
Grades and run details
Decision model 100 · LLM judge 12 of 12 checks
Decision model checks
- passUses the supplied evidence correctly70%
- passAddresses the actual decision93%
- passRespects explicit constraints66%
- passIdentifies material uncertainty99%
- passAvoids unsupported claims47%
- passProduces the required deliverable94%
- passFinds the load-bearing assumption99%
- passUses the interviews faithfully78%
- passEngages the counter-evidence95%
- passA cheap test that can actually read out66%
- passTigers, not paper tigers95%
Run
- Run
- #1
- API response time
- 27 s
- Submitted
- 29 Sept 2026