Usable with a quick edit10 of 11 checks passed
The memo is structurally complete and well‑reasoned; a quick edit to shift the test gate to an early signal (e.g., qualified proposals) would make it fully usable without rechecking evidence or redoing the reasoning.
Output
To: Founding team Subject: We risk automating a channel our customers cannot turn into revenue
Do not commit a year to this product yet. The strongest reason it fails is that finding prospects and sending personalised emails does not remove the constraint on small agencies’ sales: the human work of establishing trust and converting interest. Our promise assumes an agent can keep revenue flowing while the founder stays busy delivering. The interviews suggest we would generate more work for that same founder.
Eleven of 14 agencies get most new revenue from referrals and repeat clients, averaging 86% across those eleven. Of nine that tried outbound, seven stopped within six months without an attributable deal. This is more than unfamiliarity with the channel: most have bought or attempted it and abandoned it.
A05 shows the mechanism. Its lead-generation agency booked six or seven meetings a month, but prospects lacked budgets or wanted evidence of relevant experience. With a referred prospect, “the first call was about scope, not about whether we were any good”. Personalisation may earn attention; these interviews do not establish that it supplies the credibility needed to buy.
The successful case makes the constraint clearer. A04 generates 30% of new revenue through outbound, but has a dedicated BD lead who follows emails with site teardowns, videos and calls: “They don’t have the time to follow up, and the follow-up is the job.” Automating the first step leaves that job with the founder.
Some agencies do not need additional demand at all. A08 says, “Getting them isn’t. Handling them is,” and needs a senior designer. Three agencies regularly turn work away. Feast or famine is not a universal unmet need.
What cuts against this conclusion
Six agencies do report feast-or-famine swings. A04 and A07 prove outbound can work, although both have dedicated business-development capacity. A04 would pay $600 if we replaced research tools, saved time, integrated with HubSpot and required email approval. That supports a narrower research assistant proposition.
Past failures also do not prove our product would fail: targeting and execution may improve, and A05’s two-month trial barely covered the typical 6–10-week sales cycle. Fourteen interviews cannot establish market prevalence.
But “one extra client pays for it” skips the disputed step: winning that client. The $7,200 annual subscription is 40% of the $18,000 average deal’s revenue, before delivery costs and selling time.
What would prove us wrong, and the cheapest test
Run a paid concierge pilot before building. Recruit five agencies with spare delivery capacity, a stated growth goal and no dedicated BD person. Charge the planned $600 monthly price. Using existing tools, manually deliver prospect research and personalised drafts, with approval before sending. Founders must handle follow-up; we must not secretly supply the missing salesperson.
Run for 12 weeks to cover the stated sales cycle.1 Agree a founder-time ceiling upfront, record all selling time and track qualified opportunities through to attributable signed deals, contribution margin and renewal.
A proposed gate: at least three of five win business whose expected contribution covers subscription and selling costs, remain within the time ceiling and renew at full price. Meetings alone do not pass.
If only agencies with BD staff succeed, narrow the customer and promise. A05’s request for past-client reminders suggests another test, but “I’d pay for that tomorrow” is not payment. Do not turn one appealing interview into the next year-long commitment.
What a PM had to fix
1Test or gate too weakTighten the testQuick edit
What we’d changeSigning and renewing inside 12 weeks is a bar that quietly favours the memo's own conclusion. Extend the window or add a leading indicator such as qualified proposals, and include the risk to agencies' LinkedIn channel.
Check by check
Got wrong · 1
- A cheap test that can actually read outThe proposed test gate relies on signed deals (won business), but a 6–10 week sales cycle means the 12-week test likely cannot observe this late signal; the criteria require measuring an early signal like qualified meetings or proposals.
Got right · 10
- Uses the supplied evidence correctlyAll factual claims about the current situation are drawn accurately from the brief and supplied context.
- Addresses the actual decisionThe output gives a clear 'do not commit a year' answer early, framed for the founding team, and specifies what would change it.
- Respects explicit constraintsThe memo is under 600 words, addresses the founding team, and includes the required one reason, evidence, counter-evidence, proof-wrong condition, and cheapest test.
- Identifies material uncertaintyIt acknowledges that past failures do not prove the product will fail, the small sample limits prevalence estimates, and states what results would change the call.
- Avoids unsupported claimsInterpretations are clearly labelled as hypotheses or suggestions, not presented as established facts.
- Produces the required deliverableThe memo is a complete, usable document in the requested form, length, and reader, requiring only light editing.
- Finds the load-bearing assumptionThe critique centers on the assumption that small agencies' growth is limited by a shortage of new leads, showing trust and capacity as the real constraints.
- Uses the interviews faithfullyQuotes and paraphrases from A05, A04, A08 are accurate and correctly attributed, with no invented material.
- Engages the counter-evidenceA04 and A07 are named, their dedicated BD capacity and follow-up are highlighted, and the implication of a narrow addressable segment is stated.
- Tigers, not paper tigersIt dismisses the worry that a small sample invalidates the finding by noting that 14 interviews cannot establish prevalence, and surfaces the unspoken issue that outbound adds follow‑up work the founder cannot do.
Grades and run details
Decision model 95 · LLM judge 11 of 12 checks
Decision model checks
- passUses the supplied evidence correctly59%
- passAddresses the actual decision99%
- passRespects explicit constraints40%
- passIdentifies material uncertainty99%
- passAvoids unsupported claims21%
- passProduces the required deliverable86%
- passFinds the load-bearing assumption99%
- passUses the interviews faithfully81%
- passEngages the counter-evidence81%
- partialA cheap test that can actually read out42%
- passTigers, not paper tigers98%
Run
- Run
- #1
- Time to output
- 72 s
- Submitted
- 25 Sept 2026