Needs more than a quick edit9 of 12 checks passed
The guide is strong and mostly usable, but it needs evidence fixes and a must-ask/cut plan before a PM could run it without re-checking assumptions.
Output
Autochase Call Guide: 8 calls, 30 minutes each
What we're trying to learn
Sam's hypothesis has three parts, and each can fail separately: 1. Cost: Late payment hurts in money or time, not just annoyance. (34% of invoices are late, but that doesn't mean it hurts.) 2. Behavior: People already try to fix it. Only 12% of accounts used the manual Send reminder button, so we need to know why the other 88% didn't. 3. Willingness: They'd let software message their clients, and pay $10 a month on top of $12 for Pro.
Interviews can't settle price. They can tell us whether it's worth testing.
Who to talk to
Sam's plan (the 300 heaviest button users, first 8 to reply) would only find people who already like reminders and respond fast to email. Recruit against quotas instead:
- 2 heavy button users (Sam's list is fine for these)
- 3 accounts with several late invoices in the last 90 days who never used the button
- 1 who filed an Autochase support ticket
- 2 agencies (3+ people) and 2 solo freelancers. These overlap with the groups above.
- Mix of Pro and free accounts, with at least 2 free.
Screen for at least one invoice paid late in the last 90 days. Pick people to fill the quotas, not first come first served. Offer a $50 gift card.
Call structure
0:00–0:03 Intro. "We're learning how people get paid, and we're not selling anything. There are no right answers, and critical feedback helps most. Okay to record?"
0:03–0:05 Context. What do you do, how many clients, how many invoices a month?
0:05–0:15 The last late invoice (the core of the call) 1. Tell me about the last time a client paid you late. Probe: how late, how much, what happened first, what did you do, how did it feel? 2. Walk me through what you do between sending an invoice and getting paid. Probe: what tool, how often, who does it? 3. What did that lateness cost you? Probe: hours, cash flow, bills or contractors you had to delay.
0:15–0:21 Current behavior 4. (Button users) What made you use Send reminder? (Non-users) What do you do instead, or why not just let it go? 5. Is there a client you would never chase? Why? 6. Have you tried anything else, such as late fees, deposits, shorter terms or other tools? What happened?
0:21–0:27 Concept. Read this, don't demo it: "Suppose Paydeck automatically emailed your client at set points after the due date, getting firmer each time." 7. What's your first reaction? What worries you? 8. What would have to be true for you to switch it on? Probe: tone, control, per-client exceptions, seeing the emails first.
0:27–0:30 Money and wrap 9. What do you pay for today that helps you get paid or run your finances? 10. "We'd likely charge extra for this. Would you want early access when it's ready?" Note the exact response. Hesitation counts as data. 11. Anything I should have asked? Anyone else we should talk to?
Why Sam's questions changed
- "How frustrating is chasing?" and "Would it be useful?" invite polite yes answers. Past-behavior questions don't.
- "$5, $10 or $15?" anchors people on our options and measures nothing real. Questions 9–10 are better evidence.
- "Would your clients be annoyed?" asks for a guess about someone else. Question 5 asks what they've actually avoided.
- "What other features?" is off-topic. Drop it.
Notes for first-time interviewers
- Split roles. One person asks and the other takes verbatim notes. Swap halfway through the 8 calls.
- Ask about the past, not the future. "What did you do last time?" beats "Would you ever...?"
- Don't pitch or rescue. Don't name Autochase until minute 21. If they struggle to answer, wait, then ask "Can you say more?" Silence is fine.
- Follow the story. Ask "why?" and "what happened next?" more than you ask the next scripted question. The script is a backstop.
- Discount compliments. "That sounds great" is worth little. Specific past actions, money spent and workarounds built are worth a lot.
- Stay neutral if they criticize the product or the idea. Don't defend it.
- Debrief within 30 minutes of each call. Write down three surprises and the best quotes. Do a practice run with a colleague first.
What we'd need to hear
Count only the people who had a late invoice.
Go (build, or run a paid pre-sale test): - At least 5 of 8 describe a specific recent late payment that cost real money, stress or more than about an hour a month. - At least 5 already chase manually or with workarounds, including people who skip the button, and can explain why the button doesn't work for them. - At least 4 would turn on automation given controls we could build, and at least 3 say yes to early access for a paid add-on.
Stop or rethink: - Most shrug lateness off, or say it comes from a few clients they'd never automate against. - The real problem is something else, such as clients not opening invoices, wanting deposits, or needing a payment link, and reminders wouldn't fix it. - Most won't let software contact their clients, or wouldn't pay anything for it.
Mixed results (for example, it only matters for agencies, or only with tone controls): narrow the target segment and run a second round of calls.
Even a clear "go" from 8 calls only justifies a cheap next test, such as a fake-door button or a pre-sale. It doesn't justify committing the quarter.
Check by check
Got wrong · 1
- Fits the callTimings add to 30 minutes, but it does not mark must-ask questions or say what to cut if time runs short.
Mixed · 2
- Uses the supplied evidence correctlyIt uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.The two graders disagreed on this one.
- Avoids unsupported claimsIt presents some interpretations as established, especially that Sam's list would only find reminder-likers/fast email responders and that Autochase tickets exist.The two graders disagreed on this one.
Got right · 9
- Addresses the actual decisionIt commits to a clear call: 8 calls can justify only a cheap next test, not committing the quarter, and gives go/stop conditions.
- Respects explicit constraintsIt is a call guide under 900 words with learning goals, timed questions, interviewer notes, recruiting, and go/stop signals.
- Identifies material uncertaintyIt names material unknowns—cost, behavior, willingness, client-contact concerns—and says how calls plus a cheap test would resolve them.
- Produces the required deliverableMei and Tomas could run the calls from it with light edits.
- Talks to the right peopleIt explains the heaviest-button-user bias and gives a quota mix including non-button users with late invoices, ticket requesters, agencies, solos, and free accounts.
- Coaches first-time interviewersIt gives specific first-timer instructions: role split, verbatim notes, silence, follow-ups, no pitching/defending, debriefs, and a practice run.
- Designed to falsifyEach learning goal has disconfirming questions about recent late payments, actual chasing behavior, client-contact limits, and past spending/early-access response.
- Asks what people did, not what they thinkThe core asks for the last late invoice, what happened, what they did, what it cost, and what they already tried.
- Doesn't lead the witnessIt removes Sam's leading/hypothetical questions, keeps the concept late, and avoids pitching before current-behavior questions.
Claims the judge couldn’t find in the brief
- There is an 'Autochase support ticket' population to recruit from.
- Sam's plan would only find people who already like reminders and respond fast to email.
- We can identify accounts with several late invoices in the last 90 days who never used the button.
Grades and run details
Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
- passUses the supplied evidence correctly26%
- passAddresses the actual decision42%
- passRespects explicit constraints47%
- passIdentifies material uncertainty100%
- passAvoids unsupported claims62%
- passProduces the required deliverable96%
- passTalks to the right people99%
- passCoaches first-time interviewers100%
- passDesigned to falsify95%
- passAsks what people did, not what they think82%
- passDoesn't lead the witness60%
- partialFits the call44%
Run
- Run
- #1
- API response time
- 32 s
- Submitted
- 30 Sept 2026