Needs more than a quick edit9 of 12 checks passed
A PM could use most of it with light edits, but the prohibited hypothetical price/trial questions and missing must-ask/cut guidance need correction before it is safe to run.
Output
Autochase customer call guide (30 minutes)
What we need to learn
- How often late payments happen, and what customers actually do when they do.
- Whether chasing creates meaningful time, cash-flow, or relationship costs—and how customers handle those today.
- When automatic reminders would help, and what control or safeguards customers would need.
- Whether the value could justify a $10/month add-on.
Who to talk to
Recruit 8 account owners or purchase decision-makers who send invoices through Paydeck and have had at least one late invoice in the last 90 days. Aim for: - 4 who used “Send reminder” and 4 with late invoices who rarely or never used it. - A mix of freelancers and small agencies, account sizes, and Pro/free plans.
Don’t recruit only the 300 most frequent button users: they’re likely to overrepresent people already motivated to chase. Ask about recent late invoices; don’t tell people the feature idea before the call.
30-minute call
0–2 min | Set the tone “Thanks for speaking with us. We’re learning how invoicing works for you; there are no right answers. I’m interested in what you’ve actually done, not what you think we want to hear. We’re exploring ideas, not selling a product.” Ask permission before recording; otherwise take notes.
2–8 min | Recent experience “Think of the most recent invoice that was paid late. Can you walk me through what happened, from sending it to getting paid?” Probe: When was it due and paid? When did you realize it was late? What did you do next? What happened?
8–15 min | Current workflow “What, if anything, did you do to follow up?” Probe: How did you decide when and how to contact them? What did you send, and through what channel? How often does this happen? What do you do when you don’t send a reminder? Have you tried tools or changed your process?
15–20 min | Impact and trade-offs “What effect did that late payment have on your work or business?” Probe: Time spent? Cash-flow impact? Client relationship? What was most frustrating—or not especially difficult? How do you prioritize which invoices to chase?
20–25 min | Explore the idea After hearing their story, say neutrally: “We’re exploring a paid add-on that sends escalating reminders automatically when an invoice becomes overdue. There isn’t a product yet.” “What’s your first reaction?” Probe: What would be useful or unhelpful? Which invoices, if any, would you trust it with? What would you need to control—timing, wording, recipients, stopping reminders? What situations would make you not use it?
25–28 min | Value and price “How do you currently spend time or money dealing with late invoices?” “If this cost $10 a month, how would you decide whether it was worth it?” Probe: What would it need to save or improve? Would you trial it at that price? Why or why not? Treat stated intent as a clue, not proof of demand.
28–30 min | Close “What have I not asked that matters about getting paid late?” Thank them. Don’t promise a launch or specific functionality.
Notes for first-time interviewers
- Mei leads; Tomas takes notes. Agree beforehand who will ask follow-ups. Don’t interrupt or tag-team.
- Let the customer finish. Use silence; ask “Can you tell me more?” or “What happened next?”
- Ask about a specific past event before opinions or hypotheticals. Avoid “Would this be useful?”, “How frustrating is it?”, feature wishlists, and leading follow-ups.
- Don’t defend the idea, pitch solutions, or share Paydeck’s late-payment statistics; these can bias answers.
- Capture examples, exact phrases, workarounds, and contradictions—not just “yes/no.” Separate what people did from what they say they might do.
Decision after 8 calls
These calls are directional, not a market-size estimate. Go ahead with a narrowly scoped MVP/pilot next quarter if at least 5 of 8 describe recurring late-payment problems with concrete business or time costs, and at least 4 of 8 see a credible reason to try automation at $10/month and can explain the controls they’d need. Check that interest isn’t confined to one narrow user type.
Stop or defer if 2 or fewer describe meaningful recurring burden, most already handle it easily, or trust/client concerns make automation unacceptable even with reasonable controls. If evidence is mixed or concentrated in one segment, don’t commit broadly: refine the target or concept and test again.
Check by check
Got wrong · 1
- Fits the callThe timings add to 30 minutes, but the guide does not mark must-ask questions or say what to cut if time runs short.
Mixed · 2
- Respects explicit constraintsIt violates the explicit prohibition on asking people to predict what they would pay or use by asking whether they would trial it at $10/month.The two graders disagreed on this one.
- Doesn't lead the witnessThe price section asks hypothetical willingness-to-pay/trial questions, and the idea section asks for first reactions to a described product, which are leading relative to the prohibited practices.The two graders disagreed on this one.
Got right · 9
- Uses the supplied evidence correctlyThe output uses the supplied facts correctly and labels recruiting assumptions and likely bias as hypotheses rather than established facts.
- Addresses the actual decisionIt gives clear go, stop/defer, and mixed-evidence decision rules tied to specific call outcomes.
- Identifies material uncertaintyIt names key unknowns such as recurring burden, controls/trust concerns, and segment concentration, and says how calls would resolve them.
- Avoids unsupported claimsIt avoids presenting unproven causes or demand as fact, using language such as “likely” and “directional.”
- Produces the required deliverableIt provides a usable 30-minute call guide with learning goals, questions, timings, interviewer notes, recruiting guidance, and decision criteria within the length limit.
- Talks to the right peopleIt explicitly avoids recruiting only the 300 heaviest reminder users and balances the sample with 4 reminder users and 4 non/rare users with late invoices.
- Coaches first-time interviewersIt gives first-time interviewers concrete follow-up prompts, silence guidance, note-taking rules, and prohibitions on pitching or defending.
- Designed to falsifyEach learning goal has questions that could falsify the hypothesis, such as asking what they actually did, what it cost, and what would make them not use automation.
- Asks what people did, not what they thinkThe core questions ask for the most recent late invoice, what happened, what they did, and current time or money spent.
Claims the judge couldn’t find in the brief
- The output’s recruiting filter assumes there are at least 8 account owners or purchase decision-makers with a late invoice in the last 90 days.
- The heaviest reminder users are likely to overrepresent people already motivated to chase.
Grades and run details
Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
- passUses the supplied evidence correctly66%
- passAddresses the actual decision29%
- passRespects explicit constraints19%
- passIdentifies material uncertainty92%
- passAvoids unsupported claims78%
- passProduces the required deliverable95%
- passTalks to the right people100%
- passCoaches first-time interviewers100%
- passDesigned to falsify85%
- passAsks what people did, not what they think66%
- passDoesn't lead the witness70%
- partialFits the call32%
Run
- Run
- #1
- API response time
- 31 s
- Submitted
- 30 Sept 2026