Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 9 graded outputs by 5 models. 50% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty100% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Designed to falsify100% pass
    Every learning goal has disconfirming questions, e.g., 'What did you do next?', 'When do you avoid sending another reminder?'
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Asks what people did, not what they think100% pass
    Core questions ask for the most recent overdue invoice and specific actions taken, not opinions or predictions.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Fits the call33% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints81% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Different questions for user and signer81% pass
    The guide does not provide distinct question tracks for the daily warehouse operations manager and the executive who signed; it only differentiates churned vs at-risk customers, with the same questions for both roles within each track.
    GPT-6 Luna · API · Why are warehouses leaving? Two execs, two theories

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Paydeck. Next week, Mei (a designer) and Tomas (an engineer) will run 8 customer calls to test whether we should build Autochase next quarter. Neither has run a customer interview before. Write the call guide they'll use: the learning goals, the questions with rough timings for a 30-minute call, notes for first-time interviewers, who we should talk to, and what we'd need to hear to go ahead or to stop. Keep it under 900 words. Sam, the PM who owns Autochase, has drafted some questions and a recruiting plan. They're below with everything else we know.

About PaydeckInvoicing software for freelancers and small agencies (1 to 10 people). 48,000 accounts, 6,100 of them paying $12 a month for Pro.
The ideaAutochase: automatic, escalating reminders for overdue invoices, sold as a $10-a-month add-on. There is no prototype yet.
Sam's hypothesis“Late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it.”
What we know34% of invoices sent through Paydeck are paid after their due date; the median is 16 days late. The manual 'Send reminder' button was used by 12% of accounts in the last 90 days. 40 support tickets last quarter asked for automatic reminders.
Sam's draft questions1. How frustrating do you find chasing late payments? 2. Would an automatic reminder feature be useful to you? 3. How much would you pay for it: $5, $10 or $15 a month? 4. Do you think your clients would be annoyed by automated reminders? 5. What other features would you like us to build?
Sam's recruiting planEmail the 300 accounts that used the 'Send reminder' button most last month, and book the first 8 who reply.
What a strong answer does

A guide that splits Sam's hypothesis into goals that can each fail: do late payments actually cost users time or money, how do they deal with them today, would they pay to fix it, and does chasing harm client relationships. The core asks for the story of the last overdue invoice (what happened, what they did, how long it took, what it cost) and what they've already tried or paid for, instead of opinions or predictions. It replaces Sam's leading and hypothetical questions: no 'how frustrating', no 'would it be useful', no price menu; payment is tested through past spending or a real commitment (say, a paid pilot sign-up) at the end, after the problem questions and without pitching Autochase first. It fits 30 minutes with timings, coaches first-timers (ask for specifics, follow up with 'what happened next', stay quiet, don't pitch or defend), and fixes the recruiting plan: the heaviest reminder users will over-state demand, so it adds accounts with overdue invoices that never send reminders, and says who to exclude or balance. It ends with clear go and stop signals tied to what people did, such as most calls describing recent chasing that took real time plus a workaround they've tried or paid for, against most saying one nudge does the job.

Critical failures (cap the score)
  • Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision
  • Describes or pitches Autochase before asking how the person handles late payments today
Case

v1.0 · synthetic · B2B SaaS, freelancers, willingness to pay, first-time interviewers

Usable with a quick edit11 of 12 checks passed

The guide is thorough and actionable; a PM could use it with minor edits like marking priority questions for time management.

Output

Autochase discovery call guide

Interviewers: Mei and Tomas · 8 calls, 30 minutes each Decision: Should Paydeck build automatic, escalating overdue-invoice reminders next quarter, priced at $10/month?

1. What we need to learn

Test Sam’s hypothesis—not whether people like the idea:

  • Problem: How often do late payments create meaningful cash-flow problems or chasing work?
  • Current behavior: How do users chase invoices today, and why do some not use Paydeck’s reminder button?
  • Fit and risk: Which reminders could safely be automated? Where might escalation damage client relationships?
  • Value: Is solving this worth an additional $10/month, given existing workarounds and alternatives?

Our data establishes lateness, not demand: 34% of invoices are late, but only 12% of accounts used reminders recently. Support requests and frequent reminder users are useful leads, not representative evidence.

2. Who to recruit

Don’t book the first eight respondents from the top reminder users. That would disproportionately select enthusiastic, high-pain users.

Recruit invoicing-active accounts from these groups:

  • 3: Recurring overdue invoices; frequent Paydeck reminder use.
  • 3: Recurring overdue invoices; little or no Paydeck reminder use.
  • 2: Few overdue invoices, or late payments that appear manageable.

Across the eight, include freelancers and small agencies, plus both Pro and free accounts. Aim for roughly four of each plan; this is purposeful sampling, not a representative survey. Include at most two customers who requested automatic reminders.

Speak to the person who actually manages invoices and chasing; establish whether they also approve software spending. Recruit within each group rather than simply taking the fastest replies. Use a neutral invitation: “Help us understand how you manage invoices and payment follow-up.” Offer the same incentive regardless of feedback.

3. Call questions and timings

0–2 minutes: Welcome and context

“We’re learning how people handle overdue invoices. We’re testing an idea, not selling anything; honest criticism is helpful. There are no right answers.”

Ask permission to record. Confirm their role, business size, and who handles invoices and software purchases.

2–12 minutes: Reconstruct a real example

“Tell me about the most recent invoice that went past its due date.”

Follow the sequence: - When was it due, and when did you notice? - What did you do next? Then what happened? - Who followed up, using what tools, messages, and timing? - How did you decide whether to remind the client again? - Was it paid? What do you think caused the delay?

If comfortable, ask them to show a redacted invoice or follow-up message. Don’t collect client-identifying information.

For low-lateness participants: “Tell me about the last late invoice—or how you usually prevent late payments.”

12–18 minutes: Frequency, consequences, alternatives

  • “Over the last three months, how many invoices needed follow-up?”
  • “Roughly how much time did you spend chasing them?”
  • “What, if anything, did the delay affect?” Probe for concrete consequences, not just frustration.
  • “What have you tried to make this easier? What did it cost?”
  • “Have you used Paydeck’s reminder button? Walk me through why or why not.”
  • “When do you deliberately avoid sending another reminder?”

18–26 minutes: Test the concept, then price

Only now introduce it:

“We’re considering Autochase: scheduled reminders for overdue invoices, with reminders becoming firmer over time. It isn’t built, and details aren’t decided.”

Ask: - “Thinking about that invoice, where would this fit—or not fit?” - “Which steps would you automate, and which would you keep manual?” - “What would you need to control or check before enabling it?” - “What could go wrong? Tell me about a client situation where it would.” - “What would you use instead?”

Then disclose: “We’re considering $10/month, additional to your current plan.”

“What would make that worth paying—or not? Which current effort or expense would it replace? Who would approve it?”

Don’t offer a $5/$10/$15 menu or treat “yes, useful” as purchase evidence.

26–30 minutes: Commitment and close

“If we offered a paid pilot at $10/month, what would you need before deciding?”

Ask whether they’d join a follow-up evaluation or discuss a pilot with the budget owner. Record the specific next step; don’t imply availability or collect payment.

“What important part of chasing payments haven’t we covered?”

4. Notes for first-time interviewers

Alternate moderator and note-taker roles; do one practice call together. Ask one question at a time, allow silence, and follow concrete examples. Avoid praise, pitching, defending, and questions that assume pain. Separate observed behavior, exact quotes, estimates, and your interpretations.

Debrief immediately: frequency, cost, workaround, automation boundaries, price reaction, commitment, and contradictory evidence.

5. Decision rules

Agree these directional gates before interviewing; eight calls cannot estimate market demand.

Go toward a scoped build if at least five describe recurring, consequential pain—including two non-reminder users—and at least three connect $10 to concrete value and take a specific pilot-evaluation step. There must also be a credible way to address relationship and control risks. Validate paid adoption before committing substantial engineering.

Stop or reshape if lateness rarely causes meaningful cost, chasing is easy, automation cannot address the underlying causes, or users prefer manual control and won’t pay.

Mixed findings mean targeted follow-up—not “eight customers validated it.”

Check by check

Got wrong · 1
  • Fits the callTimings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
Got right · 11
  • Uses the supplied evidence correctlyAll factual claims about the current situation are directly from the supplied context.
  • Addresses the actual decisionProvides clear go/stop conditions tied to specific call outcomes, framed for the team.
  • Respects explicit constraintsIncludes all required sections, stays under 900 words, and is written for Mei and Tomas.
  • Identifies material uncertaintyNames specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
  • Avoids unsupported claimsInterpretations are clearly separated from facts; no unsupported causal claims are presented as established.
  • Produces the required deliverableComplete call guide with learning goals, timed questions, coaching, recruiting plan, and decision rules, usable as-is.
  • Talks to the right peopleExplicitly avoids only heavy reminder users, gives a balanced split with rationale and includes non-reminder users.
  • Coaches first-time interviewersProvides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
  • Designed to falsifyEvery learning goal has disconfirming questions, e.g., 'What did you do next?', 'When do you avoid sending another reminder?'
  • Asks what people did, not what they thinkCore questions ask for the most recent overdue invoice and specific actions taken, not opinions or predictions.
  • Doesn't lead the witnessQuestions are neutral and open; Autochase is introduced only after 18 minutes of problem exploration.

Grades and run details

Decision model 96 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly69%
  • passAddresses the actual decision51%
  • passRespects explicit constraints54%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims76%
  • passProduces the required deliverable93%
  • passTalks to the right people100%
  • passCoaches first-time interviewers87%
  • passDesigned to falsify98%
  • passAsks what people did, not what they think82%
  • passDoesn't lead the witness80%
  • partialFits the call56%
Run
Run
#1
API response time
44 s
Submitted
30 Sept 2026

Usable with a quick edit11 of 12 checks passed

The guide is complete and usable; a quick edit to add a note about cutting the check-understanding section if time is short would fix the minor omission.

Output

Autochase: 30-minute customer call guide

Interviewers: Mei leads; Tomas takes notes and asks brief follow-ups at the end. Keep roles consistent across the eight calls.

Decision: Should Paydeck prioritize automatic, escalating overdue-invoice reminders as a $10/month add-on next quarter?

1. Learning goals

Test whether: - Late payments create recurring, material costs—not just irritation. - Chasing takes enough time or money that customers actively try to reduce it. - Reminders address the cause of delays, rather than disputes, approval processes, or clients lacking funds. - Customers would trust automation, and what control they require. - There is credible demand at $10/month, beyond polite enthusiasm.

Starting evidence is suggestive, not conclusive: 34% of invoices are late, with a median delay of 16 days. That does not tell us how many accounts experience serious pain. Reminder usage (12% of accounts) may miss chasing outside Paydeck. The 40 support requests show interest, not willingness to pay.

2. Who to recruit

Do not book the first eight respondents from the 300 heaviest reminder users. That would favor highly engaged users already chasing through Paydeck.

Recruit purposively: - 3 accounts: recurring overdue invoices; frequent manual reminders. - 3 accounts: recurring overdue invoices; little or no manual reminder use. - 2 accounts: mostly on-time invoices, as a contrast group.

Across those groups, aim for six Pro and two free accounts, with both freelancers and small agencies. Include no more than two support-ticket requesters. Speak to whoever actually handles collections; establish whether they also approve software spending.

Randomly invite accounts within each group, then fill group slots rather than accepting the first eight overall. Replace no-shows within the same group. Use a neutral invitation: “Help us understand how you manage invoices and payments.” If offering an incentive, keep it fixed and unrelated to feedback.

3. Call script and timings

0–3 minutes: Welcome and permission

“Thanks for helping. We’re learning how people manage payments—not testing you or selling anything. Honest criticism is useful. You can skip anything confidential.”

Ask permission before recording. Otherwise, take notes.

“What’s your role, who handles overdue invoices, and roughly how many invoices do you send each month?”

3–12 minutes: Reconstruct a real incident

“Tell me about the most recent invoice that wasn’t paid by its due date.”

Follow the timeline: - “When was it due, and when did you notice?” - “What did you do next? Then what happened?” - “Who else was involved?” - “What reason did the client give for the delay?” - “Was it paid? When?”

“Could you walk us through a redacted reminder or your tracking process?” Do not request client-identifying information.

If they have no recent late invoice: “Walk us through how your last invoice got paid. What helps payments arrive on time?”

12–19 minutes: Frequency, consequences, alternatives

“In the last three months, about how often has this happened?”

“For that example, roughly how much time did you spend chasing? What, if anything, did the delay cost you?”

“What happens if you don’t follow up?”

“How do you decide when and how firmly to chase? Have you ever chosen not to?”

“What have you tried to improve this? What did it cost, and why did you keep or stop using it?”

If relevant: “Do you use Paydeck’s reminder button? Why or why not?”

19–23 minutes: Automation fit and risks

“Which parts of that process need your judgment?”

“Tell me about a follow-up that went badly—or worked particularly well.”

“What situations would make sending another reminder inappropriate?”

Probe only after open answers: payments already made, disputed invoices, important relationships, tone, timing, approvals.

23–28 minutes: Concept and price

Only now introduce the idea:

“We’re considering Autochase: automatic reminders that escalate while an invoice remains overdue. There’s no prototype, and details aren’t decided.”

“For the invoice we discussed, where would this help, fail, or create extra work?”

“What would you need to control before trusting it?”

“The proposed price is $10/month, on top of your current plan. How would you decide whether to buy it? What would it replace? Who approves that spending?”

Ask interested participants whether they will join a follow-up pilot discussion. Record this as interest, not a purchase.

28–30 minutes: Check understanding

Summarize their situation and ask what you misunderstood. Tomas asks outstanding clarifications. Thank them.

4. First-time interviewer notes

  • Ask one question at a time; allow silence.
  • Prefer “Tell me about the last time” to predictions.
  • Don’t defend Autochase, suggest answers, or praise positive reactions.
  • Separate quotes, observed behavior, estimates, and interpretation.
  • Debrief immediately: pain, frequency, cost, workaround, automation fit, price evidence, contradictions.

5. Decision rules

Agree these directional thresholds before interviewing:

Advance: At least four of six overdue-account participants demonstrate recurring, consequential pain and reminder-suitable workflows; at least three Pro participants show credible $10 value grounded in current costs or spending and accept a concrete follow-up. Proceed to a scoped pilot and real pricing test—not a full-build commitment from interviews alone.

Stop or rethink: Pain is mostly minor, delays are not reminder-solvable, relationship risks outweigh savings, or interest disappears at $10.

Mixed evidence: Narrow the target segment and recruit more. Eight purposive interviews cannot estimate market demand.

Check by check

Got wrong · 1
  • Fits the callThe guide does not mark must-ask questions or say what to cut if time runs short, as required by the criterion.
Got right · 11
  • Uses the supplied evidence correctlyAll factual claims about the current situation are taken directly from the brief or supplied context.
  • Addresses the actual decisionThe guide commits to clear go/stop decision rules with thresholds and says what would change the answer.
  • Respects explicit constraintsThe output is under 900 words, includes all requested sections, and respects the form and reader.
  • Identifies material uncertaintyIt names specific unknowns (e.g., whether pain is minor, delays not reminder-solvable) and says how the interviews will resolve them.
  • Avoids unsupported claimsInterpretations are clearly labelled as suggestive, and no unsupported claims are presented as established fact.
  • Produces the required deliverableThe guide is a complete, usable call guide with learning goals, timed questions, interviewer notes, recruiting plan, and decision rules.
  • Talks to the right peopleIt explicitly avoids only the heaviest reminder users, explains why, and balances with users who have overdue invoices but don't chase in Paydeck.
  • Coaches first-time interviewersIt gives specific, usable instructions: ask one question at a time, allow silence, prefer past-behaviour questions, don't defend or pitch, and separate notes.
  • Designed to falsifyEvery learning goal has disconfirming questions, such as reconstructing a real late invoice and asking what happens if they don't follow up.
  • Asks what people did, not what they thinkCore questions ask for the most recent overdue invoice and what they did, with opinions and predictions kept secondary and at the end.
  • Doesn't lead the witnessQuestions are neutral and open; Autochase is introduced only after the problem exploration, with no pitching beforehand.

Grades and run details

Decision model 96 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly64%
  • passAddresses the actual decision52%
  • passRespects explicit constraints45%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims92%
  • passProduces the required deliverable94%
  • passTalks to the right people100%
  • passCoaches first-time interviewers86%
  • passDesigned to falsify89%
  • passAsks what people did, not what they think90%
  • passDoesn't lead the witness86%
  • partialFits the call51%
Run
Run
#1
Time to output
38 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.892.31None
2GPT-6.1 SolwithAPI92.192.62None
3Opus 5.5withClaude94.169.22None
4Sonnet 5.5withAPI86.484.62None
5GPT-6 LunawithAPI90.273.921 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review