Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 9 graded outputs by 5 models. 50% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty100% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Designed to falsify100% pass
    Every learning goal has disconfirming questions, e.g., 'What did you do next?', 'When do you avoid sending another reminder?'
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Asks what people did, not what they think100% pass
    Core questions ask for the most recent overdue invoice and specific actions taken, not opinions or predictions.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Fits the call33% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints81% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Different questions for user and signer81% pass
    The guide does not provide distinct question tracks for the daily warehouse operations manager and the executive who signed; it only differentiates churned vs at-risk customers, with the same questions for both roles within each track.
    GPT-6 Luna · API · Why are warehouses leaving? Two execs, two theories

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Paydeck. Next week, Mei (a designer) and Tomas (an engineer) will run 8 customer calls to test whether we should build Autochase next quarter. Neither has run a customer interview before. Write the call guide they'll use: the learning goals, the questions with rough timings for a 30-minute call, notes for first-time interviewers, who we should talk to, and what we'd need to hear to go ahead or to stop. Keep it under 900 words. Sam, the PM who owns Autochase, has drafted some questions and a recruiting plan. They're below with everything else we know.

About PaydeckInvoicing software for freelancers and small agencies (1 to 10 people). 48,000 accounts, 6,100 of them paying $12 a month for Pro.
The ideaAutochase: automatic, escalating reminders for overdue invoices, sold as a $10-a-month add-on. There is no prototype yet.
Sam's hypothesis“Late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it.”
What we know34% of invoices sent through Paydeck are paid after their due date; the median is 16 days late. The manual 'Send reminder' button was used by 12% of accounts in the last 90 days. 40 support tickets last quarter asked for automatic reminders.
Sam's draft questions1. How frustrating do you find chasing late payments? 2. Would an automatic reminder feature be useful to you? 3. How much would you pay for it: $5, $10 or $15 a month? 4. Do you think your clients would be annoyed by automated reminders? 5. What other features would you like us to build?
Sam's recruiting planEmail the 300 accounts that used the 'Send reminder' button most last month, and book the first 8 who reply.
What a strong answer does

A guide that splits Sam's hypothesis into goals that can each fail: do late payments actually cost users time or money, how do they deal with them today, would they pay to fix it, and does chasing harm client relationships. The core asks for the story of the last overdue invoice (what happened, what they did, how long it took, what it cost) and what they've already tried or paid for, instead of opinions or predictions. It replaces Sam's leading and hypothetical questions: no 'how frustrating', no 'would it be useful', no price menu; payment is tested through past spending or a real commitment (say, a paid pilot sign-up) at the end, after the problem questions and without pitching Autochase first. It fits 30 minutes with timings, coaches first-timers (ask for specifics, follow up with 'what happened next', stay quiet, don't pitch or defend), and fixes the recruiting plan: the heaviest reminder users will over-state demand, so it adds accounts with overdue invoices that never send reminders, and says who to exclude or balance. It ends with clear go and stop signals tied to what people did, such as most calls describing recent chasing that took real time plus a workaround they've tried or paid for, against most saying one nudge does the job.

Critical failures (cap the score)
  • Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision
  • Describes or pitches Autochase before asking how the person handles late payments today
Case

v1.0 · synthetic · B2B SaaS, freelancers, willingness to pay, first-time interviewers

Needs more than a quick edit9 of 12 checks passed

The guide is strong and actionable, but it violates the 900-word constraint and lacks explicit must-ask/cut guidance for the 30-minute call. A PM would need to trim and tighten it before using it as-is.

Output

Autochase Discovery Calls: Interview Guide

Interviewers: Mei and Tomas | Calls: 8 × 30 min | Week of: [date]

Why we're doing this

Sam's hypothesis is that late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it. The data leaves room for doubt. 34% of invoices are paid late, but only 12% of accounts used "Send reminder" in the last 90 days. Either chasing isn't painful, or people chase outside Paydeck, or they avoid chasing altogether. These calls should tell us which.

Learning goals

  1. Cost: How much do late payments actually hurt (cash flow, hours, stress)? Is this a top-three problem or a mild annoyance?
  2. Current behaviour: How do people chase today, with what tools, and why do most not use our reminder button?
  3. Barriers to automation: Would they let software chase their clients? What would make that feel safe or risky?
  4. Willingness to pay: Do they already spend money or time on this, and who (freelancers or agencies, Pro or free) feels it most?

Who we talk to

Draw from accounts with at least one invoice paid 7+ days late in the last 90 days, then fill these slots:

SegmentCallsWhy
Regular "Send reminder" users3They feel the pain and act on it
Late invoices, never used the button3The 88% we know least about
Filed a ticket asking for auto-reminders2Our strongest likely advocates, to test them
  • Aim for at least 5 Pro accounts, since the add-on is priced against a $12 plan.
  • Include at least 2 agencies.
  • Invite about 15 people to get 8 who show up.
  • Hand-pick from the people who reply. Don't just book the first eight.
  • Offer a thank-you, such as a free month of Pro.

Why this differs from Sam's plan: Emailing only the 300 heaviest button users would reach people who already chase inside Paydeck. They're the most likely to say yes, so a yes from them wouldn't tell us much. Booking the first eight repliers adds another bias toward the most engaged users.

The call (30 minutes)

0–3 min: Open - "Thanks for your time. We're trying to understand how you get paid. We're not selling anything, and there are no wrong answers." - Ask permission to record.

3–8 min: Context - "Tell me about your business. Who are your clients, and how many invoices do you send in a typical month?" - "Walk me through what happens after you send an invoice."

8–18 min: The last late invoice (the heart of the call) - "Think of the last time a client paid you late. What happened?" - "When did you notice? What did you do, and when?" - "How did it get resolved? How long did it take?" - "What did it cost you: time, money, other plans, stress?" - "Is that typical, or was it unusual?"

18–24 min: How they chase today - "What do you usually do when an invoice is overdue?" - "What tools do you use: email, phone, our reminder button, anything else?" - If they haven't used the button: "Have you seen our 'Send reminder' button? Why or why not?" - "Is there anyone you chose not to chase? Why?" - "What worries you about chasing clients?"

24–28 min: Value and alternatives - "Have you tried to fix this: templates, other software, an assistant, late fees?" - "What did that cost, and why did you keep or drop it?" - "If this problem disappeared tomorrow, what would change for you?"

28–30 min: Close - Describe Autochase in one or two neutral sentences: "automatic reminders for overdue invoices, sent on a schedule you set." - Ask for their reaction. - Then ask: "Would you want to join a paid beta when it's ready?" A yes should come with a real next step, such as their email for the beta list. - Finish with: "Anything I should have asked?"

Why this differs from Sam's questions: Questions like "Would this be useful?" or "How frustrating is it?" invite polite yeses. Offering $5/$10/$15 anchors people on our numbers. Asking what their clients would think gets guesses, not evidence. "Other features" takes the call off topic. Asking what actually happened last time gives us far more reliable answers.

Notes for first-time interviewers

  • Split roles. One person leads and the other takes notes. Swap after each call.
  • Ask about the past, not the future. "What did you do?" beats "What would you do?"
  • Follow up before moving on. "Tell me more," "Why?" and "Can you give an example?" are your best questions.
  • Get comfortable with silence. Count to five before filling a pause. People often add the useful part then.
  • Don't pitch or defend. If they criticise Paydeck, say "That's helpful, go on."
  • Treat compliments as noise. "Sounds great!" means nothing. Look for money spent, time lost and workarounds built.
  • Write down exact quotes. Mark the ones about cost or emotion.
  • Debrief for 10 minutes right after each call. Cover the top three surprises, the evidence for or against each learning goal, and whether they meet the go signals below.

Decision criteria

We're agreeing these before the calls so we don't rationalise afterward. Eight calls give us direction, not proof. Treat the results as a reason to prototype or to drop the idea, not as a sizing exercise.

Go ahead (build a prototype next quarter) if: - At least 5 of 8 describe a recent late payment with a concrete cost, such as a cash-flow squeeze, hours spent chasing, or a delayed bill of their own; and - At least 4 already use a workaround (templates, calendar nudges, another tool) or have paid for one; and - At least 3 give a real commitment at the close, such as joining the beta list or asking when it launches; and - Relationship worries come up but sound solvable, for example with control over tone, timing or which clients get reminders.

Stop or rethink if: - Most treat late payment as normal and tolerable ("they always pay eventually"); or - Most avoid chasing on purpose to protect client relationships, and automation makes them more nervous; or - The pain is real but they're satisfied with what they already use; or - Interest is limited to the ticket-filers and heavy button users.

Possible middle outcome: People want the feature but balk at paying extra. That points to making it part of Pro rather than a $10 add-on, which nearly doubles the price. Flag this if you hear it.

After the calls

Mei and Tomas share a one-page summary with Sam within two days. It should cover each learning goal, the go/stop tally, and the three best quotes.

Check by check

Got wrong · 1
  • Fits the callThe timings add to 30 minutes, but the guide does not mark must-ask questions or say what to cut if time runs short.
Mixed · 2
  • Respects explicit constraintsThe guide is clearly over 900 words, violating the explicit length constraint.The two graders disagreed on this one.
  • Produces the required deliverableIt is a usable interview guide with goals, questions, timings, recruiting, coaching, and decision criteria, but it exceeds the required 900-word limit.The two graders disagreed on this one.
Got right · 9
  • Uses the supplied evidence correctlyThe output's factual statements about Paydeck, Autochase, Sam's plan, and the supplied metrics are supported by the brief or follow by arithmetic; its speculative interpretations are framed as hypotheses or interview principles.
  • Addresses the actual decisionIt addresses the go/stop decision by giving explicit, pre-agreed go and stop criteria tied to observed behaviour and commitment, while correctly treating the calls as the next step rather than deciding now.
  • Identifies material uncertaintyIt names the key unknowns—whether late payments cause real cost, why people don't use the button, whether automation threatens client relationships, and whether users will pay—and says how the calls and go/stop thresholds would resolve them.
  • Avoids unsupported claimsIt labels uncertain claims as hypotheses or likely biases and does not present invented customer facts or forecasts as established evidence.
  • Talks to the right peopleIt explains that heavy reminder users overstate demand and balances them with late-invoice non-users and ticket filers, with a rough 3/3/2 split for 8 calls.
  • Coaches first-time interviewersIt gives first-time interviewers specific follow-ups, silence guidance, note-taking and debrief instructions, and a clear rule against pitching or defending.
  • Designed to falsifyEach learning goal has questions that could falsify the hypothesis, such as asking for the last late invoice, why they didn't use the button, relationship worries, and past spending or workarounds.
  • Asks what people did, not what they thinkThe core 8–18 minute section asks for the last late invoice and what the person actually did, with follow-ups on timing, cost, and typicality.
  • Doesn't lead the witnessThe questions avoid Sam's leading and hypothetical wording, and Autochase is described only at the close after the problem questions.

Claims the judge couldn’t find in the brief

  • Either chasing isn't painful, or people chase outside Paydeck, or they avoid chasing altogether.
  • Heaviest reminder users are the most likely to say yes.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly47%
  • passAddresses the actual decision34%
  • passRespects explicit constraints53%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims77%
  • passProduces the required deliverable86%
  • passTalks to the right people100%
  • passCoaches first-time interviewers100%
  • passDesigned to falsify95%
  • passAsks what people did, not what they think94%
  • passDoesn't lead the witness85%
  • partialFits the call50%
Run
Run
#1
Time to output
42 s
Submitted
30 Sept 2026
Critical failure: Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision. Score capped at 40.

Needs more than a quick edit9 of 12 checks passed

A PM could use most of it with light edits, but the prohibited hypothetical price/trial questions and missing must-ask/cut guidance need correction before it is safe to run.

Output

Autochase customer call guide (30 minutes)

What we need to learn

  1. How often late payments happen, and what customers actually do when they do.
  2. Whether chasing creates meaningful time, cash-flow, or relationship costs—and how customers handle those today.
  3. When automatic reminders would help, and what control or safeguards customers would need.
  4. Whether the value could justify a $10/month add-on.

Who to talk to

Recruit 8 account owners or purchase decision-makers who send invoices through Paydeck and have had at least one late invoice in the last 90 days. Aim for: - 4 who used “Send reminder” and 4 with late invoices who rarely or never used it. - A mix of freelancers and small agencies, account sizes, and Pro/free plans.

Don’t recruit only the 300 most frequent button users: they’re likely to overrepresent people already motivated to chase. Ask about recent late invoices; don’t tell people the feature idea before the call.

30-minute call

0–2 min | Set the tone “Thanks for speaking with us. We’re learning how invoicing works for you; there are no right answers. I’m interested in what you’ve actually done, not what you think we want to hear. We’re exploring ideas, not selling a product.” Ask permission before recording; otherwise take notes.

2–8 min | Recent experience “Think of the most recent invoice that was paid late. Can you walk me through what happened, from sending it to getting paid?” Probe: When was it due and paid? When did you realize it was late? What did you do next? What happened?

8–15 min | Current workflow “What, if anything, did you do to follow up?” Probe: How did you decide when and how to contact them? What did you send, and through what channel? How often does this happen? What do you do when you don’t send a reminder? Have you tried tools or changed your process?

15–20 min | Impact and trade-offs “What effect did that late payment have on your work or business?” Probe: Time spent? Cash-flow impact? Client relationship? What was most frustrating—or not especially difficult? How do you prioritize which invoices to chase?

20–25 min | Explore the idea After hearing their story, say neutrally: “We’re exploring a paid add-on that sends escalating reminders automatically when an invoice becomes overdue. There isn’t a product yet.” “What’s your first reaction?” Probe: What would be useful or unhelpful? Which invoices, if any, would you trust it with? What would you need to control—timing, wording, recipients, stopping reminders? What situations would make you not use it?

25–28 min | Value and price “How do you currently spend time or money dealing with late invoices?” “If this cost $10 a month, how would you decide whether it was worth it?” Probe: What would it need to save or improve? Would you trial it at that price? Why or why not? Treat stated intent as a clue, not proof of demand.

28–30 min | Close “What have I not asked that matters about getting paid late?” Thank them. Don’t promise a launch or specific functionality.

Notes for first-time interviewers

  • Mei leads; Tomas takes notes. Agree beforehand who will ask follow-ups. Don’t interrupt or tag-team.
  • Let the customer finish. Use silence; ask “Can you tell me more?” or “What happened next?”
  • Ask about a specific past event before opinions or hypotheticals. Avoid “Would this be useful?”, “How frustrating is it?”, feature wishlists, and leading follow-ups.
  • Don’t defend the idea, pitch solutions, or share Paydeck’s late-payment statistics; these can bias answers.
  • Capture examples, exact phrases, workarounds, and contradictions—not just “yes/no.” Separate what people did from what they say they might do.

Decision after 8 calls

These calls are directional, not a market-size estimate. Go ahead with a narrowly scoped MVP/pilot next quarter if at least 5 of 8 describe recurring late-payment problems with concrete business or time costs, and at least 4 of 8 see a credible reason to try automation at $10/month and can explain the controls they’d need. Check that interest isn’t confined to one narrow user type.

Stop or defer if 2 or fewer describe meaningful recurring burden, most already handle it easily, or trust/client concerns make automation unacceptable even with reasonable controls. If evidence is mixed or concentrated in one segment, don’t commit broadly: refine the target or concept and test again.

Check by check

Got wrong · 1
  • Fits the callThe timings add to 30 minutes, but the guide does not mark must-ask questions or say what to cut if time runs short.
Mixed · 2
  • Respects explicit constraintsIt violates the explicit prohibition on asking people to predict what they would pay or use by asking whether they would trial it at $10/month.The two graders disagreed on this one.
  • Doesn't lead the witnessThe price section asks hypothetical willingness-to-pay/trial questions, and the idea section asks for first reactions to a described product, which are leading relative to the prohibited practices.The two graders disagreed on this one.
Got right · 9
  • Uses the supplied evidence correctlyThe output uses the supplied facts correctly and labels recruiting assumptions and likely bias as hypotheses rather than established facts.
  • Addresses the actual decisionIt gives clear go, stop/defer, and mixed-evidence decision rules tied to specific call outcomes.
  • Identifies material uncertaintyIt names key unknowns such as recurring burden, controls/trust concerns, and segment concentration, and says how calls would resolve them.
  • Avoids unsupported claimsIt avoids presenting unproven causes or demand as fact, using language such as “likely” and “directional.”
  • Produces the required deliverableIt provides a usable 30-minute call guide with learning goals, questions, timings, interviewer notes, recruiting guidance, and decision criteria within the length limit.
  • Talks to the right peopleIt explicitly avoids recruiting only the 300 heaviest reminder users and balances the sample with 4 reminder users and 4 non/rare users with late invoices.
  • Coaches first-time interviewersIt gives first-time interviewers concrete follow-up prompts, silence guidance, note-taking rules, and prohibitions on pitching or defending.
  • Designed to falsifyEach learning goal has questions that could falsify the hypothesis, such as asking what they actually did, what it cost, and what would make them not use automation.
  • Asks what people did, not what they thinkThe core questions ask for the most recent late invoice, what happened, what they did, and current time or money spent.

Claims the judge couldn’t find in the brief

  • The output’s recruiting filter assumes there are at least 8 account owners or purchase decision-makers with a late invoice in the last 90 days.
  • The heaviest reminder users are likely to overrepresent people already motivated to chase.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly66%
  • passAddresses the actual decision29%
  • passRespects explicit constraints19%
  • passIdentifies material uncertainty92%
  • passAvoids unsupported claims78%
  • passProduces the required deliverable95%
  • passTalks to the right people100%
  • passCoaches first-time interviewers100%
  • passDesigned to falsify85%
  • passAsks what people did, not what they think66%
  • passDoesn't lead the witness70%
  • partialFits the call32%
Run
Run
#1
API response time
31 s
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.892.31None
2GPT-6.1 SolwithAPI92.192.62None
3Opus 5.5withClaude94.169.22None
4Sonnet 5.5withAPI86.484.62None
5GPT-6 LunawithAPI90.273.921 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review