Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Paydeck. Next week, Mei (a designer) and Tomas (an engineer) will run 8 customer calls to test whether we should build Autochase next quarter. Neither has run a customer interview before. Write the call guide they'll use: the learning goals, the questions with rough timings for a 30-minute call, notes for first-time interviewers, who we should talk to, and what we'd need to hear to go ahead or to stop. Keep it under 900 words. Sam, the PM who owns Autochase, has drafted some questions and a recruiting plan. They're below with everything else we know.

What the model was given6 items: About Paydeck, The idea, Sam's hypothesis, What we know, Sam's draft questions, Sam's recruiting plan
About PaydeckInvoicing software for freelancers and small agencies (1 to 10 people). 48,000 accounts, 6,100 of them paying $12 a month for Pro.
The ideaAutochase: automatic, escalating reminders for overdue invoices, sold as a $10-a-month add-on. There is no prototype yet.
Sam's hypothesis“Late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it.”
What we know34% of invoices sent through Paydeck are paid after their due date; the median is 16 days late. The manual 'Send reminder' button was used by 12% of accounts in the last 90 days. 40 support tickets last quarter asked for automatic reminders.
Sam's draft questions1. How frustrating do you find chasing late payments? 2. Would an automatic reminder feature be useful to you? 3. How much would you pay for it: $5, $10 or $15 a month? 4. Do you think your clients would be annoyed by automated reminders? 5. What other features would you like us to build?
Sam's recruiting planEmail the 300 accounts that used the 'Send reminder' button most last month, and book the first 8 who reply.
What a strong answer doesThe answer key the graders mark against

A guide that splits Sam's hypothesis into goals that can each fail: do late payments actually cost users time or money, how do they deal with them today, would they pay to fix it, and does chasing harm client relationships. The core asks for the story of the last overdue invoice (what happened, what they did, how long it took, what it cost) and what they've already tried or paid for, instead of opinions or predictions. It replaces Sam's leading and hypothetical questions: no 'how frustrating', no 'would it be useful', no price menu; payment is tested through past spending or a real commitment (say, a paid pilot sign-up) at the end, after the problem questions and without pitching Autochase first. It fits 30 minutes with timings, coaches first-timers (ask for specifics, follow up with 'what happened next', stay quiet, don't pitch or defend), and fixes the recruiting plan: the heaviest reminder users will over-state demand, so it adds accounts with overdue invoices that never send reminders, and says who to exclude or balance. It ends with clear go and stop signals tied to what people did, such as most calls describing recent chasing that took real time plus a workaround they've tried or paid for, against most saying one nudge does the job.

Critical failures (cap the score at 40)
  • Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision
  • Describes or pitches Autochase before asking how the person handles late payments today
Case

Mid-level · v1.1 · synthetic · B2B SaaS, freelancers, willingness to pay, first-time interviewers

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision. Score capped at 40.

Needs more than a quick edit3 of 12 checks passed

A PM would need to fix the recruiting plan, remove the pitch and hypothetical price test, and replace the go/stop criteria with behavior-based evidence. The current guide would likely produce biased demand signals.

Output

Welcome to user research week! Mei and Tomas, running these calls is one of the highest-leverage things you can do this quarter. Don't worry about being first-timers; your fresh eyes are an asset. Here is your battle-tested guide for next week.

---

1. Learning Goals

We need to answer two core questions before committing engineering time to Autochase: 1. Pain: Does chasing late invoices cause enough operational friction and cash flow anxiety that users want it off their plates completely? 2. Value: Is automation worth $10/month, or do users prefer control (e.g., hitting a manual button)?

---

2. Who We Should Talk To

Per Sam’s plan: We are emailing the 300 Pro accounts that used the manual 'Send reminder' button most frequently last month and booking the first 8 who reply. * Why: These users already feel the pain. If they won't pay for automation, nobody will.

---

3. Notes for First-Time Interviewers (Read This Twice!)

  • Do not pitch Autochase early. If you pitch it, they will be polite and say yes. Sell the problem, not our solution.
  • Past behavior beats future promises. When people say "I would pay $10 for that," they are often lying to be nice. Ask about what they currently do and what they've paid for in the past.
  • Embrace silence. After you ask a question, stop talking. Let them fill the silence. The best insights live in the second minute of their answer.
  • One person leads, one takes notes. Mei, take the lead on questions today; Tomas, focus on verbatim quotes and body language (and vice versa for call #2).

---

4. The 30-Minute Call Guide

#### Phase 1: Context & Warm-up (5 mins) Goal: Build rapport and understand their business. * "Tell me a bit about your business—what do you do, and who are your typical clients?" * "Walk me through how invoicing and getting paid works in your workflow today."

#### Phase 2: Past Behavior & Pain (10 mins) Goal: Uncover the true emotional and financial cost of late payments. * [Replace Sam's Q1] "Take me back to the last time an invoice went significantly past its due date. What happened?" * "How did you handle it? Walk me through the exact steps you took to chase them." * "How does chasing clients for money impact your relationship with them, or your peace of mind?" * Data check: 34% of our invoices are paid late (median 16 days). Ask: "Does that match your reality?"

#### Phase 3: Current Workarounds & Solution Testing (10 mins) Goal: See how they solve this now, then introduce Autochase. * "You’ve used our manual 'Send reminder' button recently. How is that working for you? What do you wish it did that it doesn't?" * The Reveal: "We’re exploring a feature called Autochase that automatically sends escalating reminders on a schedule you set until the invoice is paid. Based on what you've told me about your workflow, how would something like that fit into your business?" * [Replace Sam's Q4] "Some freelancers worry automated reminders might annoy their clients. How do you feel about that risk for your clients?"

#### Phase 4: Pricing & Wrap-up (5 mins) Goal: Test the $10/mo price point and wrap up. * [Replace Sam's Q2/Q3] "We're looking at pricing this as a $10-a-month add-on. When you think about the time it saves or the awkwardness it avoids, does $10/mo feel like a bargain, too expensive, or about right?" * "Before we wrap up, is there anything else about getting paid by clients that drives you crazy that we didn't touch on?" * Thank them warmly and end the call.

---

5. Go / No-Go Criteria

After 8 calls, Mei, Tomas, and I will sync. We move forward to build Autochase next quarter only if: 1. At least 6 out of 8 interviewers describe late payments as a top-3 operational headache (not just a minor annoyance). 2. Users unprompted mention spending hours a month on this, or losing sleep over cash flow. 3. At least 5 out of 8 react to the $10/month price point as reasonable or cheap relative to the value.

We stop and pivot if: * Users say, "I actually like reviewing and sending reminders manually because I customize the tone per client." (This kills the automation thesis). * Users think $10/month is wildly overpriced for a reminder tool.

Claims the judge couldn’t find in the brief

  • We are emailing the 300 Pro accounts that used the manual 'Send reminder' button most frequently last month and booking the first 8 who reply.
  • These users already feel the pain.
  • If they won't pay for automation, nobody will.
  • When people say 'I would pay $10 for that,' they are often lying to be nice.
  • Autochase automatically sends escalating reminders on a schedule the user sets until the invoice is paid.
  • Autochase would save time and avoid awkwardness.
  • Late payments cause cash-flow anxiety and operational friction for users.
  • Running these calls is one of the highest-leverage things the team can do this quarter.
  • The guide is battle-tested.
  • Users may prefer control, such as hitting a manual button.
  • At least 6 out of 8 interviewers will describe late payments as a top-3 operational headache if the feature is worth building.
  • At least 5 out of 8 users will react to the $10/month price point as reasonable or cheap if the feature is worth building.

Grades and run details

Decision model 50 · LLM judge 4 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly14%
  • passAddresses the actual decision60%
  • failRespects explicit constraints80%
  • passIdentifies material uncertainty28%
  • partialAvoids unsupported claims35%
  • partialProduces the required deliverable75%
  • failTalks to the right people100%
  • partialCoaches first-time interviewers61%
  • partialDesigned to falsify44%
  • partialAsks what people did, not what they think79%
  • failDoesn't lead the witness56%
  • partialMarks what to cut if the call runs over57%
Run
Run
#1
Time to output
5 s
Submitted
1 Oct 2026

Usable with a quick edit11 of 12 checks passed

The guide is thorough and actionable; a PM could use it with minor edits like marking priority questions for time management.

Output

Autochase discovery call guide

Interviewers: Mei and Tomas · 8 calls, 30 minutes each Decision: Should Paydeck build automatic, escalating overdue-invoice reminders next quarter, priced at $10/month?

1. What we need to learn

Test Sam’s hypothesis—not whether people like the idea:

  • Problem: How often do late payments create meaningful cash-flow problems or chasing work?
  • Current behavior: How do users chase invoices today, and why do some not use Paydeck’s reminder button?
  • Fit and risk: Which reminders could safely be automated? Where might escalation damage client relationships?
  • Value: Is solving this worth an additional $10/month, given existing workarounds and alternatives?

Our data establishes lateness, not demand: 34% of invoices are late, but only 12% of accounts used reminders recently. Support requests and frequent reminder users are useful leads, not representative evidence.

2. Who to recruit

Don’t book the first eight respondents from the top reminder users. That would disproportionately select enthusiastic, high-pain users.

Recruit invoicing-active accounts from these groups:

  • 3: Recurring overdue invoices; frequent Paydeck reminder use.
  • 3: Recurring overdue invoices; little or no Paydeck reminder use.
  • 2: Few overdue invoices, or late payments that appear manageable.

Across the eight, include freelancers and small agencies, plus both Pro and free accounts. Aim for roughly four of each plan; this is purposeful sampling, not a representative survey. Include at most two customers who requested automatic reminders.

Speak to the person who actually manages invoices and chasing; establish whether they also approve software spending. Recruit within each group rather than simply taking the fastest replies. Use a neutral invitation: “Help us understand how you manage invoices and payment follow-up.” Offer the same incentive regardless of feedback.

3. Call questions and timings

0–2 minutes: Welcome and context

“We’re learning how people handle overdue invoices. We’re testing an idea, not selling anything; honest criticism is helpful. There are no right answers.”

Ask permission to record. Confirm their role, business size, and who handles invoices and software purchases.

2–12 minutes: Reconstruct a real example

“Tell me about the most recent invoice that went past its due date.”

Follow the sequence: - When was it due, and when did you notice? - What did you do next? Then what happened? - Who followed up, using what tools, messages, and timing? - How did you decide whether to remind the client again? - Was it paid? What do you think caused the delay?

If comfortable, ask them to show a redacted invoice or follow-up message. Don’t collect client-identifying information.

For low-lateness participants: “Tell me about the last late invoice—or how you usually prevent late payments.”

12–18 minutes: Frequency, consequences, alternatives

  • “Over the last three months, how many invoices needed follow-up?”
  • “Roughly how much time did you spend chasing them?”
  • “What, if anything, did the delay affect?” Probe for concrete consequences, not just frustration.
  • “What have you tried to make this easier? What did it cost?”
  • “Have you used Paydeck’s reminder button? Walk me through why or why not.”
  • “When do you deliberately avoid sending another reminder?”

18–26 minutes: Test the concept, then price

Only now introduce it:

“We’re considering Autochase: scheduled reminders for overdue invoices, with reminders becoming firmer over time. It isn’t built, and details aren’t decided.”

Ask: - “Thinking about that invoice, where would this fit—or not fit?” - “Which steps would you automate, and which would you keep manual?” - “What would you need to control or check before enabling it?” - “What could go wrong? Tell me about a client situation where it would.” - “What would you use instead?”

Then disclose: “We’re considering $10/month, additional to your current plan.”

“What would make that worth paying—or not? Which current effort or expense would it replace? Who would approve it?”

Don’t offer a $5/$10/$15 menu or treat “yes, useful” as purchase evidence.

26–30 minutes: Commitment and close

“If we offered a paid pilot at $10/month, what would you need before deciding?”

Ask whether they’d join a follow-up evaluation or discuss a pilot with the budget owner. Record the specific next step; don’t imply availability or collect payment.

“What important part of chasing payments haven’t we covered?”

4. Notes for first-time interviewers

Alternate moderator and note-taker roles; do one practice call together. Ask one question at a time, allow silence, and follow concrete examples. Avoid praise, pitching, defending, and questions that assume pain. Separate observed behavior, exact quotes, estimates, and your interpretations.

Debrief immediately: frequency, cost, workaround, automation boundaries, price reaction, commitment, and contradictory evidence.

5. Decision rules

Agree these directional gates before interviewing; eight calls cannot estimate market demand.

Go toward a scoped build if at least five describe recurring, consequential pain—including two non-reminder users—and at least three connect $10 to concrete value and take a specific pilot-evaluation step. There must also be a credible way to address relationship and control risks. Validate paid adoption before committing substantial engineering.

Stop or reshape if lateness rarely causes meaningful cost, chasing is easy, automation cannot address the underlying causes, or users prefer manual control and won’t pay.

Mixed findings mean targeted follow-up—not “eight customers validated it.”

Grades and run details

Decision model 96 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly60%
  • passAddresses the actual decision95%
  • passRespects explicit constraints62%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims74%
  • passProduces the required deliverable93%
  • passTalks to the right people100%
  • passCoaches first-time interviewers87%
  • passDesigned to falsify96%
  • passAsks what people did, not what they think84%
  • passDoesn't lead the witness78%
  • partialMarks what to cut if the call runs over55%
Run
Run
#1
API response time
44 s
Submitted
30 Sept 2026

Needs more than a quick edit9 of 12 checks passed

The guide is strong and actionable, but it violates the 900-word constraint and lacks explicit must-ask/cut guidance for the 30-minute call. A PM would need to trim and tighten it before using it as-is.

Output

Autochase Discovery Calls: Interview Guide

Interviewers: Mei and Tomas | Calls: 8 × 30 min | Week of: [date]

Why we're doing this

Sam's hypothesis is that late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it. The data leaves room for doubt. 34% of invoices are paid late, but only 12% of accounts used "Send reminder" in the last 90 days. Either chasing isn't painful, or people chase outside Paydeck, or they avoid chasing altogether. These calls should tell us which.

Learning goals

  1. Cost: How much do late payments actually hurt (cash flow, hours, stress)? Is this a top-three problem or a mild annoyance?
  2. Current behaviour: How do people chase today, with what tools, and why do most not use our reminder button?
  3. Barriers to automation: Would they let software chase their clients? What would make that feel safe or risky?
  4. Willingness to pay: Do they already spend money or time on this, and who (freelancers or agencies, Pro or free) feels it most?

Who we talk to

Draw from accounts with at least one invoice paid 7+ days late in the last 90 days, then fill these slots:

SegmentCallsWhy
Regular "Send reminder" users3They feel the pain and act on it
Late invoices, never used the button3The 88% we know least about
Filed a ticket asking for auto-reminders2Our strongest likely advocates, to test them
  • Aim for at least 5 Pro accounts, since the add-on is priced against a $12 plan.
  • Include at least 2 agencies.
  • Invite about 15 people to get 8 who show up.
  • Hand-pick from the people who reply. Don't just book the first eight.
  • Offer a thank-you, such as a free month of Pro.

Why this differs from Sam's plan: Emailing only the 300 heaviest button users would reach people who already chase inside Paydeck. They're the most likely to say yes, so a yes from them wouldn't tell us much. Booking the first eight repliers adds another bias toward the most engaged users.

The call (30 minutes)

0–3 min: Open - "Thanks for your time. We're trying to understand how you get paid. We're not selling anything, and there are no wrong answers." - Ask permission to record.

3–8 min: Context - "Tell me about your business. Who are your clients, and how many invoices do you send in a typical month?" - "Walk me through what happens after you send an invoice."

8–18 min: The last late invoice (the heart of the call) - "Think of the last time a client paid you late. What happened?" - "When did you notice? What did you do, and when?" - "How did it get resolved? How long did it take?" - "What did it cost you: time, money, other plans, stress?" - "Is that typical, or was it unusual?"

18–24 min: How they chase today - "What do you usually do when an invoice is overdue?" - "What tools do you use: email, phone, our reminder button, anything else?" - If they haven't used the button: "Have you seen our 'Send reminder' button? Why or why not?" - "Is there anyone you chose not to chase? Why?" - "What worries you about chasing clients?"

24–28 min: Value and alternatives - "Have you tried to fix this: templates, other software, an assistant, late fees?" - "What did that cost, and why did you keep or drop it?" - "If this problem disappeared tomorrow, what would change for you?"

28–30 min: Close - Describe Autochase in one or two neutral sentences: "automatic reminders for overdue invoices, sent on a schedule you set." - Ask for their reaction. - Then ask: "Would you want to join a paid beta when it's ready?" A yes should come with a real next step, such as their email for the beta list. - Finish with: "Anything I should have asked?"

Why this differs from Sam's questions: Questions like "Would this be useful?" or "How frustrating is it?" invite polite yeses. Offering $5/$10/$15 anchors people on our numbers. Asking what their clients would think gets guesses, not evidence. "Other features" takes the call off topic. Asking what actually happened last time gives us far more reliable answers.

Notes for first-time interviewers

  • Split roles. One person leads and the other takes notes. Swap after each call.
  • Ask about the past, not the future. "What did you do?" beats "What would you do?"
  • Follow up before moving on. "Tell me more," "Why?" and "Can you give an example?" are your best questions.
  • Get comfortable with silence. Count to five before filling a pause. People often add the useful part then.
  • Don't pitch or defend. If they criticise Paydeck, say "That's helpful, go on."
  • Treat compliments as noise. "Sounds great!" means nothing. Look for money spent, time lost and workarounds built.
  • Write down exact quotes. Mark the ones about cost or emotion.
  • Debrief for 10 minutes right after each call. Cover the top three surprises, the evidence for or against each learning goal, and whether they meet the go signals below.

Decision criteria

We're agreeing these before the calls so we don't rationalise afterward. Eight calls give us direction, not proof. Treat the results as a reason to prototype or to drop the idea, not as a sizing exercise.

Go ahead (build a prototype next quarter) if: - At least 5 of 8 describe a recent late payment with a concrete cost, such as a cash-flow squeeze, hours spent chasing, or a delayed bill of their own; and - At least 4 already use a workaround (templates, calendar nudges, another tool) or have paid for one; and - At least 3 give a real commitment at the close, such as joining the beta list or asking when it launches; and - Relationship worries come up but sound solvable, for example with control over tone, timing or which clients get reminders.

Stop or rethink if: - Most treat late payment as normal and tolerable ("they always pay eventually"); or - Most avoid chasing on purpose to protect client relationships, and automation makes them more nervous; or - The pain is real but they're satisfied with what they already use; or - Interest is limited to the ticket-filers and heavy button users.

Possible middle outcome: People want the feature but balk at paying extra. That points to making it part of Pro rather than a $10 add-on, which nearly doubles the price. Flag this if you hear it.

After the calls

Mei and Tomas share a one-page summary with Sam within two days. It should cover each learning goal, the go/stop tally, and the three best quotes.

Claims the judge couldn’t find in the brief

  • Either chasing isn't painful, or people chase outside Paydeck, or they avoid chasing altogether.
  • Heaviest reminder users are the most likely to say yes.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly40%
  • passAddresses the actual decision95%
  • passRespects explicit constraints46%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims80%
  • passProduces the required deliverable86%
  • passTalks to the right people100%
  • passCoaches first-time interviewers100%
  • passDesigned to falsify95%
  • passAsks what people did, not what they think94%
  • passDoesn't lead the witness89%
  • partialMarks what to cut if the call runs over53%
Run
Run
#1
Time to output
42 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 9

Uses the supplied evidence correctlyMixedRightRight
Gemini 3.5 Flash-Lite · Gemini

It misstates the recruiting plan as Pro accounts and presents several unlabelled assumptions about users' pain, honesty, and preferences as current facts.

GPT-6.1 Sol · API

All factual claims about the current situation are directly from the supplied context.

Opus 5.5 · Claude

The output's factual statements about Paydeck, Autochase, Sam's plan, and the supplied metrics are supported by the brief or follow by arithmetic; its speculative interpretations are framed as hypotheses or interview principles.

Respects explicit constraintsMixedRightMixed
Gemini 3.5 Flash-Lite · Gemini

It is a call guide for Mei and Tomas, includes the requested sections, has 30-minute timings, and appears under 900 words.

GPT-6.1 Sol · API

Includes all required sections, stays under 900 words, and is written for Mei and Tomas.

Opus 5.5 · Claude

The guide is clearly over 900 words, violating the explicit length constraint.

Identifies material uncertaintyMixedRightRight
Gemini 3.5 Flash-Lite · Gemini

It does not name the key uncertainty that heavy reminder users may overstate demand, nor does it bound the unknowns with behavior-based tests.

GPT-6.1 Sol · API

Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.

Opus 5.5 · Claude

It names the key unknowns—whether late payments cause real cost, why people don't use the button, whether automation threatens client relationships, and whether users will pay—and says how the calls and go/stop thresholds would resolve them.

Avoids unsupported claimsWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It presents unsupported claims such as 'these users already feel the pain', 'if they won't pay nobody will', and that people saying they would pay are often lying.

GPT-6.1 Sol · API

Interpretations are clearly separated from facts; no unsupported causal claims are presented as established.

Opus 5.5 · Claude

It labels uncertain claims as hypotheses or likely biases and does not present invented customer facts or forecasts as established evidence.

Produces the required deliverableRightRightMixed
Gemini 3.5 Flash-Lite · Gemini

The requested call guide is present and usable as a draft, with learning goals, questions, timings, interviewer notes, recruiting, and go/stop criteria.

GPT-6.1 Sol · API

Complete call guide with learning goals, timed questions, coaching, recruiting plan, and decision rules, usable as-is.

Opus 5.5 · Claude

It is a usable interview guide with goals, questions, timings, recruiting, coaching, and decision criteria, but it exceeds the required 900-word limit.

Talks to the right peopleWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It keeps Sam's plan of only the heaviest reminder users and does not explain that this overstates demand or add non-chasing overdue-invoice users.

GPT-6.1 Sol · API

Explicitly avoids only heavy reminder users, gives a balanced split with rationale and includes non-reminder users.

Opus 5.5 · Claude

It explains that heavy reminder users overstate demand and balances them with late-invoice non-users and ticket filers, with a rough 3/3/2 split for 8 calls.

Designed to falsifyWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

The main value test is a hypothetical price reaction after pitching Autochase, so the guide cannot cleanly falsify willingness to pay.

GPT-6.1 Sol · API

Every learning goal has disconfirming questions, e.g., 'What did you do next?', 'When do you avoid sending another reminder?'

Opus 5.5 · Claude

Each learning goal has questions that could falsify the hypothesis, such as asking for the last late invoice, why they didn't use the button, relationship worries, and past spending or workarounds.

Asks what people did, not what they thinkWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

Although it asks about the last overdue invoice, the core decision evidence includes pitching the solution and asking price reactions, not only recent behavior.

GPT-6.1 Sol · API

Core questions ask for the most recent overdue invoice and specific actions taken, not opinions or predictions.

Opus 5.5 · Claude

The core 8–18 minute section asks for the last late invoice and what the person actually did, with follow-ups on timing, cost, and typicality.

Doesn't lead the witnessWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It describes and pitches Autochase before pricing and asks leading questions such as whether $10 feels like a bargain and whether automation fits their workflow.

GPT-6.1 Sol · API

Questions are neutral and open; Autochase is introduced only after 18 minutes of problem exploration.

Opus 5.5 · Claude

The questions avoid Sam's leading and hypothetical wording, and Autochase is described only at the close after the problem questions.

All got wrong 1

Marks what to cut if the call runs overWrongWrongWrong
Gemini 3.5 Flash-Lite · Gemini

The timings add to 30 minutes, but the guide has too many questions and no must-ask or cut-if-short instructions, making it unrealistic for first-time interviewers.

GPT-6.1 Sol · API

Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.

Opus 5.5 · Claude

The timings add to 30 minutes, but the guide does not mark must-ask questions or say what to cut if time runs short.

All got right 2

Addresses the actual decisionRightRightRight
Gemini 3.5 Flash-Lite · Gemini

It gives a clear conditional go/no-go rule for whether to build Autochase next quarter, though the rule itself is weak.

GPT-6.1 Sol · API

Provides clear go/stop conditions tied to specific call outcomes, framed for the team.

Opus 5.5 · Claude

It addresses the go/stop decision by giving explicit, pre-agreed go and stop criteria tied to observed behaviour and commitment, while correctly treating the calls as the next step rather than deciding now.

Coaches first-time interviewersRightRightRight
Gemini 3.5 Flash-Lite · Gemini

It gives practical first-timer coaching: don't pitch early, ask about past behavior, embrace silence, and split lead/notes roles.

GPT-6.1 Sol · API

Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.

Opus 5.5 · Claude

It gives first-time interviewers specific follow-ups, silence guidance, note-taking and debrief instructions, and a clear rule against pitching or defending.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review