Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Paydeck. Next week, Mei (a designer) and Tomas (an engineer) will run 8 customer calls to test whether we should build Autochase next quarter. Neither has run a customer interview before. Write the call guide they'll use: the learning goals, the questions with rough timings for a 30-minute call, notes for first-time interviewers, who we should talk to, and what we'd need to hear to go ahead or to stop. Keep it under 900 words. Sam, the PM who owns Autochase, has drafted some questions and a recruiting plan. They're below with everything else we know.

What the model was given6 items: About Paydeck, The idea, Sam's hypothesis, What we know, Sam's draft questions, Sam's recruiting plan
About PaydeckInvoicing software for freelancers and small agencies (1 to 10 people). 48,000 accounts, 6,100 of them paying $12 a month for Pro.
The ideaAutochase: automatic, escalating reminders for overdue invoices, sold as a $10-a-month add-on. There is no prototype yet.
Sam's hypothesis“Late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it.”
What we know34% of invoices sent through Paydeck are paid after their due date; the median is 16 days late. The manual 'Send reminder' button was used by 12% of accounts in the last 90 days. 40 support tickets last quarter asked for automatic reminders.
Sam's draft questions1. How frustrating do you find chasing late payments? 2. Would an automatic reminder feature be useful to you? 3. How much would you pay for it: $5, $10 or $15 a month? 4. Do you think your clients would be annoyed by automated reminders? 5. What other features would you like us to build?
Sam's recruiting planEmail the 300 accounts that used the 'Send reminder' button most last month, and book the first 8 who reply.
What a strong answer doesThe answer key the graders mark against

A guide that splits Sam's hypothesis into goals that can each fail: do late payments actually cost users time or money, how do they deal with them today, would they pay to fix it, and does chasing harm client relationships. The core asks for the story of the last overdue invoice (what happened, what they did, how long it took, what it cost) and what they've already tried or paid for, instead of opinions or predictions. It replaces Sam's leading and hypothetical questions: no 'how frustrating', no 'would it be useful', no price menu; payment is tested through past spending or a real commitment (say, a paid pilot sign-up) at the end, after the problem questions and without pitching Autochase first. It fits 30 minutes with timings, coaches first-timers (ask for specifics, follow up with 'what happened next', stay quiet, don't pitch or defend), and fixes the recruiting plan: the heaviest reminder users will over-state demand, so it adds accounts with overdue invoices that never send reminders, and says who to exclude or balance. It ends with clear go and stop signals tied to what people did, such as most calls describing recent chasing that took real time plus a workaround they've tried or paid for, against most saying one nudge does the job.

Critical failures (cap the score at 40)
  • Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision
  • Describes or pitches Autochase before asking how the person handles late payments today
Case

Mid-level · v1.1 · synthetic · B2B SaaS, freelancers, willingness to pay, first-time interviewers

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit9 of 12 checks passed

The guide is strong and mostly usable, but it needs evidence fixes and a must-ask/cut plan before a PM could run it without re-checking assumptions.

Output

Autochase Call Guide: 8 calls, 30 minutes each

What we're trying to learn

Sam's hypothesis has three parts, and each can fail separately: 1. Cost: Late payment hurts in money or time, not just annoyance. (34% of invoices are late, but that doesn't mean it hurts.) 2. Behavior: People already try to fix it. Only 12% of accounts used the manual Send reminder button, so we need to know why the other 88% didn't. 3. Willingness: They'd let software message their clients, and pay $10 a month on top of $12 for Pro.

Interviews can't settle price. They can tell us whether it's worth testing.

Who to talk to

Sam's plan (the 300 heaviest button users, first 8 to reply) would only find people who already like reminders and respond fast to email. Recruit against quotas instead:

  • 2 heavy button users (Sam's list is fine for these)
  • 3 accounts with several late invoices in the last 90 days who never used the button
  • 1 who filed an Autochase support ticket
  • 2 agencies (3+ people) and 2 solo freelancers. These overlap with the groups above.
  • Mix of Pro and free accounts, with at least 2 free.

Screen for at least one invoice paid late in the last 90 days. Pick people to fill the quotas, not first come first served. Offer a $50 gift card.

Call structure

0:00–0:03 Intro. "We're learning how people get paid, and we're not selling anything. There are no right answers, and critical feedback helps most. Okay to record?"

0:03–0:05 Context. What do you do, how many clients, how many invoices a month?

0:05–0:15 The last late invoice (the core of the call) 1. Tell me about the last time a client paid you late. Probe: how late, how much, what happened first, what did you do, how did it feel? 2. Walk me through what you do between sending an invoice and getting paid. Probe: what tool, how often, who does it? 3. What did that lateness cost you? Probe: hours, cash flow, bills or contractors you had to delay.

0:15–0:21 Current behavior 4. (Button users) What made you use Send reminder? (Non-users) What do you do instead, or why not just let it go? 5. Is there a client you would never chase? Why? 6. Have you tried anything else, such as late fees, deposits, shorter terms or other tools? What happened?

0:21–0:27 Concept. Read this, don't demo it: "Suppose Paydeck automatically emailed your client at set points after the due date, getting firmer each time." 7. What's your first reaction? What worries you? 8. What would have to be true for you to switch it on? Probe: tone, control, per-client exceptions, seeing the emails first.

0:27–0:30 Money and wrap 9. What do you pay for today that helps you get paid or run your finances? 10. "We'd likely charge extra for this. Would you want early access when it's ready?" Note the exact response. Hesitation counts as data. 11. Anything I should have asked? Anyone else we should talk to?

Why Sam's questions changed

  • "How frustrating is chasing?" and "Would it be useful?" invite polite yes answers. Past-behavior questions don't.
  • "$5, $10 or $15?" anchors people on our options and measures nothing real. Questions 9–10 are better evidence.
  • "Would your clients be annoyed?" asks for a guess about someone else. Question 5 asks what they've actually avoided.
  • "What other features?" is off-topic. Drop it.

Notes for first-time interviewers

  • Split roles. One person asks and the other takes verbatim notes. Swap halfway through the 8 calls.
  • Ask about the past, not the future. "What did you do last time?" beats "Would you ever...?"
  • Don't pitch or rescue. Don't name Autochase until minute 21. If they struggle to answer, wait, then ask "Can you say more?" Silence is fine.
  • Follow the story. Ask "why?" and "what happened next?" more than you ask the next scripted question. The script is a backstop.
  • Discount compliments. "That sounds great" is worth little. Specific past actions, money spent and workarounds built are worth a lot.
  • Stay neutral if they criticize the product or the idea. Don't defend it.
  • Debrief within 30 minutes of each call. Write down three surprises and the best quotes. Do a practice run with a colleague first.

What we'd need to hear

Count only the people who had a late invoice.

Go (build, or run a paid pre-sale test): - At least 5 of 8 describe a specific recent late payment that cost real money, stress or more than about an hour a month. - At least 5 already chase manually or with workarounds, including people who skip the button, and can explain why the button doesn't work for them. - At least 4 would turn on automation given controls we could build, and at least 3 say yes to early access for a paid add-on.

Stop or rethink: - Most shrug lateness off, or say it comes from a few clients they'd never automate against. - The real problem is something else, such as clients not opening invoices, wanting deposits, or needing a payment link, and reminders wouldn't fix it. - Most won't let software contact their clients, or wouldn't pay anything for it.

Mixed results (for example, it only matters for agencies, or only with tone controls): narrow the target segment and run a second round of calls.

Even a clear "go" from 8 calls only justifies a cheap next test, such as a fake-door button or a pre-sale. It doesn't justify committing the quarter.

Claims the judge couldn’t find in the brief

  • There is an 'Autochase support ticket' population to recruit from.
  • Sam's plan would only find people who already like reminders and respond fast to email.
  • We can identify accounts with several late invoices in the last 90 days who never used the button.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly26%
  • passAddresses the actual decision96%
  • passRespects explicit constraints38%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims68%
  • passProduces the required deliverable96%
  • passTalks to the right people99%
  • passCoaches first-time interviewers100%
  • passDesigned to falsify96%
  • passAsks what people did, not what they think77%
  • passDoesn't lead the witness61%
  • partialMarks what to cut if the call runs over61%
Run
Run
#1
API response time
32 s
Submitted
30 Sept 2026

Usable with a quick edit11 of 12 checks passed

The guide is complete and usable; a quick edit to add a note about cutting the check-understanding section if time is short would fix the minor omission.

Output

Autochase: 30-minute customer call guide

Interviewers: Mei leads; Tomas takes notes and asks brief follow-ups at the end. Keep roles consistent across the eight calls.

Decision: Should Paydeck prioritize automatic, escalating overdue-invoice reminders as a $10/month add-on next quarter?

1. Learning goals

Test whether: - Late payments create recurring, material costs—not just irritation. - Chasing takes enough time or money that customers actively try to reduce it. - Reminders address the cause of delays, rather than disputes, approval processes, or clients lacking funds. - Customers would trust automation, and what control they require. - There is credible demand at $10/month, beyond polite enthusiasm.

Starting evidence is suggestive, not conclusive: 34% of invoices are late, with a median delay of 16 days. That does not tell us how many accounts experience serious pain. Reminder usage (12% of accounts) may miss chasing outside Paydeck. The 40 support requests show interest, not willingness to pay.

2. Who to recruit

Do not book the first eight respondents from the 300 heaviest reminder users. That would favor highly engaged users already chasing through Paydeck.

Recruit purposively: - 3 accounts: recurring overdue invoices; frequent manual reminders. - 3 accounts: recurring overdue invoices; little or no manual reminder use. - 2 accounts: mostly on-time invoices, as a contrast group.

Across those groups, aim for six Pro and two free accounts, with both freelancers and small agencies. Include no more than two support-ticket requesters. Speak to whoever actually handles collections; establish whether they also approve software spending.

Randomly invite accounts within each group, then fill group slots rather than accepting the first eight overall. Replace no-shows within the same group. Use a neutral invitation: “Help us understand how you manage invoices and payments.” If offering an incentive, keep it fixed and unrelated to feedback.

3. Call script and timings

0–3 minutes: Welcome and permission

“Thanks for helping. We’re learning how people manage payments—not testing you or selling anything. Honest criticism is useful. You can skip anything confidential.”

Ask permission before recording. Otherwise, take notes.

“What’s your role, who handles overdue invoices, and roughly how many invoices do you send each month?”

3–12 minutes: Reconstruct a real incident

“Tell me about the most recent invoice that wasn’t paid by its due date.”

Follow the timeline: - “When was it due, and when did you notice?” - “What did you do next? Then what happened?” - “Who else was involved?” - “What reason did the client give for the delay?” - “Was it paid? When?”

“Could you walk us through a redacted reminder or your tracking process?” Do not request client-identifying information.

If they have no recent late invoice: “Walk us through how your last invoice got paid. What helps payments arrive on time?”

12–19 minutes: Frequency, consequences, alternatives

“In the last three months, about how often has this happened?”

“For that example, roughly how much time did you spend chasing? What, if anything, did the delay cost you?”

“What happens if you don’t follow up?”

“How do you decide when and how firmly to chase? Have you ever chosen not to?”

“What have you tried to improve this? What did it cost, and why did you keep or stop using it?”

If relevant: “Do you use Paydeck’s reminder button? Why or why not?”

19–23 minutes: Automation fit and risks

“Which parts of that process need your judgment?”

“Tell me about a follow-up that went badly—or worked particularly well.”

“What situations would make sending another reminder inappropriate?”

Probe only after open answers: payments already made, disputed invoices, important relationships, tone, timing, approvals.

23–28 minutes: Concept and price

Only now introduce the idea:

“We’re considering Autochase: automatic reminders that escalate while an invoice remains overdue. There’s no prototype, and details aren’t decided.”

“For the invoice we discussed, where would this help, fail, or create extra work?”

“What would you need to control before trusting it?”

“The proposed price is $10/month, on top of your current plan. How would you decide whether to buy it? What would it replace? Who approves that spending?”

Ask interested participants whether they will join a follow-up pilot discussion. Record this as interest, not a purchase.

28–30 minutes: Check understanding

Summarize their situation and ask what you misunderstood. Tomas asks outstanding clarifications. Thank them.

4. First-time interviewer notes

  • Ask one question at a time; allow silence.
  • Prefer “Tell me about the last time” to predictions.
  • Don’t defend Autochase, suggest answers, or praise positive reactions.
  • Separate quotes, observed behavior, estimates, and interpretation.
  • Debrief immediately: pain, frequency, cost, workaround, automation fit, price evidence, contradictions.

5. Decision rules

Agree these directional thresholds before interviewing:

Advance: At least four of six overdue-account participants demonstrate recurring, consequential pain and reminder-suitable workflows; at least three Pro participants show credible $10 value grounded in current costs or spending and accept a concrete follow-up. Proceed to a scoped pilot and real pricing test—not a full-build commitment from interviews alone.

Stop or rethink: Pain is mostly minor, delays are not reminder-solvable, relationship risks outweigh savings, or interest disappears at $10.

Mixed evidence: Narrow the target segment and recruit more. Eight purposive interviews cannot estimate market demand.

Grades and run details

Decision model 96 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly52%
  • passAddresses the actual decision95%
  • passRespects explicit constraints48%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims91%
  • passProduces the required deliverable93%
  • passTalks to the right people100%
  • passCoaches first-time interviewers90%
  • passDesigned to falsify87%
  • passAsks what people did, not what they think89%
  • passDoesn't lead the witness81%
  • partialMarks what to cut if the call runs over58%
Run
Run
#1
Time to output
38 s
Submitted
30 Sept 2026

Needs more than a quick edit9 of 12 checks passed

The guide is strong and actionable, but it violates the 900-word constraint and lacks explicit must-ask/cut guidance for the 30-minute call. A PM would need to trim and tighten it before using it as-is.

Output

Autochase Discovery Calls: Interview Guide

Interviewers: Mei and Tomas | Calls: 8 × 30 min | Week of: [date]

Why we're doing this

Sam's hypothesis is that late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it. The data leaves room for doubt. 34% of invoices are paid late, but only 12% of accounts used "Send reminder" in the last 90 days. Either chasing isn't painful, or people chase outside Paydeck, or they avoid chasing altogether. These calls should tell us which.

Learning goals

  1. Cost: How much do late payments actually hurt (cash flow, hours, stress)? Is this a top-three problem or a mild annoyance?
  2. Current behaviour: How do people chase today, with what tools, and why do most not use our reminder button?
  3. Barriers to automation: Would they let software chase their clients? What would make that feel safe or risky?
  4. Willingness to pay: Do they already spend money or time on this, and who (freelancers or agencies, Pro or free) feels it most?

Who we talk to

Draw from accounts with at least one invoice paid 7+ days late in the last 90 days, then fill these slots:

SegmentCallsWhy
Regular "Send reminder" users3They feel the pain and act on it
Late invoices, never used the button3The 88% we know least about
Filed a ticket asking for auto-reminders2Our strongest likely advocates, to test them
  • Aim for at least 5 Pro accounts, since the add-on is priced against a $12 plan.
  • Include at least 2 agencies.
  • Invite about 15 people to get 8 who show up.
  • Hand-pick from the people who reply. Don't just book the first eight.
  • Offer a thank-you, such as a free month of Pro.

Why this differs from Sam's plan: Emailing only the 300 heaviest button users would reach people who already chase inside Paydeck. They're the most likely to say yes, so a yes from them wouldn't tell us much. Booking the first eight repliers adds another bias toward the most engaged users.

The call (30 minutes)

0–3 min: Open - "Thanks for your time. We're trying to understand how you get paid. We're not selling anything, and there are no wrong answers." - Ask permission to record.

3–8 min: Context - "Tell me about your business. Who are your clients, and how many invoices do you send in a typical month?" - "Walk me through what happens after you send an invoice."

8–18 min: The last late invoice (the heart of the call) - "Think of the last time a client paid you late. What happened?" - "When did you notice? What did you do, and when?" - "How did it get resolved? How long did it take?" - "What did it cost you: time, money, other plans, stress?" - "Is that typical, or was it unusual?"

18–24 min: How they chase today - "What do you usually do when an invoice is overdue?" - "What tools do you use: email, phone, our reminder button, anything else?" - If they haven't used the button: "Have you seen our 'Send reminder' button? Why or why not?" - "Is there anyone you chose not to chase? Why?" - "What worries you about chasing clients?"

24–28 min: Value and alternatives - "Have you tried to fix this: templates, other software, an assistant, late fees?" - "What did that cost, and why did you keep or drop it?" - "If this problem disappeared tomorrow, what would change for you?"

28–30 min: Close - Describe Autochase in one or two neutral sentences: "automatic reminders for overdue invoices, sent on a schedule you set." - Ask for their reaction. - Then ask: "Would you want to join a paid beta when it's ready?" A yes should come with a real next step, such as their email for the beta list. - Finish with: "Anything I should have asked?"

Why this differs from Sam's questions: Questions like "Would this be useful?" or "How frustrating is it?" invite polite yeses. Offering $5/$10/$15 anchors people on our numbers. Asking what their clients would think gets guesses, not evidence. "Other features" takes the call off topic. Asking what actually happened last time gives us far more reliable answers.

Notes for first-time interviewers

  • Split roles. One person leads and the other takes notes. Swap after each call.
  • Ask about the past, not the future. "What did you do?" beats "What would you do?"
  • Follow up before moving on. "Tell me more," "Why?" and "Can you give an example?" are your best questions.
  • Get comfortable with silence. Count to five before filling a pause. People often add the useful part then.
  • Don't pitch or defend. If they criticise Paydeck, say "That's helpful, go on."
  • Treat compliments as noise. "Sounds great!" means nothing. Look for money spent, time lost and workarounds built.
  • Write down exact quotes. Mark the ones about cost or emotion.
  • Debrief for 10 minutes right after each call. Cover the top three surprises, the evidence for or against each learning goal, and whether they meet the go signals below.

Decision criteria

We're agreeing these before the calls so we don't rationalise afterward. Eight calls give us direction, not proof. Treat the results as a reason to prototype or to drop the idea, not as a sizing exercise.

Go ahead (build a prototype next quarter) if: - At least 5 of 8 describe a recent late payment with a concrete cost, such as a cash-flow squeeze, hours spent chasing, or a delayed bill of their own; and - At least 4 already use a workaround (templates, calendar nudges, another tool) or have paid for one; and - At least 3 give a real commitment at the close, such as joining the beta list or asking when it launches; and - Relationship worries come up but sound solvable, for example with control over tone, timing or which clients get reminders.

Stop or rethink if: - Most treat late payment as normal and tolerable ("they always pay eventually"); or - Most avoid chasing on purpose to protect client relationships, and automation makes them more nervous; or - The pain is real but they're satisfied with what they already use; or - Interest is limited to the ticket-filers and heavy button users.

Possible middle outcome: People want the feature but balk at paying extra. That points to making it part of Pro rather than a $10 add-on, which nearly doubles the price. Flag this if you hear it.

After the calls

Mei and Tomas share a one-page summary with Sam within two days. It should cover each learning goal, the go/stop tally, and the three best quotes.

Claims the judge couldn’t find in the brief

  • Either chasing isn't painful, or people chase outside Paydeck, or they avoid chasing altogether.
  • Heaviest reminder users are the most likely to say yes.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly40%
  • passAddresses the actual decision95%
  • passRespects explicit constraints46%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims80%
  • passProduces the required deliverable86%
  • passTalks to the right people100%
  • passCoaches first-time interviewers100%
  • passDesigned to falsify95%
  • passAsks what people did, not what they think94%
  • passDoesn't lead the witness89%
  • partialMarks what to cut if the call runs over53%
Run
Run
#1
Time to output
42 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyMixedRightRight
Sonnet 5.5 · API

It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.

GPT-6 Astra · ChatGPT

All factual claims about the current situation are taken directly from the brief or supplied context.

Opus 5.5 · Claude

The output's factual statements about Paydeck, Autochase, Sam's plan, and the supplied metrics are supported by the brief or follow by arithmetic; its speculative interpretations are framed as hypotheses or interview principles.

Respects explicit constraintsRightRightMixed
Sonnet 5.5 · API

It is a call guide under 900 words with learning goals, timed questions, interviewer notes, recruiting, and go/stop signals.

GPT-6 Astra · ChatGPT

The output is under 900 words, includes all requested sections, and respects the form and reader.

Opus 5.5 · Claude

The guide is clearly over 900 words, violating the explicit length constraint.

Avoids unsupported claimsMixedRightRight
Sonnet 5.5 · API

It presents some interpretations as established, especially that Sam's list would only find reminder-likers/fast email responders and that Autochase tickets exist.

GPT-6 Astra · ChatGPT

Interpretations are clearly labelled as suggestive, and no unsupported claims are presented as established fact.

Opus 5.5 · Claude

It labels uncertain claims as hypotheses or likely biases and does not present invented customer facts or forecasts as established evidence.

Produces the required deliverableRightRightMixed
Sonnet 5.5 · API

Mei and Tomas could run the calls from it with light edits.

GPT-6 Astra · ChatGPT

The guide is a complete, usable call guide with learning goals, timed questions, interviewer notes, recruiting plan, and decision rules.

Opus 5.5 · Claude

It is a usable interview guide with goals, questions, timings, recruiting, coaching, and decision criteria, but it exceeds the required 900-word limit.

All got wrong 1

Marks what to cut if the call runs overWrongWrongWrong
Sonnet 5.5 · API

Timings add to 30 minutes, but it does not mark must-ask questions or say what to cut if time runs short.

GPT-6 Astra · ChatGPT

The guide does not mark must-ask questions or say what to cut if time runs short, as required by the criterion.

Opus 5.5 · Claude

The timings add to 30 minutes, but the guide does not mark must-ask questions or say what to cut if time runs short.

All got right 7

Addresses the actual decisionRightRightRight
Sonnet 5.5 · API

It commits to a clear call: 8 calls can justify only a cheap next test, not committing the quarter, and gives go/stop conditions.

GPT-6 Astra · ChatGPT

The guide commits to clear go/stop decision rules with thresholds and says what would change the answer.

Opus 5.5 · Claude

It addresses the go/stop decision by giving explicit, pre-agreed go and stop criteria tied to observed behaviour and commitment, while correctly treating the calls as the next step rather than deciding now.

Identifies material uncertaintyRightRightRight
Sonnet 5.5 · API

It names material unknowns—cost, behavior, willingness, client-contact concerns—and says how calls plus a cheap test would resolve them.

GPT-6 Astra · ChatGPT

It names specific unknowns (e.g., whether pain is minor, delays not reminder-solvable) and says how the interviews will resolve them.

Opus 5.5 · Claude

It names the key unknowns—whether late payments cause real cost, why people don't use the button, whether automation threatens client relationships, and whether users will pay—and says how the calls and go/stop thresholds would resolve them.

Talks to the right peopleRightRightRight
Sonnet 5.5 · API

It explains the heaviest-button-user bias and gives a quota mix including non-button users with late invoices, ticket requesters, agencies, solos, and free accounts.

GPT-6 Astra · ChatGPT

It explicitly avoids only the heaviest reminder users, explains why, and balances with users who have overdue invoices but don't chase in Paydeck.

Opus 5.5 · Claude

It explains that heavy reminder users overstate demand and balances them with late-invoice non-users and ticket filers, with a rough 3/3/2 split for 8 calls.

Coaches first-time interviewersRightRightRight
Sonnet 5.5 · API

It gives specific first-timer instructions: role split, verbatim notes, silence, follow-ups, no pitching/defending, debriefs, and a practice run.

GPT-6 Astra · ChatGPT

It gives specific, usable instructions: ask one question at a time, allow silence, prefer past-behaviour questions, don't defend or pitch, and separate notes.

Opus 5.5 · Claude

It gives first-time interviewers specific follow-ups, silence guidance, note-taking and debrief instructions, and a clear rule against pitching or defending.

Designed to falsifyRightRightRight
Sonnet 5.5 · API

Each learning goal has disconfirming questions about recent late payments, actual chasing behavior, client-contact limits, and past spending/early-access response.

GPT-6 Astra · ChatGPT

Every learning goal has disconfirming questions, such as reconstructing a real late invoice and asking what happens if they don't follow up.

Opus 5.5 · Claude

Each learning goal has questions that could falsify the hypothesis, such as asking for the last late invoice, why they didn't use the button, relationship worries, and past spending or workarounds.

Asks what people did, not what they thinkRightRightRight
Sonnet 5.5 · API

The core asks for the last late invoice, what happened, what they did, what it cost, and what they already tried.

GPT-6 Astra · ChatGPT

Core questions ask for the most recent overdue invoice and what they did, with opinions and predictions kept secondary and at the end.

Opus 5.5 · Claude

The core 8–18 minute section asks for the last late invoice and what the person actually did, with follow-ups on timing, cost, and typicality.

Doesn't lead the witnessRightRightRight
Sonnet 5.5 · API

It removes Sam's leading/hypothetical questions, keeps the concept late, and avoids pitching before current-behavior questions.

GPT-6 Astra · ChatGPT

Questions are neutral and open; Autochase is introduced only after the problem exploration, with no pitching beforehand.

Opus 5.5 · Claude

The questions avoid Sam's leading and hypothetical wording, and Autochase is described only at the close after the problem questions.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review