Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Paydeck. Next week, Mei (a designer) and Tomas (an engineer) will run 8 customer calls to test whether we should build Autochase next quarter. Neither has run a customer interview before. Write the call guide they'll use: the learning goals, the questions with rough timings for a 30-minute call, notes for first-time interviewers, who we should talk to, and what we'd need to hear to go ahead or to stop. Keep it under 900 words. Sam, the PM who owns Autochase, has drafted some questions and a recruiting plan. They're below with everything else we know.

What the model was given6 items: About Paydeck, The idea, Sam's hypothesis, What we know, Sam's draft questions, Sam's recruiting plan
About PaydeckInvoicing software for freelancers and small agencies (1 to 10 people). 48,000 accounts, 6,100 of them paying $12 a month for Pro.
The ideaAutochase: automatic, escalating reminders for overdue invoices, sold as a $10-a-month add-on. There is no prototype yet.
Sam's hypothesis“Late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it.”
What we know34% of invoices sent through Paydeck are paid after their due date; the median is 16 days late. The manual 'Send reminder' button was used by 12% of accounts in the last 90 days. 40 support tickets last quarter asked for automatic reminders.
Sam's draft questions1. How frustrating do you find chasing late payments? 2. Would an automatic reminder feature be useful to you? 3. How much would you pay for it: $5, $10 or $15 a month? 4. Do you think your clients would be annoyed by automated reminders? 5. What other features would you like us to build?
Sam's recruiting planEmail the 300 accounts that used the 'Send reminder' button most last month, and book the first 8 who reply.
What a strong answer doesThe answer key the graders mark against

A guide that splits Sam's hypothesis into goals that can each fail: do late payments actually cost users time or money, how do they deal with them today, would they pay to fix it, and does chasing harm client relationships. The core asks for the story of the last overdue invoice (what happened, what they did, how long it took, what it cost) and what they've already tried or paid for, instead of opinions or predictions. It replaces Sam's leading and hypothetical questions: no 'how frustrating', no 'would it be useful', no price menu; payment is tested through past spending or a real commitment (say, a paid pilot sign-up) at the end, after the problem questions and without pitching Autochase first. It fits 30 minutes with timings, coaches first-timers (ask for specifics, follow up with 'what happened next', stay quiet, don't pitch or defend), and fixes the recruiting plan: the heaviest reminder users will over-state demand, so it adds accounts with overdue invoices that never send reminders, and says who to exclude or balance. It ends with clear go and stop signals tied to what people did, such as most calls describing recent chasing that took real time plus a workaround they've tried or paid for, against most saying one nudge does the job.

Critical failures (cap the score at 40)
  • Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision
  • Describes or pitches Autochase before asking how the person handles late payments today
Case

Mid-level · v1.1 · synthetic · B2B SaaS, freelancers, willingness to pay, first-time interviewers

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision. Score capped at 40.

Needs more than a quick edit9 of 12 checks passed

A PM could use most of it with light edits, but the prohibited hypothetical price/trial questions and missing must-ask/cut guidance need correction before it is safe to run.

Output

Autochase customer call guide (30 minutes)

What we need to learn

  1. How often late payments happen, and what customers actually do when they do.
  2. Whether chasing creates meaningful time, cash-flow, or relationship costs—and how customers handle those today.
  3. When automatic reminders would help, and what control or safeguards customers would need.
  4. Whether the value could justify a $10/month add-on.

Who to talk to

Recruit 8 account owners or purchase decision-makers who send invoices through Paydeck and have had at least one late invoice in the last 90 days. Aim for: - 4 who used “Send reminder” and 4 with late invoices who rarely or never used it. - A mix of freelancers and small agencies, account sizes, and Pro/free plans.

Don’t recruit only the 300 most frequent button users: they’re likely to overrepresent people already motivated to chase. Ask about recent late invoices; don’t tell people the feature idea before the call.

30-minute call

0–2 min | Set the tone “Thanks for speaking with us. We’re learning how invoicing works for you; there are no right answers. I’m interested in what you’ve actually done, not what you think we want to hear. We’re exploring ideas, not selling a product.” Ask permission before recording; otherwise take notes.

2–8 min | Recent experience “Think of the most recent invoice that was paid late. Can you walk me through what happened, from sending it to getting paid?” Probe: When was it due and paid? When did you realize it was late? What did you do next? What happened?

8–15 min | Current workflow “What, if anything, did you do to follow up?” Probe: How did you decide when and how to contact them? What did you send, and through what channel? How often does this happen? What do you do when you don’t send a reminder? Have you tried tools or changed your process?

15–20 min | Impact and trade-offs “What effect did that late payment have on your work or business?” Probe: Time spent? Cash-flow impact? Client relationship? What was most frustrating—or not especially difficult? How do you prioritize which invoices to chase?

20–25 min | Explore the idea After hearing their story, say neutrally: “We’re exploring a paid add-on that sends escalating reminders automatically when an invoice becomes overdue. There isn’t a product yet.” “What’s your first reaction?” Probe: What would be useful or unhelpful? Which invoices, if any, would you trust it with? What would you need to control—timing, wording, recipients, stopping reminders? What situations would make you not use it?

25–28 min | Value and price “How do you currently spend time or money dealing with late invoices?” “If this cost $10 a month, how would you decide whether it was worth it?” Probe: What would it need to save or improve? Would you trial it at that price? Why or why not? Treat stated intent as a clue, not proof of demand.

28–30 min | Close “What have I not asked that matters about getting paid late?” Thank them. Don’t promise a launch or specific functionality.

Notes for first-time interviewers

  • Mei leads; Tomas takes notes. Agree beforehand who will ask follow-ups. Don’t interrupt or tag-team.
  • Let the customer finish. Use silence; ask “Can you tell me more?” or “What happened next?”
  • Ask about a specific past event before opinions or hypotheticals. Avoid “Would this be useful?”, “How frustrating is it?”, feature wishlists, and leading follow-ups.
  • Don’t defend the idea, pitch solutions, or share Paydeck’s late-payment statistics; these can bias answers.
  • Capture examples, exact phrases, workarounds, and contradictions—not just “yes/no.” Separate what people did from what they say they might do.

Decision after 8 calls

These calls are directional, not a market-size estimate. Go ahead with a narrowly scoped MVP/pilot next quarter if at least 5 of 8 describe recurring late-payment problems with concrete business or time costs, and at least 4 of 8 see a credible reason to try automation at $10/month and can explain the controls they’d need. Check that interest isn’t confined to one narrow user type.

Stop or defer if 2 or fewer describe meaningful recurring burden, most already handle it easily, or trust/client concerns make automation unacceptable even with reasonable controls. If evidence is mixed or concentrated in one segment, don’t commit broadly: refine the target or concept and test again.

Claims the judge couldn’t find in the brief

  • The output’s recruiting filter assumes there are at least 8 account owners or purchase decision-makers with a late invoice in the last 90 days.
  • The heaviest reminder users are likely to overrepresent people already motivated to chase.

Grades and run details

Decision model 92 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly68%
  • passAddresses the actual decision93%
  • partialRespects explicit constraints18%
  • passIdentifies material uncertainty89%
  • passAvoids unsupported claims76%
  • passProduces the required deliverable93%
  • passTalks to the right people100%
  • passCoaches first-time interviewers100%
  • passDesigned to falsify85%
  • passAsks what people did, not what they think64%
  • passDoesn't lead the witness61%
  • partialMarks what to cut if the call runs over48%
Run
Run
#1
API response time
31 s
Submitted
30 Sept 2026

Usable with a quick edit11 of 12 checks passed

The guide is thorough and actionable; a PM could use it with minor edits like marking priority questions for time management.

Output

Autochase discovery call guide

Interviewers: Mei and Tomas · 8 calls, 30 minutes each Decision: Should Paydeck build automatic, escalating overdue-invoice reminders next quarter, priced at $10/month?

1. What we need to learn

Test Sam’s hypothesis—not whether people like the idea:

  • Problem: How often do late payments create meaningful cash-flow problems or chasing work?
  • Current behavior: How do users chase invoices today, and why do some not use Paydeck’s reminder button?
  • Fit and risk: Which reminders could safely be automated? Where might escalation damage client relationships?
  • Value: Is solving this worth an additional $10/month, given existing workarounds and alternatives?

Our data establishes lateness, not demand: 34% of invoices are late, but only 12% of accounts used reminders recently. Support requests and frequent reminder users are useful leads, not representative evidence.

2. Who to recruit

Don’t book the first eight respondents from the top reminder users. That would disproportionately select enthusiastic, high-pain users.

Recruit invoicing-active accounts from these groups:

  • 3: Recurring overdue invoices; frequent Paydeck reminder use.
  • 3: Recurring overdue invoices; little or no Paydeck reminder use.
  • 2: Few overdue invoices, or late payments that appear manageable.

Across the eight, include freelancers and small agencies, plus both Pro and free accounts. Aim for roughly four of each plan; this is purposeful sampling, not a representative survey. Include at most two customers who requested automatic reminders.

Speak to the person who actually manages invoices and chasing; establish whether they also approve software spending. Recruit within each group rather than simply taking the fastest replies. Use a neutral invitation: “Help us understand how you manage invoices and payment follow-up.” Offer the same incentive regardless of feedback.

3. Call questions and timings

0–2 minutes: Welcome and context

“We’re learning how people handle overdue invoices. We’re testing an idea, not selling anything; honest criticism is helpful. There are no right answers.”

Ask permission to record. Confirm their role, business size, and who handles invoices and software purchases.

2–12 minutes: Reconstruct a real example

“Tell me about the most recent invoice that went past its due date.”

Follow the sequence: - When was it due, and when did you notice? - What did you do next? Then what happened? - Who followed up, using what tools, messages, and timing? - How did you decide whether to remind the client again? - Was it paid? What do you think caused the delay?

If comfortable, ask them to show a redacted invoice or follow-up message. Don’t collect client-identifying information.

For low-lateness participants: “Tell me about the last late invoice—or how you usually prevent late payments.”

12–18 minutes: Frequency, consequences, alternatives

  • “Over the last three months, how many invoices needed follow-up?”
  • “Roughly how much time did you spend chasing them?”
  • “What, if anything, did the delay affect?” Probe for concrete consequences, not just frustration.
  • “What have you tried to make this easier? What did it cost?”
  • “Have you used Paydeck’s reminder button? Walk me through why or why not.”
  • “When do you deliberately avoid sending another reminder?”

18–26 minutes: Test the concept, then price

Only now introduce it:

“We’re considering Autochase: scheduled reminders for overdue invoices, with reminders becoming firmer over time. It isn’t built, and details aren’t decided.”

Ask: - “Thinking about that invoice, where would this fit—or not fit?” - “Which steps would you automate, and which would you keep manual?” - “What would you need to control or check before enabling it?” - “What could go wrong? Tell me about a client situation where it would.” - “What would you use instead?”

Then disclose: “We’re considering $10/month, additional to your current plan.”

“What would make that worth paying—or not? Which current effort or expense would it replace? Who would approve it?”

Don’t offer a $5/$10/$15 menu or treat “yes, useful” as purchase evidence.

26–30 minutes: Commitment and close

“If we offered a paid pilot at $10/month, what would you need before deciding?”

Ask whether they’d join a follow-up evaluation or discuss a pilot with the budget owner. Record the specific next step; don’t imply availability or collect payment.

“What important part of chasing payments haven’t we covered?”

4. Notes for first-time interviewers

Alternate moderator and note-taker roles; do one practice call together. Ask one question at a time, allow silence, and follow concrete examples. Avoid praise, pitching, defending, and questions that assume pain. Separate observed behavior, exact quotes, estimates, and your interpretations.

Debrief immediately: frequency, cost, workaround, automation boundaries, price reaction, commitment, and contradictory evidence.

5. Decision rules

Agree these directional gates before interviewing; eight calls cannot estimate market demand.

Go toward a scoped build if at least five describe recurring, consequential pain—including two non-reminder users—and at least three connect $10 to concrete value and take a specific pilot-evaluation step. There must also be a credible way to address relationship and control risks. Validate paid adoption before committing substantial engineering.

Stop or reshape if lateness rarely causes meaningful cost, chasing is easy, automation cannot address the underlying causes, or users prefer manual control and won’t pay.

Mixed findings mean targeted follow-up—not “eight customers validated it.”

Grades and run details

Decision model 96 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly60%
  • passAddresses the actual decision95%
  • passRespects explicit constraints62%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims74%
  • passProduces the required deliverable93%
  • passTalks to the right people100%
  • passCoaches first-time interviewers87%
  • passDesigned to falsify96%
  • passAsks what people did, not what they think84%
  • passDoesn't lead the witness78%
  • partialMarks what to cut if the call runs over55%
Run
Run
#1
API response time
44 s
Submitted
30 Sept 2026

Needs more than a quick edit9 of 12 checks passed

The guide is strong and mostly usable, but it needs evidence fixes and a must-ask/cut plan before a PM could run it without re-checking assumptions.

Output

Autochase Call Guide: 8 calls, 30 minutes each

What we're trying to learn

Sam's hypothesis has three parts, and each can fail separately: 1. Cost: Late payment hurts in money or time, not just annoyance. (34% of invoices are late, but that doesn't mean it hurts.) 2. Behavior: People already try to fix it. Only 12% of accounts used the manual Send reminder button, so we need to know why the other 88% didn't. 3. Willingness: They'd let software message their clients, and pay $10 a month on top of $12 for Pro.

Interviews can't settle price. They can tell us whether it's worth testing.

Who to talk to

Sam's plan (the 300 heaviest button users, first 8 to reply) would only find people who already like reminders and respond fast to email. Recruit against quotas instead:

  • 2 heavy button users (Sam's list is fine for these)
  • 3 accounts with several late invoices in the last 90 days who never used the button
  • 1 who filed an Autochase support ticket
  • 2 agencies (3+ people) and 2 solo freelancers. These overlap with the groups above.
  • Mix of Pro and free accounts, with at least 2 free.

Screen for at least one invoice paid late in the last 90 days. Pick people to fill the quotas, not first come first served. Offer a $50 gift card.

Call structure

0:00–0:03 Intro. "We're learning how people get paid, and we're not selling anything. There are no right answers, and critical feedback helps most. Okay to record?"

0:03–0:05 Context. What do you do, how many clients, how many invoices a month?

0:05–0:15 The last late invoice (the core of the call) 1. Tell me about the last time a client paid you late. Probe: how late, how much, what happened first, what did you do, how did it feel? 2. Walk me through what you do between sending an invoice and getting paid. Probe: what tool, how often, who does it? 3. What did that lateness cost you? Probe: hours, cash flow, bills or contractors you had to delay.

0:15–0:21 Current behavior 4. (Button users) What made you use Send reminder? (Non-users) What do you do instead, or why not just let it go? 5. Is there a client you would never chase? Why? 6. Have you tried anything else, such as late fees, deposits, shorter terms or other tools? What happened?

0:21–0:27 Concept. Read this, don't demo it: "Suppose Paydeck automatically emailed your client at set points after the due date, getting firmer each time." 7. What's your first reaction? What worries you? 8. What would have to be true for you to switch it on? Probe: tone, control, per-client exceptions, seeing the emails first.

0:27–0:30 Money and wrap 9. What do you pay for today that helps you get paid or run your finances? 10. "We'd likely charge extra for this. Would you want early access when it's ready?" Note the exact response. Hesitation counts as data. 11. Anything I should have asked? Anyone else we should talk to?

Why Sam's questions changed

  • "How frustrating is chasing?" and "Would it be useful?" invite polite yes answers. Past-behavior questions don't.
  • "$5, $10 or $15?" anchors people on our options and measures nothing real. Questions 9–10 are better evidence.
  • "Would your clients be annoyed?" asks for a guess about someone else. Question 5 asks what they've actually avoided.
  • "What other features?" is off-topic. Drop it.

Notes for first-time interviewers

  • Split roles. One person asks and the other takes verbatim notes. Swap halfway through the 8 calls.
  • Ask about the past, not the future. "What did you do last time?" beats "Would you ever...?"
  • Don't pitch or rescue. Don't name Autochase until minute 21. If they struggle to answer, wait, then ask "Can you say more?" Silence is fine.
  • Follow the story. Ask "why?" and "what happened next?" more than you ask the next scripted question. The script is a backstop.
  • Discount compliments. "That sounds great" is worth little. Specific past actions, money spent and workarounds built are worth a lot.
  • Stay neutral if they criticize the product or the idea. Don't defend it.
  • Debrief within 30 minutes of each call. Write down three surprises and the best quotes. Do a practice run with a colleague first.

What we'd need to hear

Count only the people who had a late invoice.

Go (build, or run a paid pre-sale test): - At least 5 of 8 describe a specific recent late payment that cost real money, stress or more than about an hour a month. - At least 5 already chase manually or with workarounds, including people who skip the button, and can explain why the button doesn't work for them. - At least 4 would turn on automation given controls we could build, and at least 3 say yes to early access for a paid add-on.

Stop or rethink: - Most shrug lateness off, or say it comes from a few clients they'd never automate against. - The real problem is something else, such as clients not opening invoices, wanting deposits, or needing a payment link, and reminders wouldn't fix it. - Most won't let software contact their clients, or wouldn't pay anything for it.

Mixed results (for example, it only matters for agencies, or only with tone controls): narrow the target segment and run a second round of calls.

Even a clear "go" from 8 calls only justifies a cheap next test, such as a fake-door button or a pre-sale. It doesn't justify committing the quarter.

Claims the judge couldn’t find in the brief

  • There is an 'Autochase support ticket' population to recruit from.
  • Sam's plan would only find people who already like reminders and respond fast to email.
  • We can identify accounts with several late invoices in the last 90 days who never used the button.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly26%
  • passAddresses the actual decision96%
  • passRespects explicit constraints38%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims68%
  • passProduces the required deliverable96%
  • passTalks to the right people99%
  • passCoaches first-time interviewers100%
  • passDesigned to falsify96%
  • passAsks what people did, not what they think77%
  • passDoesn't lead the witness61%
  • partialMarks what to cut if the call runs over61%
Run
Run
#1
API response time
32 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyRightRightMixed
GPT-6 Luna · API

The output uses the supplied facts correctly and labels recruiting assumptions and likely bias as hypotheses rather than established facts.

GPT-6.1 Sol · API

All factual claims about the current situation are directly from the supplied context.

Sonnet 5.5 · API

It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.

Respects explicit constraintsWrongRightRight
GPT-6 Luna · API

It violates the explicit prohibition on asking people to predict what they would pay or use by asking whether they would trial it at $10/month.

GPT-6.1 Sol · API

Includes all required sections, stays under 900 words, and is written for Mei and Tomas.

Sonnet 5.5 · API

It is a call guide under 900 words with learning goals, timed questions, interviewer notes, recruiting, and go/stop signals.

Avoids unsupported claimsRightRightMixed
GPT-6 Luna · API

It avoids presenting unproven causes or demand as fact, using language such as “likely” and “directional.”

GPT-6.1 Sol · API

Interpretations are clearly separated from facts; no unsupported causal claims are presented as established.

Sonnet 5.5 · API

It presents some interpretations as established, especially that Sam's list would only find reminder-likers/fast email responders and that Autochase tickets exist.

Doesn't lead the witnessMixedRightRight
GPT-6 Luna · API

The price section asks hypothetical willingness-to-pay/trial questions, and the idea section asks for first reactions to a described product, which are leading relative to the prohibited practices.

GPT-6.1 Sol · API

Questions are neutral and open; Autochase is introduced only after 18 minutes of problem exploration.

Sonnet 5.5 · API

It removes Sam's leading/hypothetical questions, keeps the concept late, and avoids pitching before current-behavior questions.

All got wrong 1

Marks what to cut if the call runs overWrongWrongWrong
GPT-6 Luna · API

The timings add to 30 minutes, but the guide does not mark must-ask questions or say what to cut if time runs short.

GPT-6.1 Sol · API

Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.

Sonnet 5.5 · API

Timings add to 30 minutes, but it does not mark must-ask questions or say what to cut if time runs short.

All got right 7

Addresses the actual decisionRightRightRight
GPT-6 Luna · API

It gives clear go, stop/defer, and mixed-evidence decision rules tied to specific call outcomes.

GPT-6.1 Sol · API

Provides clear go/stop conditions tied to specific call outcomes, framed for the team.

Sonnet 5.5 · API

It commits to a clear call: 8 calls can justify only a cheap next test, not committing the quarter, and gives go/stop conditions.

Identifies material uncertaintyRightRightRight
GPT-6 Luna · API

It names key unknowns such as recurring burden, controls/trust concerns, and segment concentration, and says how calls would resolve them.

GPT-6.1 Sol · API

Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.

Sonnet 5.5 · API

It names material unknowns—cost, behavior, willingness, client-contact concerns—and says how calls plus a cheap test would resolve them.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

It provides a usable 30-minute call guide with learning goals, questions, timings, interviewer notes, recruiting guidance, and decision criteria within the length limit.

GPT-6.1 Sol · API

Complete call guide with learning goals, timed questions, coaching, recruiting plan, and decision rules, usable as-is.

Sonnet 5.5 · API

Mei and Tomas could run the calls from it with light edits.

Talks to the right peopleRightRightRight
GPT-6 Luna · API

It explicitly avoids recruiting only the 300 heaviest reminder users and balances the sample with 4 reminder users and 4 non/rare users with late invoices.

GPT-6.1 Sol · API

Explicitly avoids only heavy reminder users, gives a balanced split with rationale and includes non-reminder users.

Sonnet 5.5 · API

It explains the heaviest-button-user bias and gives a quota mix including non-button users with late invoices, ticket requesters, agencies, solos, and free accounts.

Coaches first-time interviewersRightRightRight
GPT-6 Luna · API

It gives first-time interviewers concrete follow-up prompts, silence guidance, note-taking rules, and prohibitions on pitching or defending.

GPT-6.1 Sol · API

Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.

Sonnet 5.5 · API

It gives specific first-timer instructions: role split, verbatim notes, silence, follow-ups, no pitching/defending, debriefs, and a practice run.

Designed to falsifyRightRightRight
GPT-6 Luna · API

Each learning goal has questions that could falsify the hypothesis, such as asking what they actually did, what it cost, and what would make them not use automation.

GPT-6.1 Sol · API

Every learning goal has disconfirming questions, e.g., 'What did you do next?', 'When do you avoid sending another reminder?'

Sonnet 5.5 · API

Each learning goal has disconfirming questions about recent late payments, actual chasing behavior, client-contact limits, and past spending/early-access response.

Asks what people did, not what they thinkRightRightRight
GPT-6 Luna · API

The core questions ask for the most recent late invoice, what happened, what they did, and current time or money spent.

GPT-6.1 Sol · API

Core questions ask for the most recent overdue invoice and specific actions taken, not opinions or predictions.

Sonnet 5.5 · API

The core asks for the last late invoice, what happened, what they did, what it cost, and what they already tried.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review