Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a PM at Paydeck. Next week, Mei (a designer) and Tomas (an engineer) will run 8 customer calls to test whether we should build Autochase next quarter. Neither has run a customer interview before. Write the call guide they'll use: the learning goals, the questions with rough timings for a 30-minute call, notes for first-time interviewers, who we should talk to, and what we'd need to hear to go ahead or to stop. Keep it under 900 words. Sam, the PM who owns Autochase, has drafted some questions and a recruiting plan. They're below with everything else we know.

What the model was given6 items: About Paydeck, The idea, Sam's hypothesis, What we know, Sam's draft questions, Sam's recruiting plan
About PaydeckInvoicing software for freelancers and small agencies (1 to 10 people). 48,000 accounts, 6,100 of them paying $12 a month for Pro.
The ideaAutochase: automatic, escalating reminders for overdue invoices, sold as a $10-a-month add-on. There is no prototype yet.
Sam's hypothesis“Late payments cost our users real money, and chasing them is painful enough that they'd pay $10 a month to automate it.”
What we know34% of invoices sent through Paydeck are paid after their due date; the median is 16 days late. The manual 'Send reminder' button was used by 12% of accounts in the last 90 days. 40 support tickets last quarter asked for automatic reminders.
Sam's draft questions1. How frustrating do you find chasing late payments? 2. Would an automatic reminder feature be useful to you? 3. How much would you pay for it: $5, $10 or $15 a month? 4. Do you think your clients would be annoyed by automated reminders? 5. What other features would you like us to build?
Sam's recruiting planEmail the 300 accounts that used the 'Send reminder' button most last month, and book the first 8 who reply.
What a strong answer doesThe answer key the graders mark against

A guide that splits Sam's hypothesis into goals that can each fail: do late payments actually cost users time or money, how do they deal with them today, would they pay to fix it, and does chasing harm client relationships. The core asks for the story of the last overdue invoice (what happened, what they did, how long it took, what it cost) and what they've already tried or paid for, instead of opinions or predictions. It replaces Sam's leading and hypothetical questions: no 'how frustrating', no 'would it be useful', no price menu; payment is tested through past spending or a real commitment (say, a paid pilot sign-up) at the end, after the problem questions and without pitching Autochase first. It fits 30 minutes with timings, coaches first-timers (ask for specifics, follow up with 'what happened next', stay quiet, don't pitch or defend), and fixes the recruiting plan: the heaviest reminder users will over-state demand, so it adds accounts with overdue invoices that never send reminders, and says who to exclude or balance. It ends with clear go and stop signals tied to what people did, such as most calls describing recent chasing that took real time plus a workaround they've tried or paid for, against most saying one nudge does the job.

Critical failures (cap the score at 40)
  • Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision
  • Describes or pitches Autochase before asking how the person handles late payments today
Case

Mid-level · v1.1 · synthetic · B2B SaaS, freelancers, willingness to pay, first-time interviewers

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 12 checks passed

The guide is well-structured and complete; a PM could add a note about cutting questions if time runs short in a quick edit, without re-checking evidence or redoing reasoning.

Output

Customer Interview Guide: Autochase Discovery

Interviewers: Mei & Tomas | Call length: 30 min | Target: 8 calls

---

1. Learning Goals

  • Understand the real cost: Does chasing late payments cost users actual time/cash flow, or is it just a minor nuisance?
  • Current behaviors & workarounds: How do they handle late invoices today? What tools or manual steps do they use?
  • Fear vs. relief: Are users worried automated reminders will harm client relationships?
  • Willingness to pay: Is the pain acute enough to warrant a $10/mo add-on (nearly doubling their $12/mo Pro fee)?

---

2. Who We Are Talking To (Recruiting Update)

Sam’s plan to recruit only the top 'Send reminder' clickers creates selection bias—we’d only hear from people comfortable nudging clients. We need 8 users split across two segments: 1. Active Chasers (4 users): Have late invoices and regularly use the manual reminder button. 2. Reluctant Chasers (4 users): Have late invoices (>14 days overdue) but rarely or never use the reminder button. (Why aren't they sending them? Fear? Embarrassment?) Include a mix of solo freelancers and 2–10 person agencies.

---

3. Call Structure & Questions (30 Minutes)

#### 00:00–03:00 | Welcome & Context * "Thanks for joining! I’m Mei/Tomas. We’re working on how Paydeck handles invoicing workflows. We aren't selling anything today—we just want to understand how you run your business. There are no right or wrong answers." * Ask permission to record.

#### 03:00–10:00 | Past Behavior & The Real Pain Instead of asking "How frustrating is this?", anchor on their actual past experiences. * "Think back to the last time a client was late on an invoice. Can you walk me through what happened?" * "How late was the payment, and how did it impact your business or cash flow?" * "How often does this happen across all your clients?"

#### 10:00–18:00 | Current Workflow & Emotional Friction * "When an invoice became overdue, what exact steps did you take? What did you say, and when?" * "What made you decide to reach out at that specific moment rather than earlier or later?" * (For non-clickers): "We noticed you rarely use the 'Send reminder' button in Paydeck. How come?" * "How do your clients react when you remind them? What worries do you have, if any, about nudging them?"

#### 18:00–26:00 | Solution Testing & Willingness to Pay Instead of pitching Autochase directly or asking hypothetical pricing, test the concept against their current spend. * "Have you tried any other tools, templates, or processes to get paid faster? What did you like or dislike about them?" * "If Paydeck could automatically send reminders on a schedule you choose, would you turn it on? Why or why not?" * "What would hold you back from letting software message your clients automatically?" * "If this existed today as a paid add-on, how would you evaluate whether it’s worth paying for?"

#### 26:00–30:00 | Wrap-up * "Is there anything else about getting paid that we didn’t cover?" * Thank them; note that their input will directly shape the product.

---

4. Tips for First-Time Interviewers

  • Split roles: One person leads the conversation; the other takes notes. Alternate roles across calls.
  • Never ask hypotheticals: People lie when predicting the future ("Would you use X?"). Ask what they actually did in the past.
  • Embrace awkward silence: After a user answers, wait 3 seconds. They will almost always add crucial context.
  • Don't defend the product: If they criticize Paydeck, say: "Tell me more about that."

---

5. Go / No-Go Decision Criteria

#### ✅ GO (Green light to build Autochase): * Acute pain confirmed: At least 6 of 8 users report overdue invoices directly disrupting their business (e.g., missed personal paychecks, hours lost drafting awkward emails). * High trust/low fear: Users are eager to automate reminders, provided they have control over tone, cadence, and exclusions. * Clear WTP: Users view late payment collection as an ROI-positive problem where paying ~$10/mo to recover thousands in cash flow is an obvious tradeoff.

#### 🛑 NO-GO (Do not build, or pivot the concept): * Relationship sensitivity: Users refuse automated outreach because they require bespoke, gentle communication to preserve client relationships. * Trivial problem: Late payments are accepted as "cost of doing business," resolved with one quick text, or infrequent enough that manual follow-up is sufficient. * Price mismatch: Users balk at a $10 add-on for a product that costs $12/month (signal: consider baking basic reminders into the core Pro plan to reduce churn rather than selling it as a standalone paid add-on).

Claims the judge couldn’t find in the brief

  • Sam’s plan to recruit only the top 'Send reminder' clickers creates selection bias—we’d only hear from people comfortable nudging clients.

Grades and run details

Decision model 83 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly24%
  • passAddresses the actual decision85%
  • partialRespects explicit constraints29%
  • passIdentifies material uncertainty86%
  • passAvoids unsupported claims23%
  • passProduces the required deliverable41%
  • passTalks to the right people97%
  • passCoaches first-time interviewers33%
  • passDesigned to falsify56%
  • partialAsks what people did, not what they think72%
  • partialDoesn't lead the witness37%
  • partialMarks what to cut if the call runs over58%
Run
Run
#1
API response time
20 s
Submitted
1 Oct 2026
Critical failure: Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision. Score capped at 40.

Needs more than a quick edit9 of 12 checks passed

A PM could use most of it with light edits, but the prohibited hypothetical price/trial questions and missing must-ask/cut guidance need correction before it is safe to run.

Output

Autochase customer call guide (30 minutes)

What we need to learn

  1. How often late payments happen, and what customers actually do when they do.
  2. Whether chasing creates meaningful time, cash-flow, or relationship costs—and how customers handle those today.
  3. When automatic reminders would help, and what control or safeguards customers would need.
  4. Whether the value could justify a $10/month add-on.

Who to talk to

Recruit 8 account owners or purchase decision-makers who send invoices through Paydeck and have had at least one late invoice in the last 90 days. Aim for: - 4 who used “Send reminder” and 4 with late invoices who rarely or never used it. - A mix of freelancers and small agencies, account sizes, and Pro/free plans.

Don’t recruit only the 300 most frequent button users: they’re likely to overrepresent people already motivated to chase. Ask about recent late invoices; don’t tell people the feature idea before the call.

30-minute call

0–2 min | Set the tone “Thanks for speaking with us. We’re learning how invoicing works for you; there are no right answers. I’m interested in what you’ve actually done, not what you think we want to hear. We’re exploring ideas, not selling a product.” Ask permission before recording; otherwise take notes.

2–8 min | Recent experience “Think of the most recent invoice that was paid late. Can you walk me through what happened, from sending it to getting paid?” Probe: When was it due and paid? When did you realize it was late? What did you do next? What happened?

8–15 min | Current workflow “What, if anything, did you do to follow up?” Probe: How did you decide when and how to contact them? What did you send, and through what channel? How often does this happen? What do you do when you don’t send a reminder? Have you tried tools or changed your process?

15–20 min | Impact and trade-offs “What effect did that late payment have on your work or business?” Probe: Time spent? Cash-flow impact? Client relationship? What was most frustrating—or not especially difficult? How do you prioritize which invoices to chase?

20–25 min | Explore the idea After hearing their story, say neutrally: “We’re exploring a paid add-on that sends escalating reminders automatically when an invoice becomes overdue. There isn’t a product yet.” “What’s your first reaction?” Probe: What would be useful or unhelpful? Which invoices, if any, would you trust it with? What would you need to control—timing, wording, recipients, stopping reminders? What situations would make you not use it?

25–28 min | Value and price “How do you currently spend time or money dealing with late invoices?” “If this cost $10 a month, how would you decide whether it was worth it?” Probe: What would it need to save or improve? Would you trial it at that price? Why or why not? Treat stated intent as a clue, not proof of demand.

28–30 min | Close “What have I not asked that matters about getting paid late?” Thank them. Don’t promise a launch or specific functionality.

Notes for first-time interviewers

  • Mei leads; Tomas takes notes. Agree beforehand who will ask follow-ups. Don’t interrupt or tag-team.
  • Let the customer finish. Use silence; ask “Can you tell me more?” or “What happened next?”
  • Ask about a specific past event before opinions or hypotheticals. Avoid “Would this be useful?”, “How frustrating is it?”, feature wishlists, and leading follow-ups.
  • Don’t defend the idea, pitch solutions, or share Paydeck’s late-payment statistics; these can bias answers.
  • Capture examples, exact phrases, workarounds, and contradictions—not just “yes/no.” Separate what people did from what they say they might do.

Decision after 8 calls

These calls are directional, not a market-size estimate. Go ahead with a narrowly scoped MVP/pilot next quarter if at least 5 of 8 describe recurring late-payment problems with concrete business or time costs, and at least 4 of 8 see a credible reason to try automation at $10/month and can explain the controls they’d need. Check that interest isn’t confined to one narrow user type.

Stop or defer if 2 or fewer describe meaningful recurring burden, most already handle it easily, or trust/client concerns make automation unacceptable even with reasonable controls. If evidence is mixed or concentrated in one segment, don’t commit broadly: refine the target or concept and test again.

Claims the judge couldn’t find in the brief

  • The output’s recruiting filter assumes there are at least 8 account owners or purchase decision-makers with a late invoice in the last 90 days.
  • The heaviest reminder users are likely to overrepresent people already motivated to chase.

Grades and run details

Decision model 92 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly68%
  • passAddresses the actual decision93%
  • partialRespects explicit constraints18%
  • passIdentifies material uncertainty89%
  • passAvoids unsupported claims76%
  • passProduces the required deliverable93%
  • passTalks to the right people100%
  • passCoaches first-time interviewers100%
  • passDesigned to falsify85%
  • passAsks what people did, not what they think64%
  • passDoesn't lead the witness61%
  • partialMarks what to cut if the call runs over48%
Run
Run
#1
API response time
31 s
Submitted
30 Sept 2026
Critical failure: Relies on hypothetical questions (would you use it, how much would you pay) as the main evidence for the decision. Score capped at 40.

Needs more than a quick edit3 of 12 checks passed

A PM would need to fix the recruiting plan, remove the pitch and hypothetical price test, and replace the go/stop criteria with behavior-based evidence. The current guide would likely produce biased demand signals.

Output

Welcome to user research week! Mei and Tomas, running these calls is one of the highest-leverage things you can do this quarter. Don't worry about being first-timers; your fresh eyes are an asset. Here is your battle-tested guide for next week.

---

1. Learning Goals

We need to answer two core questions before committing engineering time to Autochase: 1. Pain: Does chasing late invoices cause enough operational friction and cash flow anxiety that users want it off their plates completely? 2. Value: Is automation worth $10/month, or do users prefer control (e.g., hitting a manual button)?

---

2. Who We Should Talk To

Per Sam’s plan: We are emailing the 300 Pro accounts that used the manual 'Send reminder' button most frequently last month and booking the first 8 who reply. * Why: These users already feel the pain. If they won't pay for automation, nobody will.

---

3. Notes for First-Time Interviewers (Read This Twice!)

  • Do not pitch Autochase early. If you pitch it, they will be polite and say yes. Sell the problem, not our solution.
  • Past behavior beats future promises. When people say "I would pay $10 for that," they are often lying to be nice. Ask about what they currently do and what they've paid for in the past.
  • Embrace silence. After you ask a question, stop talking. Let them fill the silence. The best insights live in the second minute of their answer.
  • One person leads, one takes notes. Mei, take the lead on questions today; Tomas, focus on verbatim quotes and body language (and vice versa for call #2).

---

4. The 30-Minute Call Guide

#### Phase 1: Context & Warm-up (5 mins) Goal: Build rapport and understand their business. * "Tell me a bit about your business—what do you do, and who are your typical clients?" * "Walk me through how invoicing and getting paid works in your workflow today."

#### Phase 2: Past Behavior & Pain (10 mins) Goal: Uncover the true emotional and financial cost of late payments. * [Replace Sam's Q1] "Take me back to the last time an invoice went significantly past its due date. What happened?" * "How did you handle it? Walk me through the exact steps you took to chase them." * "How does chasing clients for money impact your relationship with them, or your peace of mind?" * Data check: 34% of our invoices are paid late (median 16 days). Ask: "Does that match your reality?"

#### Phase 3: Current Workarounds & Solution Testing (10 mins) Goal: See how they solve this now, then introduce Autochase. * "You’ve used our manual 'Send reminder' button recently. How is that working for you? What do you wish it did that it doesn't?" * The Reveal: "We’re exploring a feature called Autochase that automatically sends escalating reminders on a schedule you set until the invoice is paid. Based on what you've told me about your workflow, how would something like that fit into your business?" * [Replace Sam's Q4] "Some freelancers worry automated reminders might annoy their clients. How do you feel about that risk for your clients?"

#### Phase 4: Pricing & Wrap-up (5 mins) Goal: Test the $10/mo price point and wrap up. * [Replace Sam's Q2/Q3] "We're looking at pricing this as a $10-a-month add-on. When you think about the time it saves or the awkwardness it avoids, does $10/mo feel like a bargain, too expensive, or about right?" * "Before we wrap up, is there anything else about getting paid by clients that drives you crazy that we didn't touch on?" * Thank them warmly and end the call.

---

5. Go / No-Go Criteria

After 8 calls, Mei, Tomas, and I will sync. We move forward to build Autochase next quarter only if: 1. At least 6 out of 8 interviewers describe late payments as a top-3 operational headache (not just a minor annoyance). 2. Users unprompted mention spending hours a month on this, or losing sleep over cash flow. 3. At least 5 out of 8 react to the $10/month price point as reasonable or cheap relative to the value.

We stop and pivot if: * Users say, "I actually like reviewing and sending reminders manually because I customize the tone per client." (This kills the automation thesis). * Users think $10/month is wildly overpriced for a reminder tool.

Claims the judge couldn’t find in the brief

  • We are emailing the 300 Pro accounts that used the manual 'Send reminder' button most frequently last month and booking the first 8 who reply.
  • These users already feel the pain.
  • If they won't pay for automation, nobody will.
  • When people say 'I would pay $10 for that,' they are often lying to be nice.
  • Autochase automatically sends escalating reminders on a schedule the user sets until the invoice is paid.
  • Autochase would save time and avoid awkwardness.
  • Late payments cause cash-flow anxiety and operational friction for users.
  • Running these calls is one of the highest-leverage things the team can do this quarter.
  • The guide is battle-tested.
  • Users may prefer control, such as hitting a manual button.
  • At least 6 out of 8 interviewers will describe late payments as a top-3 operational headache if the feature is worth building.
  • At least 5 out of 8 users will react to the $10/month price point as reasonable or cheap if the feature is worth building.

Grades and run details

Decision model 50 · LLM judge 4 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly14%
  • passAddresses the actual decision60%
  • failRespects explicit constraints80%
  • passIdentifies material uncertainty28%
  • partialAvoids unsupported claims35%
  • partialProduces the required deliverable75%
  • failTalks to the right people100%
  • partialCoaches first-time interviewers61%
  • partialDesigned to falsify44%
  • partialAsks what people did, not what they think79%
  • failDoesn't lead the witness56%
  • partialMarks what to cut if the call runs over57%
Run
Run
#1
Time to output
5 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 8

Uses the supplied evidence correctlyRightRightMixed
Gemini 3.8 Flash · API

All factual statements about the current situation are taken directly from the brief or derived by simple arithmetic; the selection-bias claim is an interpretation, not a fact about the current product or data.

GPT-6 Luna · API

The output uses the supplied facts correctly and labels recruiting assumptions and likely bias as hypotheses rather than established facts.

Gemini 3.5 Flash-Lite · Gemini

It misstates the recruiting plan as Pro accounts and presents several unlabelled assumptions about users' pain, honesty, and preferences as current facts.

Respects explicit constraintsRightWrongMixed
Gemini 3.8 Flash · API

The output is under 900 words, includes all requested sections (learning goals, timed questions, interviewer notes, recruiting, decision criteria), and is addressed to Mei and Tomas.

GPT-6 Luna · API

It violates the explicit prohibition on asking people to predict what they would pay or use by asking whether they would trial it at $10/month.

Gemini 3.5 Flash-Lite · Gemini

It is a call guide for Mei and Tomas, includes the requested sections, has 30-minute timings, and appears under 900 words.

Identifies material uncertaintyRightRightMixed
Gemini 3.8 Flash · API

The go/no-go criteria implicitly identify the unknowns (real cost, fear, willingness to pay) and specify the conditions that would resolve them and change the decision.

GPT-6 Luna · API

It names key unknowns such as recurring burden, controls/trust concerns, and segment concentration, and says how calls would resolve them.

Gemini 3.5 Flash-Lite · Gemini

It does not name the key uncertainty that heavy reminder users may overstate demand, nor does it bound the unknowns with behavior-based tests.

Avoids unsupported claimsMixedRightWrong
Gemini 3.8 Flash · API

The claim that recruiting only top reminder users means 'we’d only hear from people comfortable nudging clients' is presented as fact but is an unsupported inference not established by the brief.

GPT-6 Luna · API

It avoids presenting unproven causes or demand as fact, using language such as “likely” and “directional.”

Gemini 3.5 Flash-Lite · Gemini

It presents unsupported claims such as 'these users already feel the pain', 'if they won't pay nobody will', and that people saying they would pay are often lying.

Talks to the right peopleRightRightWrong
Gemini 3.8 Flash · API

The guide explicitly calls out the selection bias in Sam's plan and proposes a balanced split of 4 active chasers and 4 reluctant chasers, with a mix of freelancers and agencies.

GPT-6 Luna · API

It explicitly avoids recruiting only the 300 heaviest reminder users and balances the sample with 4 reminder users and 4 non/rare users with late invoices.

Gemini 3.5 Flash-Lite · Gemini

It keeps Sam's plan of only the heaviest reminder users and does not explain that this overstates demand or add non-chasing overdue-invoice users.

Designed to falsifyRightRightWrong
Gemini 3.8 Flash · API

Each learning goal is paired with a disconfirming question (e.g., 'walk me through what happened' for real cost, 'what worries do you have' for fear, 'how would you evaluate' for willingness to pay).

GPT-6 Luna · API

Each learning goal has questions that could falsify the hypothesis, such as asking what they actually did, what it cost, and what would make them not use automation.

Gemini 3.5 Flash-Lite · Gemini

The main value test is a hypothetical price reaction after pitching Autochase, so the guide cannot cleanly falsify willingness to pay.

Asks what people did, not what they thinkRightRightWrong
Gemini 3.8 Flash · API

The core questions anchor on the last overdue invoice and specific actions taken, with opinions and hypotheticals kept secondary and placed after the behavioral section.

GPT-6 Luna · API

The core questions ask for the most recent late invoice, what happened, what they did, and current time or money spent.

Gemini 3.5 Flash-Lite · Gemini

Although it asks about the last overdue invoice, the core decision evidence includes pitching the solution and asking price reactions, not only recent behavior.

Doesn't lead the witnessRightMixedWrong
Gemini 3.8 Flash · API

Questions are neutral and open-ended; the Autochase concept is introduced only after exploring current behavior, and no price menu or leading 'how frustrating' questions are used.

GPT-6 Luna · API

The price section asks hypothetical willingness-to-pay/trial questions, and the idea section asks for first reactions to a described product, which are leading relative to the prohibited practices.

Gemini 3.5 Flash-Lite · Gemini

It describes and pitches Autochase before pricing and asks leading questions such as whether $10 feels like a bargain and whether automation fits their workflow.

All got wrong 1

Marks what to cut if the call runs overWrongWrongWrong
Gemini 3.8 Flash · API

The guide includes timed sections that sum to 30 minutes and a realistic number of questions, but it does not say what to cut if time runs short, as the criterion requires.

GPT-6 Luna · API

The timings add to 30 minutes, but the guide does not mark must-ask questions or say what to cut if time runs short.

Gemini 3.5 Flash-Lite · Gemini

The timings add to 30 minutes, but the guide has too many questions and no must-ask or cut-if-short instructions, making it unrealistic for first-time interviewers.

All got right 3

Addresses the actual decisionRightRightRight
Gemini 3.8 Flash · API

The guide includes clear go/no-go criteria that state what would make the team proceed or stop, fulfilling the brief's request for what we'd need to hear.

GPT-6 Luna · API

It gives clear go, stop/defer, and mixed-evidence decision rules tied to specific call outcomes.

Gemini 3.5 Flash-Lite · Gemini

It gives a clear conditional go/no-go rule for whether to build Autochase next quarter, though the rule itself is weak.

Produces the required deliverableRightRightRight
Gemini 3.8 Flash · API

The output is a complete call guide with all required elements, under 900 words, and directly usable by Mei and Tomas.

GPT-6 Luna · API

It provides a usable 30-minute call guide with learning goals, questions, timings, interviewer notes, recruiting guidance, and decision criteria within the length limit.

Gemini 3.5 Flash-Lite · Gemini

The requested call guide is present and usable as a draft, with learning goals, questions, timings, interviewer notes, recruiting, and go/stop criteria.

Coaches first-time interviewersRightRightRight
Gemini 3.8 Flash · API

The tips section gives specific, actionable advice: split roles, avoid hypotheticals, embrace silence, don't defend the product, with example responses.

GPT-6 Luna · API

It gives first-time interviewers concrete follow-up prompts, silence guidance, note-taking rules, and prohibitions on pitching or defending.

Gemini 3.5 Flash-Lite · Gemini

It gives practical first-timer coaching: don't pitch early, ask about past behavior, embrace silence, and split lead/notes roles.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review