Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Crate. In four weeks, the exec team decides which of two bets gets our two squads next half, and both bets rest on a theory of why customers are leaving. We have three weeks for 12 customer calls of 45 minutes. Write the guide for them. It should include: the learning goals; the questions, with timings, for the two kinds of people we'll talk to; guidance for whoever runs the calls; any changes you'd make to the plan below; and what we'd need to hear to back each bet, or neither. Keep it under 1,500 words. The exec team will read it before the calls start.

What the model was given8 items: About Crate, The decision, CEO, Ana Ferreira, CPO, Marcus Hill, Exit survey (41 churned customers, last 12 months), Implementation and CRM data, The call plan, Notes from Customer Success and Legal
About CrateWarehouse management software for mid-size third-party logistics companies (3PLs). 340 customers, $22M ARR. Gross logo churn rose from 11% to 18% over the last year.
The decisionBet A: integrations with warehouse robotics (autonomous carts, pick-assist robots). Bet B: rebuild implementation so customers go live faster. Two squads for the next half go to one of them.
CEO, Ana Ferreira“We're losing warehouses because they're automating and we don't talk to their robots. Every lost customer I've spoken to mentions it.”
CPO, Marcus Hill“We lose them in the first year because implementation takes forever. They're gone before they ever see the value.”
Exit survey (41 churned customers, last 12 months)A single-choice question, 'Why are you leaving?'. Missing features or integrations: 46%. Too hard to implement: 22%. Price: 17%. Other: 15% (free-text box, left blank by all but two). 'Missing features or integrations' was the first option in the list.
Implementation and CRM dataMedian time from signing to go-live: 94 days; the contract promises 45. Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year. 3 churned customers moved to a competitor with robotics integrations; robotics was cited in 9 lost new-logo deals. We don't know when in the year churned customers decided to leave.
The call plan12 calls: 6 with customers who churned in the last 9 months, 6 with at-risk customers (health score red). On each account we want to talk to the warehouse operations manager who used Crate every day and, separately, the VP of Operations or COO who signed the contract. Ana wants to run 4 of the calls with former customers she knows personally.
Notes from Customer Success and Legal5 of the 6 at-risk accounts are in renewal talks, and Customer Success is offering them discounts. Interviewers must not discuss pricing, discounts or what's on our roadmap. Every call must start by asking permission to record.
What a strong answer doesThe answer key the graders mark against

A guide that tests both theories fairly and leaves room for a third: the goals ask what actually drove each decision to leave or doubt, when it happened, and who made it, with questions that could disprove robotics (did they automate, did robots come up before they decided, what did they switch to and why) and implementation (how long go-live took, what happened in those months, whether value arrived) alike. It treats the exit survey as weak evidence (single choice, integrations listed first, 'Other' left blank) and the 2.4× figure as a correlation to probe, not a cause. The core is a timeline of the account: the story from signing to the moment they decided to leave or started doubting, with sequential probes. It has separate tracks for the daily user (what went wrong day to day) and the signer (how the decision was made, what alternatives they weighed), fits 45 minutes with timings, and opens with consent to record. It guards the calls: no pitching, pricing, discounts or roadmap promises, with a script for when an at-risk customer asks, and it addresses Ana's bias: she has a public view and personal relationships, so she should observe or take notes rather than lead, or at minimum use the same script and have her calls read separately. It ends with decision signals: what pattern across the 12 would back Bet A, what would back Bet B, and what would back neither (price, a business change, a missing capability that isn't robotics).

Critical failures (cap the score at 40)
  • Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks
  • Asks customers to choose between robotics and implementation as the reason they left
Case

Staff-level · v1.1 · synthetic · B2B SaaS, logistics, churn, exec bias, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit12 of 13 checks passed

The guide is thorough and evidence-based; a PM could add a note on which questions to prioritize if time is tight and use it immediately.

Output

Crate churn research: executive pre-read and interview guide

Decision: Allocate two squads next half to robotics integrations (A), implementation rebuild (B), or neither if the evidence does not justify either.

Research stance: Test both explanations, not confirm either executive’s account. These interviews can reveal mechanisms and decision timelines; they cannot estimate their prevalence across 340 customers or prove causality.

1. Learning goals

  1. Reconstruct the departure process: When did concerns emerge, when did leaving become likely, and when was the decision effectively made?
  2. Identify consequential problems: Separate reasons volunteered after leaving from problems that materially changed warehouse outcomes and the buying decision.
  3. Test Bet A: Was a specific robotics incompatibility a binding obstacle? Would addressing it have plausibly retained the account?
  4. Test Bet B: Did implementation delay prevent value, damage trust, or trigger departure? What caused the delay, and could Crate control it?
  5. Find alternatives and interactions: Price, service, reliability, business changes, or other missing capabilities may dominate. Robotics work may itself complicate implementation.
  6. Connect learning to an investable scope: What could two squads deliver next half that would address the demonstrated retention mechanism?

What the existing evidence does—and does not—say

  • The exit survey’s broad integration category, single-choice format, first-position placement, and sparse free text do not establish robotics as the cause.
  • The 2.4× churn association makes implementation worth investigating, but complexity, customer readiness, or other factors could drive both delay and churn.
  • Three competitor moves and nine lost deals indicate a robotics opportunity; acquisition losses are not evidence of a retention mechanism.
  • Unknown decision dates make chronology essential.

2. Changes to the call plan

Resolve the arithmetic: Twelve separate interviews cannot cover twelve accounts with two roles each. Use six accounts × two separate role interviews = twelve 45-minute calls.

Recommended mix: Four churned accounts and two at-risk accounts. This reduces reliance on renewal negotiations while retaining a prospective view. Findings will be depth-oriented, not representative.

Select accounts from the full CRM population, not executive contacts or survey answers alone:

  • Across churned accounts, include slow and faster implementations, first-year and later departures, and robotics-exposed and non-exposed customers.
  • Include at least one confirmed move to a robotics-enabled competitor and one slow implementation without an apparent robotics need.
  • For at-risk accounts, prefer one with active automation plans and one with implementation/value-realization problems. Include the non-renewal account if suitable.
  • Seek contrasting retained accounts through existing operational data, even though the call budget does not cover them.

Recruit both the daily warehouse operations manager and the contract-signing VP Operations/COO. If the original signer has left, recruit someone directly involved in the departure or renewal decision and document that substitution.

Ana should not lead calls with her personal contacts. Use a neutral interviewer without renewal responsibility. Ana can help recruit; observing requires explicit customer consent. Her existing conversations are useful leads, not independently verified findings.

CS should introduce research as separate from renewal discussions, then step out. Participation must not affect commercial treatment.

3. Interview guides: 45 minutes each

Use the relevant role column. In every section, distinguish direct experience from hearsay. Ask broad questions before naming either bet.

TimeDaily warehouse operations managerVP Operations / COO
0–3 min: permission and framingFirst ask: “May we record this conversation for internal research?” If declined, continue with notes. Then explain purpose and confidentiality limits.Same opening.
3–7 min: context and expected value“What did your warehouse handle, and what was your role with Crate? What was supposed to improve?”“What led you to buy Crate? What outcomes and deadlines mattered? How would you judge success?”
7–17 min: chronological reconstruction“Walk me from signing through onboarding, first operational use, and the point when problems became serious.” “Describe a particular shift or incident.” Establish dates, workarounds, operational impact, and who knew.“Walk me from purchase to the first concern, evaluating alternatives, and the decision.” “When did staying stop being the default?” Establish dates, decision-makers, triggers, and alternatives.
17–27 min: implementation and value“What had to happen before you could use Crate successfully? Where did work stall? Who owned each step?” “When did you first achieve useful results?” Probe migration, configuration, training, integrations, staffing, and rework after an open answer.“What go-live date did you expect, and what happened? What consequences did that have?” “What other factors accompanied the delay?” “If the same product had gone live in 45 days, what would still have put the relationship at risk?” Ask why.
27–36 min: workflow, automation, and unmet needs“Which workflows did Crate support poorly? Show or describe a recent example.” Then: “Were robots in use or planned? Which systems, what workflow, and what connection was required?” “What happened without it?”“What capabilities influenced staying or switching?” Then probe automation: vendor, deployment date, committed budget, required integration, and alternatives considered. “If that connection had existed, what else would have needed to change for you to stay?”
36–42 min: decision and competing explanations“Which problem mattered most in daily operations? What did you escalate, to whom, and when?” “What worked well?” “What have we missed?”“Which issues were necessary to the decision, and which were secondary?” “What would have had to be different for you to stay?” Ask about business changes, service, reliability, and other alternatives without forcing a category.
42–45 min: verify and closeSummarize the timeline and mechanism: “What have I misunderstood?” Request relevant artifacts and permission for a brief clarification follow-up.Same; verify whether operational problems actually influenced the commercial decision.

Status-specific wording

  • Churned: Ask what the replacement actually delivered, whether it went live, and whether the cited problem improved—not merely what the competitor promised.
  • At-risk: Ask what has already happened, current unresolved consequences, and what would cause escalation or departure. Do not imply a departure decision exists. Distinguish funded plans from aspirations.

4. Guidance for interviewers

  • Begin with recording permission; record only after consent. Explain that this is research, not a support, roadmap, or renewal discussion.
  • Do not discuss prices, discounts, or roadmap commitments. If raised, acknowledge without probing commercial terms: “I can’t discuss that here; your account team handles commercial conversations.” Record volunteered price concerns as an alternative explanation.
  • Avoid “Did missing robotics make you leave?” and “Would faster onboarding have saved you?” Start with events; use counterfactuals only afterward.
  • Ask for concrete examples, dates, artifacts, and consequences. “It was painful” is not yet a mechanism.
  • Do not defend Crate, disclose executive hypotheses, or offer fixes.
  • Interview roles separately before reconciling accounts. Treat disagreement as evidence, not an error to erase.
  • Use a note-taker and common template: timeline, problem, consequence, decision link, alternative explanation, evidence source, and confidence. Count accounts, not two interviewees as independent votes.

5. Evidence required for the decision

Back A: robotics integrations

Look for a repeated chain across independent accounts:

Committed robotics deployment → specific missing connection → material operational/business constraint → documented role in departure or credible near-term risk.

Stronger evidence includes decision-maker confirmation, pre-decision correspondence, and a replacement that solved the problem. Match the implicated vendors/workflows to current customer exposure. Before funding, establish that a feasible integration scope could address meaningful retention risk next half.

Mentions of “automation,” hypothetical plans, or competitor marketing are insufficient.

Back B: implementation rebuild

Look for:

Crate-controllable implementation bottleneck → delayed usable value or broken operational commitments → loss of confidence/avoidable cost → departure decision before value realization.

Seek repeated bottlenecks, role-level corroboration, and contrasts with faster implementations. Validate the mechanism against the full cohort, accounting where possible for complexity and customer readiness. Identify changes two squads could deliver; do not assume a wholesale rebuild is necessary.

Back neither—or declare the result inconclusive

Neither is justified if dominant causes lie elsewhere, robotics demand is narrow or speculative, delays are largely outside Crate’s control, or either intervention arrives too late to affect retention. If interviews remain contradictory or unsupported, report insufficient evidence, not “neither causes churn.”

Compare both bets on reachable ARR at risk, evidence strength, controllability, delivery feasibility, and time to impact—not mention counts.

6. Three-week execution

  • Week 1: Audit cohorts and timelines, recruit, pilot one role pair.
  • Week 2: Complete remaining interviews; synthesize after each pair without changing standards to favor an emerging winner.
  • Week 3: Verify discrepancies, triangulate with records, assess engineering scope, and prepare the decision memo with supporting evidence, counterexamples, uncertainties, and a recommendation.

Do not commit two squads merely because one story sounds more persuasive.

Grades and run details

Decision model 92 · LLM judge 13 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly54%
  • passAddresses the actual decision90%
  • partialRespects explicit constraints21%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims96%
  • passProduces the required deliverable89%
  • passTests both theories fairly100%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts95%
  • passDesigned to falsify63%
  • passAsks what people did, not what they think67%
  • passDoesn't lead the witness53%
  • partialMarks what to cut if the call runs over58%
Run
Run
#1
API response time
80 s
Submitted
30 Sept 2026

Needs more than a quick edit9 of 13 checks passed

A PM could use most of it after light edits, but the unsupported data-source claim and missing must-ask/cut guidance require re-checking and revision before running the calls.

Output

Crate churn research: decision guide

Decision: In four weeks, allocate two squads next half to robotics integrations (A), implementation redesign (B), or conclude that neither retention thesis is sufficiently supported.

Research window: Three weeks; twelve 45-minute interviews. These calls will test explanations and reconstruct decisions—not estimate how common each churn cause is across 340 customers.

1. Learning goals

  1. Reconstruct the departure sequence: What happened between signing, implementation, first value, operational problems, searching for alternatives, and deciding to leave?
  2. Identify the decisive mechanism: Was robotics incompatibility or implementation friction necessary to the departure, merely contributory, or a justification offered afterward?
  3. Find actionable scope: Which integrations or implementation changes could plausibly have prevented the problem? Could two squads address them within a half?
  4. Look for competing explanations and counterexamples: Other missing capabilities, service failures, customer-side constraints, or business changes may explain both delays and churn.
  5. Separate operational pain from buying decisions: Did daily users and economic buyers experience the same problem and agree on why the relationship failed?

Starting evidence is suggestive, not conclusive. The survey’s broad, first-listed integration option does not establish robotics demand. The 2.4× churn association does not establish that delays caused churn. Competitor destinations do not establish purchase motives; nine lost prospects are acquisition evidence, not retention evidence.

2. Changes to the call plan

Resolve the interview arithmetic. Twelve accounts with two separate interviews each would require 24 calls. Within the limit, recruit six accounts, interviewing both roles separately: twelve calls total.

Use four churned accounts and two at-risk accounts. This prioritizes actual decisions and reduces exposure to active renewal negotiations while retaining prospective evidence.

Select accounts purposively from the eligible pool, not through executive relationships:

  • Include late and on-time go-lives, plus a failed/never-live implementation if available.
  • Include known robotics signals and accounts without them.
  • Include first-year and longer-tenure departures.
  • Seek counterexamples: late implementation without abandonment; robotics discussion without an actual deployment or switch.

Do not attempt a fully balanced matrix with six accounts. Document selection, refusals, substitutions, and gaps. Check whether the nine-month churn window excludes relevant cases from the annual churn rise; expand to twelve months if needed.

For at-risk accounts, prioritize the one outside renewal talks. CS must confirm that participation—or refusal—will not affect commercial treatment. If a renewal makes independent research impractical, substitute another eligible account.

Ana should not lead interviews with former customers she knows. Use a neutral researcher or PM who owns neither bet. Ana can help recruit through a standard invitation and review consented recordings afterward. Avoid executive attendance that could suppress criticism.

Before calls, prepare a factual account timeline from CRM, implementation logs, support records, and available usage data. Keep interpretations separate. Examine cohort definitions and potential confounders behind the 2.4× figure.

3. Interview guides

Both guides total 45 minutes. Ask open questions first; introduce the two hypotheses only after the participant’s unaided account.

Guide A: Warehouse operations manager

TimeQuestions
0–4 minFirst words, before recording: “May we record this conversation for internal research? Saying no is completely fine; we can take notes instead.” Wait for permission before starting recording. Explain that this is research, not a sales or renewal conversation; participation is optional. “What did you personally own, and during which period?”
4–10 min“What was happening in the warehouse when you chose Crate? What job did you expect it to improve? How would you have recognized success?”
10–20 min“Walk me through signing to the first real production use.” Probe planned versus actual milestones, dependencies, workarounds, ownership, and first useful outcome. “Tell me about a specific day when progress stalled. What happened next?” If never live, trace the last completed milestone and stopping point.
20–30 minChurned: “When did you first think Crate might not work for you? Describe the incident. What happened between that and leaving?” At-risk: “How is Crate working today? Describe the most recent serious problem. Has anyone discussed changing systems? What has actually happened so far?” For both: “What did this cost operationally—time, throughput, errors, or customer commitments? Who saw it?”
30–39 min“What changes in equipment or workflow occurred during this period?” Then probe robotics neutrally: “Were robots evaluated or deployed? Which systems, for what workflow, and when? What specifically could Crate not do? What workaround did you try?” Separately: “Once implementation ended—or stalled—what problems remained?” Do not assume either issue existed.
39–45 min“Which issue mattered most, and what makes you say that? If only that issue had been resolved then, what would still have made Crate unsuitable?” Ask for optional, redacted supporting artifacts and names of decision participants. Summarize the timeline and invite corrections: “What important explanation have I missed?”

Guide B: VP Operations or COO

TimeQuestions
0–4 minUse the same recording-permission opening and research boundaries. “What was your role in selection, implementation oversight, and the decision to stay or leave?”
4–10 min“What business outcome justified choosing Crate? What deadline or event made that outcome important? What expectations were set?”
10–22 minChurned: “Walk me from the first concern to the decision to leave. When was that decision effectively made, rather than formally communicated? Who influenced it? What alternatives did you evaluate?” At-risk: “How are you evaluating whether Crate is working? Has a change been proposed or authorized? What evidence and next steps are involved?” Separate firsthand knowledge from reports by others.
22–31 min“What specific event most changed your confidence? What did you do afterward? What attempts were made to recover the relationship?” Churned: “What did you choose instead, and what requirement made it preferable? Is it operating successfully yet?” At-risk: “What requirements would any alternative have to meet?”
31–39 minTest both explanations, varying their order between interviews. “What role, if any, did implementation timing play? When did it affect your judgment?” “What role, if any, did warehouse automation play? Was there an approved project, named equipment, deployment date, or demonstrated integration requirement?” “Which problems would have remained if each had been solved?”
39–45 min“What would Crate have needed to change, by what date, to alter your decision?” Treat this counterfactual as weaker evidence than actions already taken. “What evidence might contradict your explanation?” Request optional artifacts, recap, and invite corrections.

4. Interviewer operating rules

  • Be curious, not persuasive. Do not quote Ana, Marcus, survey percentages, or the churn correlation.
  • Probe concrete events: “When?” “Who?” “What happened next?” “How do you know?” Distinguish recollection, inference, and observed action.
  • Do not discuss pricing, discounts, or roadmap commitments. If raised: “I can’t discuss commercial terms or future plans here; your account team can handle those separately.” Do not promise escalation outcomes.
  • Interview roles separately; do not reveal the other person’s answers. Disagreement is evidence to investigate, not average away.
  • Use a neutral note-taker. After each call, record chronology, spontaneous versus prompted reasons, supporting evidence, contradictions, and confidence.

5. Evidence required for the decision

Back A if multiple independent accounts—including actual departures—show committed robotics adoption, a specific blocking Crate incompatibility before the departure decision, failed workarounds, and credible evidence that resolving it would have preserved the relationship. Shared, feasible integration needs must emerge. Generic “missing features,” aspirational automation, or competitor capabilities alone do not qualify.

Back B if multiple accounts—including actual departures—show preventable implementation failures preceding loss of confidence, concrete operational consequences, and a credible path from faster successful activation to retention. Identify repeatable Crate-owned bottlenecks. Delays driven mainly by customer readiness, complexity, or already-failing relationships weaken this case.

Back neither if accounts remain contradictory, both theories depend on hypothetical claims, another mechanism dominates, or neither intervention fits the capacity and time available. If both qualify, compare affected retention exposure, intervention feasibility, and remaining uncertainty—not interview vote totals.

Run two pilot calls in week one, finish most interviews in week two, and reserve week three for remaining calls and synthesis. Produce one evidence sheet per account; paired interviews count as one account, not two confirmations. The final readout should combine those cases with cohort data, explicitly separate causality from correlation, and state what remains unproven before recommending either allocation.

Claims the judge couldn’t find in the brief

  • Support records and usage data exist and can be used to prepare account timelines.

Grades and run details

Decision model 96 · LLM judge 9 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly40%
  • passAddresses the actual decision92%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims94%
  • passProduces the required deliverable86%
  • passTests both theories fairly100%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts99%
  • passDesigned to falsify78%
  • passAsks what people did, not what they think74%
  • passDoesn't lead the witness61%
  • partialMarks what to cut if the call runs over52%
Run
Run
#1
Time to output
60 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Uses the supplied evidence correctlyRightMixed
GPT-6.1 Sol · API

All factual statements about the current situation are taken directly from the supplied context without invention.

GPT-6 Astra · ChatGPT

It invents current data sources by saying to prepare timelines from support records and usage data, which are not supplied.

Addresses the actual decisionRightMixed
GPT-6.1 Sol · API

The output recommends more research before deciding, states the supplied evidence cannot yet support a call, and specifies exactly what evidence would back each bet or neither.

GPT-6 Astra · ChatGPT

It does not commit to one answer early; it only states the decision frame and later gives evidence thresholds, leaving the exec team without a clear recommendation.

Respects explicit constraintsRightMixed
GPT-6.1 Sol · API

The guide is under 1,500 words, includes all requested sections (learning goals, timed questions for two roles, guidance, plan changes, decision signals), and is addressed to the exec team.

GPT-6 Astra · ChatGPT

It respects the no-pricing/roadmap and consent constraints, but it does not mark must-ask questions or say what to cut if time runs short, and it introduces unsupported data sources.

All got wrong 1

Marks what to cut if the call runs overWrongWrong
GPT-6.1 Sol · API

The guide includes timed sections that add up to 45 minutes but does not say what to cut if time runs short, nor does it mark must-ask questions.

GPT-6 Astra · ChatGPT

Timings add to 45 minutes, but the guide does not mark must-ask questions or specify what to cut if time runs short.

All got right 9

Identifies material uncertaintyRightRight
GPT-6.1 Sol · API

It names key unknowns (decision dates, whether robotics was a binding obstacle, whether delays were Crate-controllable) and says how the interview evidence would resolve them.

GPT-6 Astra · ChatGPT

It names key unknowns such as survey bias, correlation versus causation, timing of churn decisions, confounders, and renewal exposure, and says how calls and data checks would resolve them.

Avoids unsupported claimsRightRight
GPT-6.1 Sol · API

Interpretations are clearly labelled as such (e.g., the exit survey does not establish robotics as the cause), and no cause is presented as established fact.

GPT-6 Astra · ChatGPT

It labels exec theories and survey/correlation evidence as suggestive rather than established, though it still has a minor unsupported data-source claim.

Produces the required deliverableRightRight
GPT-6.1 Sol · API

The output is a complete guide with learning goals, timed questions for both roles, interviewer guidance, plan changes, and decision signals; an exec team could act on it.

GPT-6 Astra · ChatGPT

It provides learning goals, timed two-role questions, interviewer guidance, plan changes, and decision signals within the length, and the exec team could act on it with light edits.

Tests both theories fairlyRightRight
GPT-6.1 Sol · API

Questions could disprove robotics (e.g., 'Were robots in use? What happened without it?') and implementation (e.g., 'Where did work stall? If go-live had been 45 days, what would still have put the relationship at risk?'), and open questions leave room for a third cause.

GPT-6 Astra · ChatGPT

It asks open timeline questions before introducing either theory and gives both robotics and implementation questions that could disprove each, while leaving room for third causes.

Different questions for user and signerRightRight
GPT-6.1 Sol · API

The table provides distinct tracks: daily user gets questions about specific shifts, workarounds, and escalations; the signer gets questions about the purchase decision, alternatives, and go-live expectations.

GPT-6 Astra · ChatGPT

It has separate timed tracks for the daily warehouse operations manager and the VP/COO signer, matched to operational experience versus decision process.

Protects the calls and the accountsRightRight
GPT-6.1 Sol · API

It forbids discussing pricing/discounts/roadmap, provides a script to deflect, starts with recording consent, and has a concrete plan for Ana (she does not lead, observes with consent, her prior conversations are treated as leads).

GPT-6 Astra · ChatGPT

It starts with recording consent, bans pricing/discounts/roadmap, gives a script if raised, and concretely removes Ana from leading calls with people she knows.

Designed to falsifyRightRight
GPT-6.1 Sol · API

Each learning goal has disconfirming questions (e.g., 'What happened without the robot connection?' for Bet A, 'If the same product had gone live in 45 days, what would still have put the relationship at risk?' for Bet B).

GPT-6 Astra · ChatGPT

Each learning goal has disconfirming probes, including counterexamples, timing of decisions, failed workarounds, and questions about what would still have caused churn.

Asks what people did, not what they thinkRightRight
GPT-6.1 Sol · API

Core questions ask for a chronological walk-through, specific incidents, dates, and what they did, not opinions or predictions.

GPT-6 Astra · ChatGPT

The core questions ask for specific events, timelines, actions, alternatives evaluated, and what happened next, with hypotheticals clearly secondary.

Doesn't lead the witnessRightRight
GPT-6.1 Sol · API

Questions are open and neutral; robotics and implementation are not named until after broad problem questions, and no answer is hinted at.

GPT-6 Astra · ChatGPT

Questions are open and neutral, hypotheses are introduced only after unaided accounts, and the guide does not pitch either bet or ask customers to choose.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review