Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 9 graded outputs by 5 models. 50% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty100% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Designed to falsify100% pass
    Every learning goal has disconfirming questions, e.g., 'What did you do next?', 'When do you avoid sending another reminder?'
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Asks what people did, not what they think100% pass
    Core questions ask for the most recent overdue invoice and specific actions taken, not opinions or predictions.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Fits the call33% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints81% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Different questions for user and signer81% pass
    The guide does not provide distinct question tracks for the daily warehouse operations manager and the executive who signed; it only differentiates churned vs at-risk customers, with the same questions for both roles within each track.
    GPT-6 Luna · API · Why are warehouses leaving? Two execs, two theories

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Crate. In four weeks, the exec team decides which of two bets gets our two squads next half, and both bets rest on a theory of why customers are leaving. We have three weeks for 12 customer calls of 45 minutes. Write the guide for them. It should include: the learning goals; the questions, with timings, for the two kinds of people we'll talk to; guidance for whoever runs the calls; any changes you'd make to the plan below; and what we'd need to hear to back each bet, or neither. Keep it under 1,500 words. The exec team will read it before the calls start.

About CrateWarehouse management software for mid-size third-party logistics companies (3PLs). 340 customers, $22M ARR. Gross logo churn rose from 11% to 18% over the last year.
The decisionBet A: integrations with warehouse robotics (autonomous carts, pick-assist robots). Bet B: rebuild implementation so customers go live faster. Two squads for the next half go to one of them.
CEO, Ana Ferreira“We're losing warehouses because they're automating and we don't talk to their robots. Every lost customer I've spoken to mentions it.”
CPO, Marcus Hill“We lose them in the first year because implementation takes forever. They're gone before they ever see the value.”
Exit survey (41 churned customers, last 12 months)A single-choice question, 'Why are you leaving?'. Missing features or integrations: 46%. Too hard to implement: 22%. Price: 17%. Other: 15% (free-text box, left blank by all but two). 'Missing features or integrations' was the first option in the list.
Implementation and CRM dataMedian time from signing to go-live: 94 days; the contract promises 45. Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year. 3 churned customers moved to a competitor with robotics integrations; robotics was cited in 9 lost new-logo deals. We don't know when in the year churned customers decided to leave.
The call plan12 calls: 6 with customers who churned in the last 9 months, 6 with at-risk customers (health score red). On each account we want to talk to the warehouse operations manager who used Crate every day and, separately, the VP of Operations or COO who signed the contract. Ana wants to run 4 of the calls with former customers she knows personally.
Notes from Customer Success and Legal5 of the 6 at-risk accounts are in renewal talks, and Customer Success is offering them discounts. Interviewers must not discuss pricing, discounts or what's on our roadmap. Every call must start by asking permission to record.
What a strong answer does

A guide that tests both theories fairly and leaves room for a third: the goals ask what actually drove each decision to leave or doubt, when it happened, and who made it, with questions that could disprove robotics (did they automate, did robots come up before they decided, what did they switch to and why) and implementation (how long go-live took, what happened in those months, whether value arrived) alike. It treats the exit survey as weak evidence (single choice, integrations listed first, 'Other' left blank) and the 2.4× figure as a correlation to probe, not a cause. The core is a timeline of the account: the story from signing to the moment they decided to leave or started doubting, with sequential probes. It has separate tracks for the daily user (what went wrong day to day) and the signer (how the decision was made, what alternatives they weighed), fits 45 minutes with timings, and opens with consent to record. It guards the calls: no pitching, pricing, discounts or roadmap promises, with a script for when an at-risk customer asks, and it addresses Ana's bias: she has a public view and personal relationships, so she should observe or take notes rather than lead, or at minimum use the same script and have her calls read separately. It ends with decision signals: what pattern across the 12 would back Bet A, what would back Bet B, and what would back neither (price, a business change, a missing capability that isn't robotics).

Critical failures (cap the score)
  • Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks
  • Asks customers to choose between robotics and implementation as the reason they left
Case

v1.0 · synthetic · B2B SaaS, logistics, churn, exec bias, staff-level

Needs more than a quick edit10 of 13 checks passed

The guide lacks separate question tracks for the daily user and the decision-maker, and it doesn't prioritize questions for time management. A PM would need to add role-specific questions and mark must-asks, which is more than a quick edit.

Output

Crate customer research guide: why are warehouses leaving?

Purpose: In three weeks, learn whether churn is more plausibly driven by missing robotics integrations (Bet A), slow implementation (Bet B), or something else. These calls are for understanding customers’ experiences—not selling, negotiating renewals, or validating either executive’s theory.

Learning goals

For each account, establish:

  • What happened, and when? Separate implementation, go-live, first value, automation needs, and the point when leaving became likely or decided.
  • What mattered most? Distinguish a decisive problem from a feature request, contributing frustration, or reason offered after the fact.
  • What evidence supports the explanation? Look for actual workflow impact, workarounds, alternatives evaluated, and actions taken—not just opinions or hypotheticals.
  • What might change the outcome? Explore credible counterfactuals without promising anything.
  • What else explains churn? Surface causes beyond A and B.

Recommended plan changes

The proposed 12 calls cannot cover both the daily user and contract signer on all 12 accounts: that would require 24 interviews. Keep the 12-call cap, but interview two roles separately at each of six accounts: three churned accounts and three red-health accounts, with one warehouse operations manager and one VP of Operations/COO per account. This gives paired perspectives and preserves both customer situations, but is a small, directional sample—not a prevalence estimate. Select accounts for varied implementation times and automation situations where possible; don’t choose only accounts with known robotics issues.

Do not have Ana conduct four calls with former customers she knows personally. Her involvement risks courtesy bias and leading the conversation. Use an independent interviewer; ideally don’t include personal contacts in this small core sample. If one is included, disclose the relationship, and have Ana neither attend nor receive attributable notes.

Because five at-risk accounts are in renewal talks and receiving discount offers, use an interviewer outside the account/renewal team. Tell participants their answers won’t affect service or renewal discussions. Don’t share interview content with the account team in a way that could be used in negotiation. Follow Legal’s recording requirement below.

45-minute guide: churned customers

TimeQuestions
0–3Start by asking: “May I record this conversation?” If no, take notes instead. Explain the purpose, that there are no sales or renewal implications, and that we won’t discuss pricing, discounts, or roadmap.
3–7“What was your role in choosing and using Crate?” “What did your warehouse need the system to help you do?”
7–18“Take me through your experience with Crate, from signing through the decision to leave.” “When did you first consider switching? What happened just before that?” “Who was involved, and when was the decision effectively made?”
18–27“How did implementation and go-live go?” “Where, if anywhere, did it take longer or require more effort than expected?” “What was the impact on your operation?” “When did you first get value from Crate?”
27–35“Did automation or robotics play a role in your warehouse during this period?” If yes: “What equipment or workflow?” “How did it work with Crate?” “What did you do when it didn’t?” “When did this become a problem relative to the decision to leave?”
35–41“What alternatives did you consider, and what drove the choice?” “What, if anything, could realistically have changed your decision?” “What was the most important reason—not just one frustration?”
41–45“What haven’t I asked that matters?” “Is there anything we should correct about our understanding?” Thank them.

45-minute guide: red-health customers

Use the same opening and close. Do not imply that Crate can fix an issue or influence renewal terms.

TimeQuestions
0–3Start by asking: “May I record this conversation?” If no, take notes. Explain that this is research, not a sales or renewal conversation; participation won’t affect service or renewal discussions.
3–7“What is your role in using or overseeing Crate?” “What does a successful warehouse operation look like for you?”
7–17“How has your experience with Crate unfolded so far?” “When have things gone better or worse than expected?” “What is working well today?”
17–27“Where are you in implementation and go-live?” “What has taken the most effort or time?” “What value, if any, have you realized so far?” “What is the operational impact of anything that remains difficult?”
27–35“How, if at all, does automation or robotics fit into your operation?” If relevant: “What equipment or workflow?” “How does it work with Crate today?” “What workarounds do you use?” “When did this become important?”
35–41“Have you considered alternatives to Crate? What prompted that, if so?” “What would need to be true for Crate to work well for your operation?” “What are the most important unresolved issues?”
41–45“What haven’t I asked that matters?” “Is there anything we should correct?” Thank them.

Guidance for interviewers

  • Ask the questions as written, but follow the participant’s timeline. Use neutral probes: “What happened next?” “Can you give me an example?” “How do you know?”
  • Don’t offer Ana’s or Marcus’s theories, mention the exit-survey results, or ask leading questions such as “Would robots have saved you?” Ask about events before asking for interpretations.
  • Separate when an issue occurred from when the participant says it mattered. Probe for specific actions and evidence: evaluations, workarounds, escalations, or a switch decision.
  • Don’t discuss pricing, discounts, renewal negotiations, or roadmap. If asked about roadmap, say you can’t speak to future plans. Don’t make promises or relay an interview answer as a renewal signal.
  • Interview roles separately. Record role and account context in notes, but distinguish direct experience from second-hand views. Afterward, write a short timeline, leading cause(s), evidence, and uncertainties for each account.
  • Treat recordings and notes as confidential; use only with consent and Legal-approved handling. Don’t present anecdotes as representative of all customers.

What would support each bet?

Support Bet A (robotics integrations): In multiple independent accounts, a specific robotics workflow or integration gap caused meaningful operational friction, and that issue arose before the decision to leave or seriously consider leaving. Stronger evidence: an actual robot deployment or concrete plan, a costly workaround, and a competitor selected or evaluated for this capability. A feature mention alone—or robotics cited in a new-logo deal—doesn’t establish that it drove existing-customer churn.

Support Bet B (faster implementation): In multiple accounts, implementation delays or effort materially postponed go-live or first value, caused operational consequences, and preceded the decision to leave or consider leaving. Stronger evidence: specific preventable blockers, an unmet expectation, and a credible account that earlier go-live would have changed the outcome. The 94-day median, 45-day promise, and 2.4× association make this worth testing, but do not prove causation.

Support neither / remain uncertain: The call evidence points more strongly to another cause; A or B appears only after the decision; accounts describe issues without meaningful consequences; or evidence is mixed and no cause consistently precedes churn. Don’t force a winner from 12 interviews. Compare patterns across accounts, not individual interview counts, and triangulate with CRM timelines and other churn evidence. The exit survey’s single-choice format, first-listed feature option, and sparse free text limit what it can establish.

Check by check

Got wrong · 2
  • Different questions for user and signerThe guide does not provide distinct question tracks for the daily warehouse operations manager and the executive who signed; it only differentiates churned vs at-risk customers, with the same questions for both roles within each track.
  • Fits the callAlthough timed sections add up to 45 minutes, the guide does not mark must-ask questions or say what to cut if time runs short, so an interviewer could run out of time without covering the most critical probes.
Mixed · 1
  • Addresses the actual decisionThe guide clearly states what evidence from the calls would back Bet A, Bet B, or neither, which is the decision framework the exec team needs.The two graders disagreed on this one.
Got right · 10
  • Uses the supplied evidence correctlyEvery statement about the current situation is taken directly from the supplied context, with no invented facts.
  • Respects explicit constraintsThe guide respects all constraints: it includes learning goals, timed questions for both customer types, interviewer guidance, plan changes, and decision signals; it stays under 1,500 words; it enforces no pricing/roadmap discussion and recording consent.
  • Identifies material uncertaintyIt names key unknowns (exit survey weakness, correlation vs causation, timing of churn decision) and says how the calls will resolve them through patterns and timelines.
  • Avoids unsupported claimsInterpretations like the exit survey's limitations and the 2.4× association are clearly labeled as not proving causation, not presented as established fact.
  • Produces the required deliverableThe output is a complete call guide with all requested sections, written for the exec team, and well under 1,500 words.
  • Tests both theories fairlyBoth the robotics and implementation theories get questions that could disprove them, open timeline questions come first, and the guide explicitly leaves room for a third cause.
  • Protects the calls and the accountsIt gives a clear rule and script for pricing/roadmap questions, starts with recording consent, and addresses Ana's bias by recommending she not lead calls and, if a contact is included, that she not attend or receive attributable notes.
  • Designed to falsifyEvery learning goal has questions whose honest answers could disprove the hypothesis, such as asking when automation became a problem relative to the decision to leave.
  • Asks what people did, not what they thinkCore questions ask for specific past events and actions (e.g., 'When did you first consider switching? What happened just before that?'), with opinions and predictions clearly secondary.
  • Doesn't lead the witnessQuestions are neutral and open; the guide avoids mentioning either executive's theory until after the participant's own timeline, and prohibits leading probes.

Grades and run details

Decision model 85 · LLM judge 11 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly81%
  • failAddresses the actual decision50%
  • passRespects explicit constraints26%
  • passIdentifies material uncertainty95%
  • passAvoids unsupported claims91%
  • passProduces the required deliverable73%
  • passTests both theories fairly99%
  • partialDifferent questions for user and signer74%
  • passProtects the calls and the accounts79%
  • passDesigned to falsify75%
  • passAsks what people did, not what they think70%
  • passDoesn't lead the witness55%
  • partialFits the call29%
Run
Run
#1
API response time
29 s
Submitted
30 Sept 2026

Usable with a quick edit11 of 13 checks passed

The guide is thorough, respects all constraints, and provides actionable scripts and decision criteria. A PM could use it with only minor wording tweaks.

Output

Crate churn calls: interview guide

Purpose: Decide between Bet A (robotics integrations) and Bet B (faster implementation) by learning why customers actually leave. Neither theory is proven yet.

Why we're not already sure

  • Exit survey: "Missing features or integrations" bundles two things, isn't robotics-specific, was listed first, and has only 41 responses. People also tend to name the most recent or most socially easy reason rather than the first cause.
  • Robotics: 3 of roughly 60 churned accounts went to a robotics competitor. The 9 lost deals are real but are new-logo losses, not churn.
  • Implementation: The 2.4× churn gap after 90+ days is a correlation. Complex sites may be both slow to launch and likely to leave. We also don't know when customers decided to leave.
  • Ana's and Marcus's evidence: Each is anecdotal and selected. Ana hears from customers who will talk to the CEO.
  • The two theories may be linked. Robotics-heavy sites may be the slow implementations.

Learning goals

  1. When and why each account first considered leaving, relative to go-live.
  2. Whether automation was a cause, a rationalization, or irrelevant. The test is whether they took action (bought robots, evaluated replacements), not whether they mention it.
  3. Whether slow go-live caused the exit, and whether the delay was Crate's doing or the customer's (their data, staffing, customizations).
  4. What else drives exits: price, support, core product gaps, or events outside our control, such as a 3PL losing a shipper contract or being acquired.
  5. Whether the two causes overlap.

Changes to the plan

  • Call math. Two roles per account means 12 calls covers only 6 accounts. I'd use 8 accounts: 4 churned accounts with both roles (8 calls), and 4 at-risk accounts with the operations manager only (4 calls). At-risk executives are the people in renewal negotiations, which is where contamination risk is highest. Comparing what the signer and the daily user each say is itself valuable.
  • Recruit on criteria, not convenience. Among churned accounts, aim for 2 that left within 12 months of signing and 2 that left later. Include go-lives both over and under 90 days, at least 2 sites with robots or carts, and at least 1 that chose "price" in the survey. Pull this from CRM and Customer Success notes this week.
  • Ana doesn't run calls. Customers who know the CEO will be polite, and she has a stated hypothesis. Her relationships are valuable for introductions, and at most 1–2 of her contacts should be in the churned set. Someone with no stake in either bet runs the calls, with a second person taking notes. Ana and Marcus get recordings and summaries afterward.
  • At-risk accounts need Customer Success coordination. CS confirms no discount conversation is happening that week. The invitation states that this is research, separate from renewal. If a customer tries to negotiate, redirect and tell CS afterward. Weight at-risk answers less, since they describe intentions, not decisions.
  • Pre-register. Before call one, Ana and Marcus each read the last section and write what would change their mind.
  • Run a parallel data pull. Get tenure at churn, go-live days, and robotics on site for all roughly 60 churned accounts. If the calls and the data disagree, the data wins.

Guidance for whoever runs the calls

  • Consent first. The first thing you say after your name is a request to record. If they decline, take notes only and continue.
  • Off-limits: pricing, discounts, and roadmap. If they raise price, listen and ask what they compared us with and what they got, but no numbers or offers. If asked "will you build X?", say: "I can't speak to plans. I'm here to understand your experience." If they ask for a commercial conversation, point them to their account manager.
  • Do not say "robotics" or "implementation" until the prompted section. Record whether each topic came up unprompted or only when asked. These carry very different weight.
  • Alternate the order of the two prompted topics from call to call.
  • Ask for stories and dates, not opinions: "Tell me about the last time," "What happened next?" Favor actions ("Who did you call? What did you buy?") over stated reasons.
  • Don't defend Crate or fix problems. Answer criticism with "Say more about that."
  • Use silence. Wait a few seconds after an answer; the real reason often comes next.
  • Debrief within 24 hours on a one-page sheet: decision date, tenure, go-live days, robots on site, trigger event, unprompted causes, prompted causes, strongest quote, and anything that contradicts our hypotheses.

Guide: Operations manager (daily user), 45 min

TimeSectionQuestions
0:00–0:03OpenAsk permission to record. "I'm [name] from Crate's product team. I'm here to learn, not sell, and nothing you say affects your account. I can't discuss pricing or plans."
0:03–0:08ContextTell me about your site and your role. What does a normal day look like? What equipment and systems run alongside Crate?
0:08–0:20Start-up storyTake me back to when you started with Crate. What happened between signing and running real orders? What got delayed, and why? Who was involved on each side? When did it first feel like it was working?
0:20–0:32Doubt timelineWhen did you first wonder if Crate was the right system? Where were you and what happened? What did you do next, and who did you talk to? (Churned: what was the last straw? At-risk: have you looked at alternatives, and what prompted that?)
0:32–0:39Prompted topicsAutomation: Do you use, or plan to use, robots, autonomous carts, or pick-assist? What's in place, and when did it arrive? How did it work with Crate, and what did you do about it? Setup: Looking back, how did the time to get live affect your team or your view of Crate? (Skip anything already covered in depth.)
0:39–0:43CounterfactualWhat one thing would have kept you? (At-risk: what would need to change for you to be confident staying?) Has anything you've seen elsewhere worked better?
0:43–0:45CloseWhat haven't I asked that I should have? Who else should I talk to? Thank them.

Guide: VP Operations / COO (signer), 45 min

TimeSectionQuestions
0:00–0:03OpenSame as above. Ask permission to record first.
0:03–0:08ContextTell me about your business and your customers. How has the past year changed things for you?
0:08–0:18BuyingWhy did you choose Crate? What did you hope would be different after 6–12 months? What were you told about go-live, and what happened? How did you know whether it was working?
0:18–0:30DecisionWalk me through the decision to leave (or to reconsider). When did it first come up, and who raised it? What did you look at? Which alternatives did you consider? Who made the final call, and when? What would have changed the outcome?
0:30–0:38Prompted topicsAutomation: What's your automation plan over the next 2–3 years? What's happened so far? How did it figure in the decision? Setup: How did the time to go live figure in your thinking, if at all? (Skip anything already covered in depth.)
0:38–0:43Business contextWhat else changed around then: shipper contracts won or lost, sites opened or closed, leadership, budget, acquisitions? If you ranked everything that mattered, what's first? What do you use now, and why?
0:43–0:45CloseAnything I missed? Thank them.

What we'd need to hear

Eight accounts can't produce percentages. We're looking for a pattern strong enough to act on.

Back Bet A (robotics) if, in at least 3 of the 4 churned accounts: - Automation comes up unprompted, or with only light prompting, as a reason. - They acted on it: robots bought, piloted, or contracted before they decided to leave. - They were live and broadly content with Crate beforehand. - They evaluated or moved to something that connects to their robots, and say they'd likely have stayed if we had. - The same story appears in at least 2 of the 4 at-risk accounts.

Back Bet B (implementation) if, in at least 3 of the 4 churned accounts: - Their first serious doubt came during or shortly after a late go-live, typically within the first year. - They describe never reaching value, or losing their internal champion during the delay. - The delay was mainly Crate's doing (process, handoffs, configuration), not their own data or staffing. - They would say faster go-live would probably have kept them.

Neither bet if: - Exits are mainly driven by events such as a lost shipper contract, site closure, acquisition, or price. - Robotics is mentioned but no one acted on it. - Go-live was slow but they were happy afterward. - Reasons scatter with no pattern. - Core product gaps, such as reporting or inventory accuracy, show up repeatedly. That would point to a third bet we haven't framed.

Both: If the same accounts show robotics-heavy sites with long go-lives because integrations took months, that is a third reading. It favors Bet A if the cause is missing integrations, and Bet B if it's our process. Decide by asking what the customer says was the blocker.

Split results: If churned accounts split evenly, or contradict the data pull, say so. Recommend a targeted follow-up, such as segmenting churn by tenure, rather than forcing a choice.

Check by check

Mixed · 2
  • Uses the supplied evidence correctlyAll factual claims about the current situation are directly supported by the brief or derived correctly from its numbers.The two graders disagreed on this one.
  • Addresses the actual decisionThe output provides clear, actionable criteria for what would back each bet or neither after the calls, exactly as the brief requests.The two graders disagreed on this one.
Got right · 11
  • Respects explicit constraintsThe guide respects all constraints: it stays under 1,500 words, includes all requested sections, and provides concrete scripts and plans to enforce the off-limits topics and consent.
  • Identifies material uncertaintyThe output explicitly names the key unknowns (weak survey, correlation vs causation, unknown decision timing) and says how the calls and a parallel data pull would resolve them.
  • Avoids unsupported claimsHypotheses and interpretations are clearly labelled as such, and no cause or forecast is presented as established fact.
  • Produces the required deliverableThe guide is complete with learning goals, timed questions for two roles, interviewer guidance, plan changes, and decision signals, all under 1,500 words and ready for the exec team.
  • Tests both theories fairlyBoth theories get questions that could disprove them (e.g., no action on robotics, happiness after slow go-live), and open-ended timeline questions leave room for a third cause.
  • Different questions for user and signerThe guide has distinct tracks: daily user questions focus on day-to-day experience and workarounds, signer questions on the decision process and alternatives.
  • Protects the calls and the accountsIt includes a clear off-limits rule with scripts for pricing/roadmap, starts with consent, and addresses the CEO's bias by having her not lead calls and limiting her contacts.
  • Designed to falsifyEvery learning goal has a disconfirming question, such as asking what they did about robots or whether they ever reached value after a slow go-live.
  • Asks what people did, not what they thinkCore questions ask for specific recent actions and events ('What happened next?', 'Who did you call?'), not opinions or predictions.
  • Doesn't lead the witnessQuestions are neutral and open-ended; the prompted topics are introduced without pitching or hinting at the desired answer.
  • Fits the callBoth 45-minute guides have timed sections that sum correctly, a realistic number of questions, and instructions to skip covered topics if time is short.

Grades and run details

Decision model 77 · LLM judge 14 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly16%
  • failAddresses the actual decision16%
  • partialRespects explicit constraints15%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims75%
  • passProduces the required deliverable91%
  • passTests both theories fairly100%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts97%
  • passDesigned to falsify85%
  • passAsks what people did, not what they think82%
  • passDoesn't lead the witness49%
  • partialFits the call33%
Run
Run
#1
API response time
1.6 min
Submitted
30 Sept 2026

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.892.31None
2GPT-6.1 SolwithAPI92.192.62None
3Opus 5.5withClaude94.169.22None
4Sonnet 5.5withAPI86.484.62None
5GPT-6 LunawithAPI90.273.921 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review