Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Crate. In four weeks, the exec team decides which of two bets gets our two squads next half, and both bets rest on a theory of why customers are leaving. We have three weeks for 12 customer calls of 45 minutes. Write the guide for them. It should include: the learning goals; the questions, with timings, for the two kinds of people we'll talk to; guidance for whoever runs the calls; any changes you'd make to the plan below; and what we'd need to hear to back each bet, or neither. Keep it under 1,500 words. The exec team will read it before the calls start.

What the model was given8 items: About Crate, The decision, CEO, Ana Ferreira, CPO, Marcus Hill, Exit survey (41 churned customers, last 12 months), Implementation and CRM data, The call plan, Notes from Customer Success and Legal
About CrateWarehouse management software for mid-size third-party logistics companies (3PLs). 340 customers, $22M ARR. Gross logo churn rose from 11% to 18% over the last year.
The decisionBet A: integrations with warehouse robotics (autonomous carts, pick-assist robots). Bet B: rebuild implementation so customers go live faster. Two squads for the next half go to one of them.
CEO, Ana Ferreira“We're losing warehouses because they're automating and we don't talk to their robots. Every lost customer I've spoken to mentions it.”
CPO, Marcus Hill“We lose them in the first year because implementation takes forever. They're gone before they ever see the value.”
Exit survey (41 churned customers, last 12 months)A single-choice question, 'Why are you leaving?'. Missing features or integrations: 46%. Too hard to implement: 22%. Price: 17%. Other: 15% (free-text box, left blank by all but two). 'Missing features or integrations' was the first option in the list.
Implementation and CRM dataMedian time from signing to go-live: 94 days; the contract promises 45. Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year. 3 churned customers moved to a competitor with robotics integrations; robotics was cited in 9 lost new-logo deals. We don't know when in the year churned customers decided to leave.
The call plan12 calls: 6 with customers who churned in the last 9 months, 6 with at-risk customers (health score red). On each account we want to talk to the warehouse operations manager who used Crate every day and, separately, the VP of Operations or COO who signed the contract. Ana wants to run 4 of the calls with former customers she knows personally.
Notes from Customer Success and Legal5 of the 6 at-risk accounts are in renewal talks, and Customer Success is offering them discounts. Interviewers must not discuss pricing, discounts or what's on our roadmap. Every call must start by asking permission to record.
What a strong answer doesThe answer key the graders mark against

A guide that tests both theories fairly and leaves room for a third: the goals ask what actually drove each decision to leave or doubt, when it happened, and who made it, with questions that could disprove robotics (did they automate, did robots come up before they decided, what did they switch to and why) and implementation (how long go-live took, what happened in those months, whether value arrived) alike. It treats the exit survey as weak evidence (single choice, integrations listed first, 'Other' left blank) and the 2.4× figure as a correlation to probe, not a cause. The core is a timeline of the account: the story from signing to the moment they decided to leave or started doubting, with sequential probes. It has separate tracks for the daily user (what went wrong day to day) and the signer (how the decision was made, what alternatives they weighed), fits 45 minutes with timings, and opens with consent to record. It guards the calls: no pitching, pricing, discounts or roadmap promises, with a script for when an at-risk customer asks, and it addresses Ana's bias: she has a public view and personal relationships, so she should observe or take notes rather than lead, or at minimum use the same script and have her calls read separately. It ends with decision signals: what pattern across the 12 would back Bet A, what would back Bet B, and what would back neither (price, a business change, a missing capability that isn't robotics).

Critical failures (cap the score at 40)
  • Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks
  • Asks customers to choose between robotics and implementation as the reason they left
Case

Staff-level · v1.1 · synthetic · B2B SaaS, logistics, churn, exec bias, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Critical failure: Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks. Score capped at 40.

Needs more than a quick edit3 of 13 checks passed

A PM would need to fix the prohibited either/or question, add a pricing/roadmap handling script, remove unsupported claims, and clarify the revised call coverage before using it. The reasoning and evidence also need rechecking.

Output

Customer Research Guide: H2 Strategic Bets (Robotics vs. Implementation)

To: Executive Team From: Staff Product Manager, Crate Date: Today Subject: 3-Week Customer Research Plan for H2 Resource Allocation (12 Calls)

---

Executive Summary & Plan Adjustments

In four weeks, the executive team will decide where to allocate both engineering squads for the next half. Ana believes we are losing customers because we lack warehouse robotics integrations (Bet A). Marcus believes we are losing them because our implementation cycle is broken and slow (Bet B).

The exit survey heavily favors Ana’s view (46% cite missing features/integrations), but UX research 101 teaches us that exit surveys are structurally flawed. Listing "missing features" first creates primacy bias, and the survey failed to capture temporal dynamics—specifically, when the decision to leave was made. Meanwhile, our internal data tells a damning story: our median time to go-live is 94 days (vs. 45 promised), and customers crossing 90 days churn at $2.4\times$ the rate.

We have 3 weeks to conduct 12 interviews (45 minutes each) to uncover the truth.

Critical Plan Adjustments (Read Before We Start)

  1. Split the Interviewees by Persona, Not Account: The original plan called for interviewing both the Ops Manager and the VP/COO on the same accounts. In 45 minutes, splitting time between a daily user and a high-level signer dilutes the insight. Instead, we will interview Warehouse Ops Managers to understand day-to-day friction (Implementation, UI, operational bottlenecks) and VPs of Ops / COOs to understand strategic drivers (ROI, contract expectations, robotics, and when the churn decision was actually made).
  2. Revise Ana’s Participation: Ana wants to run 4 calls with former customers she knows personally. We must advise against this. Founder-led interviews with churned customers carry massive politeness bias; customers will pull punches about implementation failures or invent reasons to let Ana down gently. Ana should observe these calls anonymously or listen to the recordings, but an unbiased PM or User Researcher must conduct them.
  3. Guardrails for At-Risk Accounts: Five of our six at-risk accounts are in active renewal talks with CS discounts on the table. Interviewers are strictly prohibited from discussing pricing, commercial concessions, or our product roadmap. We are strictly here to diagnose reality.

---

Learning Goals

  1. The Churn Timeline: Determine the exact month the customer mentally decided to leave. Did they decide during a painful 90+ day onboarding (supporting Bet B), or later when evaluating modernizing warehouse automation (supporting Bet A)?
  2. The "Missing Features" Audit: Deconstruct the 46% exit survey metric. When respondents checked "missing features/integrations," were they missing critical WMS functionality (e.g., wave planning, inventory accuracy) or specifically robotics middleware?
  3. Implementation Reality Check: Measure the operational impact of our 94-day median go-live time. Did the delay create cash-flow/trust erosion that doomed the account before day one, regardless of subsequent features?
  4. Robotics Urgency: For customers who mention robotics, are they actively deploying autonomous carts today, or is it a 3-year roadmap item they use as a stick to beat us with during negotiations?

---

Interview Guide: Two Personas

Persona 1: The Daily User (Warehouse Operations Manager)

Focus: Implementation friction, daily utility, operational roadblocks.

  • 00:00–03:00 | Intro & Permission:
  • "Thanks for taking the time. I’m [Name] from Crate. We’re doing research to figure out where we failed our customers and how we can improve. I am strictly here to learn about your experience—we won't be talking about pricing, renewals, or selling you anything today. Do I have your permission to record this call for internal notes?"
  • 03:00–12:00 | Day-to-Day Workflow & Reality:
  • "Take me back to your first 30 days using Crate. What did the day-to-day onboarding look like?"
  • "How long did it actually take for your floor team to feel comfortable using Crate without holding their breath?"
  • "What was the most painful workaround your team had to invent during your first three months?"
  • 12:00–25:00 | The Implementation Experience (Bet B Probe):
  • "Our data shows it takes about 90+ days for most warehouses to go live on Crate. What was your experience? How did those delays affect your team's stress levels and your relationship with our implementation team?"
  • "Looking back, did the slow start permanently damage your trust in the software, or did you bounce back once things were running?"
  • 25:00–35:00 | Equipment & Floor Integration (Bet A Probe):
  • "What automated hardware or mobile carts are running on your floor today? How do they talk to Crate, or do they?"
  • "When a pick-assist cart or AGV moves through your warehouse, where does the process break down between the robot and Crate?"
  • 35:00–42:00 | The Breaking Point:
  • "Think about the moment you realized you wanted to leave Crate (or started looking at alternatives). What happened on the warehouse floor that week?"
  • 42:00–45:00 | Wrap-Up:
  • "If you could wave a magic wand and change one thing about your onboarding or your day-one experience with Crate, what would it be?"

---

Persona 2: The Decision Maker (VP of Operations / COO)

Focus: Strategic expectations, ROI timelines, contract promises, and robotics strategy.

  • 00:00–03:00 | Intro & Permission:
  • "Thanks for joining. I’m [Name] from Crate. We’re reviewing why customers leave us to make radical changes to our roadmap and processes. I'm not here to pitch you or talk commercial terms—just to understand your true experience. May I record this call for note-taking purposes?"
  • 03:00–10:00 | The Purchase & Promise:
  • "When you signed with Crate, what was the primary business problem you needed solved in your first 6 months?"
  • "What timeline were you given for go-live during the sales cycle, and how did reality compare?"
  • 10:00–22:00 | Implementation & Time-to-Value (Bet B Probe):
  • "For a mid-size 3PL, cash flow and client onboarding speed are everything. How did Crate's onboarding timeline impact your ability to onboard your own end-clients?"
  • "At what point in your first year did you officially consider the implementation 'failed' or deeply troubled?"
  • 22:00–34:00 | The Robotics Mandate (Bet A Probe):
  • "Let’s talk about warehouse automation. How central are autonomous carts or robotics to your strategy today versus three years from now? Are you actively running them?"
  • "When you evaluated staying vs. leaving Crate, how heavy was the absence of robotics integrations in that final calculus? Was it the straw that broke the camel's back, or a nice-to-have?"
  • 34:00–42:00 | Uncovering the Timeline:
  • "Looking back at the lifecycle of your account, when did you mentally check out and decide you wouldn't renew? Was it during implementation, or months later?"
  • 42:00–45:00 | Wrap-Up:
  • "If Crate could fix only one thing—either our onboarding speed or our hardware/robotics connectivity—which one would have saved your account?"

---

Decision Criteria: What We Need to Hear

To make our final recommendation to Ana and Marcus, we will evaluate the 12 calls against the following thresholds:

Backing Bet A (Robotics Integrations)

We back Bet A only if: 1. VPs/COOs explicitly state that the lack of robotics was the primary, active catalyst for churning (not just a grievance listed in a post-hoc survey). 2. Churned customers are actively utilizing or piloting autonomous hardware today that could not function with Crate, resulting in operational deadlocks. 3. The exit survey's 46% "missing features" bucket is proven via interviews to be overwhelmingly about hardware/robotics connectivity rather than core WMS gaps (e.g., billing, inventory tracking).

Backing Bet B (Implementation Rebuild)

We back Bet B only if: 1. Interviews reveal that the 90+ day onboarding process created an insurmountable trust deficit; customers felt abandoned and never recovered faith in the product, rendering subsequent features irrelevant. 2. Daily users report that the complexity of setup and configuration was the primary driver of operational frustration, overshadowing missing integrations. 3. Churned customers admit they mentally checked out during or immediately after the botched implementation window, long before robotics ever entered the conversation.

Backing Neither (The Pivot Option)

We back Neither if: * The qualitative data shows a split narrative where both bets are symptoms of a deeper root cause—such as core WMS unreliability, pricing/value misalignment (17% in exit survey), or poor account management—meaning that building robots or speeding up a broken implementation will simply churn customers faster down a different path.

Claims the judge couldn’t find in the brief

  • For a mid-size 3PL, cash flow and client onboarding speed are everything.

Grades and run details

Decision model 58 · LLM judge 3 of 14 checks
Decision model checks
  • partialUses the supplied evidence correctly23%
  • passAddresses the actual decision93%
  • partialRespects explicit constraints21%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims46%
  • partialProduces the required deliverable26%
  • partialTests both theories fairly47%
  • passDifferent questions for user and signer98%
  • partialProtects the calls and the accounts84%
  • partialDesigned to falsify40%
  • partialAsks what people did, not what they think52%
  • failDoesn't lead the witness30%
  • partialMarks what to cut if the call runs over45%
Run
Run
#1
Time to output
11 s
Submitted
1 Oct 2026

Usable with a quick edit12 of 13 checks passed

The guide is thorough, respects all constraints, and provides actionable scripts and decision criteria. A PM could use it with only minor wording tweaks.

Output

Crate churn calls: interview guide

Purpose: Decide between Bet A (robotics integrations) and Bet B (faster implementation) by learning why customers actually leave. Neither theory is proven yet.

Why we're not already sure

  • Exit survey: "Missing features or integrations" bundles two things, isn't robotics-specific, was listed first, and has only 41 responses. People also tend to name the most recent or most socially easy reason rather than the first cause.
  • Robotics: 3 of roughly 60 churned accounts went to a robotics competitor. The 9 lost deals are real but are new-logo losses, not churn.
  • Implementation: The 2.4× churn gap after 90+ days is a correlation. Complex sites may be both slow to launch and likely to leave. We also don't know when customers decided to leave.
  • Ana's and Marcus's evidence: Each is anecdotal and selected. Ana hears from customers who will talk to the CEO.
  • The two theories may be linked. Robotics-heavy sites may be the slow implementations.

Learning goals

  1. When and why each account first considered leaving, relative to go-live.
  2. Whether automation was a cause, a rationalization, or irrelevant. The test is whether they took action (bought robots, evaluated replacements), not whether they mention it.
  3. Whether slow go-live caused the exit, and whether the delay was Crate's doing or the customer's (their data, staffing, customizations).
  4. What else drives exits: price, support, core product gaps, or events outside our control, such as a 3PL losing a shipper contract or being acquired.
  5. Whether the two causes overlap.

Changes to the plan

  • Call math. Two roles per account means 12 calls covers only 6 accounts. I'd use 8 accounts: 4 churned accounts with both roles (8 calls), and 4 at-risk accounts with the operations manager only (4 calls). At-risk executives are the people in renewal negotiations, which is where contamination risk is highest. Comparing what the signer and the daily user each say is itself valuable.
  • Recruit on criteria, not convenience. Among churned accounts, aim for 2 that left within 12 months of signing and 2 that left later. Include go-lives both over and under 90 days, at least 2 sites with robots or carts, and at least 1 that chose "price" in the survey. Pull this from CRM and Customer Success notes this week.
  • Ana doesn't run calls. Customers who know the CEO will be polite, and she has a stated hypothesis. Her relationships are valuable for introductions, and at most 1–2 of her contacts should be in the churned set. Someone with no stake in either bet runs the calls, with a second person taking notes. Ana and Marcus get recordings and summaries afterward.
  • At-risk accounts need Customer Success coordination. CS confirms no discount conversation is happening that week. The invitation states that this is research, separate from renewal. If a customer tries to negotiate, redirect and tell CS afterward. Weight at-risk answers less, since they describe intentions, not decisions.
  • Pre-register. Before call one, Ana and Marcus each read the last section and write what would change their mind.
  • Run a parallel data pull. Get tenure at churn, go-live days, and robotics on site for all roughly 60 churned accounts. If the calls and the data disagree, the data wins.

Guidance for whoever runs the calls

  • Consent first. The first thing you say after your name is a request to record. If they decline, take notes only and continue.
  • Off-limits: pricing, discounts, and roadmap. If they raise price, listen and ask what they compared us with and what they got, but no numbers or offers. If asked "will you build X?", say: "I can't speak to plans. I'm here to understand your experience." If they ask for a commercial conversation, point them to their account manager.
  • Do not say "robotics" or "implementation" until the prompted section. Record whether each topic came up unprompted or only when asked. These carry very different weight.
  • Alternate the order of the two prompted topics from call to call.
  • Ask for stories and dates, not opinions: "Tell me about the last time," "What happened next?" Favor actions ("Who did you call? What did you buy?") over stated reasons.
  • Don't defend Crate or fix problems. Answer criticism with "Say more about that."
  • Use silence. Wait a few seconds after an answer; the real reason often comes next.
  • Debrief within 24 hours on a one-page sheet: decision date, tenure, go-live days, robots on site, trigger event, unprompted causes, prompted causes, strongest quote, and anything that contradicts our hypotheses.

Guide: Operations manager (daily user), 45 min

TimeSectionQuestions
0:00–0:03OpenAsk permission to record. "I'm [name] from Crate's product team. I'm here to learn, not sell, and nothing you say affects your account. I can't discuss pricing or plans."
0:03–0:08ContextTell me about your site and your role. What does a normal day look like? What equipment and systems run alongside Crate?
0:08–0:20Start-up storyTake me back to when you started with Crate. What happened between signing and running real orders? What got delayed, and why? Who was involved on each side? When did it first feel like it was working?
0:20–0:32Doubt timelineWhen did you first wonder if Crate was the right system? Where were you and what happened? What did you do next, and who did you talk to? (Churned: what was the last straw? At-risk: have you looked at alternatives, and what prompted that?)
0:32–0:39Prompted topicsAutomation: Do you use, or plan to use, robots, autonomous carts, or pick-assist? What's in place, and when did it arrive? How did it work with Crate, and what did you do about it? Setup: Looking back, how did the time to get live affect your team or your view of Crate? (Skip anything already covered in depth.)
0:39–0:43CounterfactualWhat one thing would have kept you? (At-risk: what would need to change for you to be confident staying?) Has anything you've seen elsewhere worked better?
0:43–0:45CloseWhat haven't I asked that I should have? Who else should I talk to? Thank them.

Guide: VP Operations / COO (signer), 45 min

TimeSectionQuestions
0:00–0:03OpenSame as above. Ask permission to record first.
0:03–0:08ContextTell me about your business and your customers. How has the past year changed things for you?
0:08–0:18BuyingWhy did you choose Crate? What did you hope would be different after 6–12 months? What were you told about go-live, and what happened? How did you know whether it was working?
0:18–0:30DecisionWalk me through the decision to leave (or to reconsider). When did it first come up, and who raised it? What did you look at? Which alternatives did you consider? Who made the final call, and when? What would have changed the outcome?
0:30–0:38Prompted topicsAutomation: What's your automation plan over the next 2–3 years? What's happened so far? How did it figure in the decision? Setup: How did the time to go live figure in your thinking, if at all? (Skip anything already covered in depth.)
0:38–0:43Business contextWhat else changed around then: shipper contracts won or lost, sites opened or closed, leadership, budget, acquisitions? If you ranked everything that mattered, what's first? What do you use now, and why?
0:43–0:45CloseAnything I missed? Thank them.

What we'd need to hear

Eight accounts can't produce percentages. We're looking for a pattern strong enough to act on.

Back Bet A (robotics) if, in at least 3 of the 4 churned accounts: - Automation comes up unprompted, or with only light prompting, as a reason. - They acted on it: robots bought, piloted, or contracted before they decided to leave. - They were live and broadly content with Crate beforehand. - They evaluated or moved to something that connects to their robots, and say they'd likely have stayed if we had. - The same story appears in at least 2 of the 4 at-risk accounts.

Back Bet B (implementation) if, in at least 3 of the 4 churned accounts: - Their first serious doubt came during or shortly after a late go-live, typically within the first year. - They describe never reaching value, or losing their internal champion during the delay. - The delay was mainly Crate's doing (process, handoffs, configuration), not their own data or staffing. - They would say faster go-live would probably have kept them.

Neither bet if: - Exits are mainly driven by events such as a lost shipper contract, site closure, acquisition, or price. - Robotics is mentioned but no one acted on it. - Go-live was slow but they were happy afterward. - Reasons scatter with no pattern. - Core product gaps, such as reporting or inventory accuracy, show up repeatedly. That would point to a third bet we haven't framed.

Both: If the same accounts show robotics-heavy sites with long go-lives because integrations took months, that is a third reading. It favors Bet A if the cause is missing integrations, and Bet B if it's our process. Decide by asking what the customer says was the blocker.

Split results: If churned accounts split evenly, or contradict the data pull, say so. Recommend a targeted follow-up, such as segmenting churn by tenure, rather than forcing a choice.

Grades and run details

Decision model 85 · LLM judge 14 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly13%
  • passAddresses the actual decision94%
  • partialRespects explicit constraints19%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims71%
  • passProduces the required deliverable91%
  • passTests both theories fairly100%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts96%
  • passDesigned to falsify82%
  • passAsks what people did, not what they think80%
  • passDoesn't lead the witness42%
  • partialMarks what to cut if the call runs over60%
Run
Run
#1
API response time
1.6 min
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 10

Uses the supplied evidence correctlyWrongMixed
Gemini 3.5 Flash-Lite · Gemini

It uses most supplied numbers correctly, but invents or overstates current-situation claims such as “cash flow and client onboarding speed are everything” and “90+ days for most warehouses” when the context only gives a 94-day median.

Sonnet 5.5 · API

All factual claims about the current situation are directly supported by the brief or derived correctly from its numbers.

Respects explicit constraintsWrongRight
Gemini 3.5 Flash-Lite · Gemini

It is under 1,500 words and includes the required sections, but it does not enforce the no-pricing/no-roadmap rule in practice because it lacks a script for when customers raise those topics and even mentions “roadmap” in the executive intro.

Sonnet 5.5 · API

The guide respects all constraints: it stays under 1,500 words, includes all requested sections, and provides concrete scripts and plans to enforce the off-limits topics and consent.

Avoids unsupported claimsWrongRight
Gemini 3.5 Flash-Lite · Gemini

It presents several interpretations as established facts, including that the exit survey “heavily favors” Ana’s view, that founder-led interviews carry “massive” politeness bias, and that cash flow/onboarding speed are “everything” for mid-size 3PLs.

Sonnet 5.5 · API

Hypotheses and interpretations are clearly labelled as such, and no cause or forecast is presented as established fact.

Produces the required deliverableWrongRight
Gemini 3.5 Flash-Lite · Gemini

The guide is mostly present, but it is not fully usable because it changes the plan to split personas without specifying how the 12 calls and 6 accounts will be covered, and it omits a required handling script for pricing/roadmap questions.

Sonnet 5.5 · API

The guide is complete with learning goals, timed questions for two roles, interviewer guidance, plan changes, and decision signals, all under 1,500 words and ready for the exec team.

Tests both theories fairlyWrongRight
Gemini 3.5 Flash-Lite · Gemini

It gives both theories probes, but the daily-user track is weighted toward implementation, the signer track asks about robotics only after implementation, and the wrap-up asks which bet would have saved the account, so it does not test both theories fairly.

Sonnet 5.5 · API

Both theories get questions that could disprove them (e.g., no action on robotics, happiness after slow go-live), and open-ended timeline questions leave room for a third cause.

Protects the calls and the accountsWrongRight
Gemini 3.5 Flash-Lite · Gemini

It includes consent to record and a no-pricing/no-roadmap rule, but it provides no script for handling pricing, discount, or roadmap questions and the executive intro mentions roadmap changes.

Sonnet 5.5 · API

It includes a clear off-limits rule with scripts for pricing/roadmap, starts with consent, and addresses the CEO's bias by having her not lead calls and limiting her contacts.

Designed to falsifyWrongRight
Gemini 3.5 Flash-Lite · Gemini

Not every learning goal has a neutral disconfirming question; for example, the implementation goal asks whether delays “permanently damaged trust,” and the robotics goal asks whether robotics was “the straw that broke the camel's back,” rather than first eliciting the decision story.

Sonnet 5.5 · API

Every learning goal has a disconfirming question, such as asking what they did about robots or whether they ever reached value after a slow go-live.

Asks what people did, not what they thinkWrongRight
Gemini 3.5 Flash-Lite · Gemini

Some questions ask about recent actions, but the core includes many opinions, hypotheticals, and leading prompts such as “magic wand,” “how heavy was the absence of robotics,” and “which one would have saved your account.”

Sonnet 5.5 · API

Core questions ask for specific recent actions and events ('What happened next?', 'Who did you call?'), not opinions or predictions.

Doesn't lead the witnessWrongRight
Gemini 3.5 Flash-Lite · Gemini

Several questions are leading or biased, including “Our data shows it takes about 90+ days,” “did the slow start permanently damage your trust,” “cash flow and client onboarding speed are everything,” and the final either/or bet question.

Sonnet 5.5 · API

Questions are neutral and open-ended; the prompted topics are introduced without pitching or hinting at the desired answer.

Marks what to cut if the call runs overWrongRight
Gemini 3.5 Flash-Lite · Gemini

Timings add to 45 minutes, but the number of questions is too high for 45 minutes, must-ask questions are not marked, and there is no guidance on what to cut if time runs short.

Sonnet 5.5 · API

Both 45-minute guides have timed sections that sum correctly, a realistic number of questions, and instructions to skip covered topics if time is short.

All got right 3

Addresses the actual decisionRightRight
Gemini 3.5 Flash-Lite · Gemini

It does not choose a bet now, but the supplied evidence cannot support a final call; it clearly says the research will settle the decision and gives explicit thresholds for Bet A, Bet B, or neither.

Sonnet 5.5 · API

The output provides clear, actionable criteria for what would back each bet or neither after the calls, exactly as the brief requests.

Identifies material uncertaintyRightRight
Gemini 3.5 Flash-Lite · Gemini

It names the key unknowns—when customers decided to leave, whether “missing features” means robotics or core WMS, and whether slow implementation caused trust loss—and says the calls and decision thresholds would resolve them.

Sonnet 5.5 · API

The output explicitly names the key unknowns (weak survey, correlation vs causation, unknown decision timing) and says how the calls and a parallel data pull would resolve them.

Different questions for user and signerRightRight
Gemini 3.5 Flash-Lite · Gemini

It has distinct tracks: the daily user is asked about onboarding, workarounds, floor equipment, and the breaking point, while the signer is asked about purchase promises, time-to-value, robotics strategy, and renewal decision timing.

Sonnet 5.5 · API

The guide has distinct tracks: daily user questions focus on day-to-day experience and workarounds, signer questions on the decision process and alternatives.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review