Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Crate. In four weeks, the exec team decides which of two bets gets our two squads next half, and both bets rest on a theory of why customers are leaving. We have three weeks for 12 customer calls of 45 minutes. Write the guide for them. It should include: the learning goals; the questions, with timings, for the two kinds of people we'll talk to; guidance for whoever runs the calls; any changes you'd make to the plan below; and what we'd need to hear to back each bet, or neither. Keep it under 1,500 words. The exec team will read it before the calls start.

What the model was given8 items: About Crate, The decision, CEO, Ana Ferreira, CPO, Marcus Hill, Exit survey (41 churned customers, last 12 months), Implementation and CRM data, The call plan, Notes from Customer Success and Legal
About CrateWarehouse management software for mid-size third-party logistics companies (3PLs). 340 customers, $22M ARR. Gross logo churn rose from 11% to 18% over the last year.
The decisionBet A: integrations with warehouse robotics (autonomous carts, pick-assist robots). Bet B: rebuild implementation so customers go live faster. Two squads for the next half go to one of them.
CEO, Ana Ferreira“We're losing warehouses because they're automating and we don't talk to their robots. Every lost customer I've spoken to mentions it.”
CPO, Marcus Hill“We lose them in the first year because implementation takes forever. They're gone before they ever see the value.”
Exit survey (41 churned customers, last 12 months)A single-choice question, 'Why are you leaving?'. Missing features or integrations: 46%. Too hard to implement: 22%. Price: 17%. Other: 15% (free-text box, left blank by all but two). 'Missing features or integrations' was the first option in the list.
Implementation and CRM dataMedian time from signing to go-live: 94 days; the contract promises 45. Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year. 3 churned customers moved to a competitor with robotics integrations; robotics was cited in 9 lost new-logo deals. We don't know when in the year churned customers decided to leave.
The call plan12 calls: 6 with customers who churned in the last 9 months, 6 with at-risk customers (health score red). On each account we want to talk to the warehouse operations manager who used Crate every day and, separately, the VP of Operations or COO who signed the contract. Ana wants to run 4 of the calls with former customers she knows personally.
Notes from Customer Success and Legal5 of the 6 at-risk accounts are in renewal talks, and Customer Success is offering them discounts. Interviewers must not discuss pricing, discounts or what's on our roadmap. Every call must start by asking permission to record.
What a strong answer doesThe answer key the graders mark against

A guide that tests both theories fairly and leaves room for a third: the goals ask what actually drove each decision to leave or doubt, when it happened, and who made it, with questions that could disprove robotics (did they automate, did robots come up before they decided, what did they switch to and why) and implementation (how long go-live took, what happened in those months, whether value arrived) alike. It treats the exit survey as weak evidence (single choice, integrations listed first, 'Other' left blank) and the 2.4× figure as a correlation to probe, not a cause. The core is a timeline of the account: the story from signing to the moment they decided to leave or started doubting, with sequential probes. It has separate tracks for the daily user (what went wrong day to day) and the signer (how the decision was made, what alternatives they weighed), fits 45 minutes with timings, and opens with consent to record. It guards the calls: no pitching, pricing, discounts or roadmap promises, with a script for when an at-risk customer asks, and it addresses Ana's bias: she has a public view and personal relationships, so she should observe or take notes rather than lead, or at minimum use the same script and have her calls read separately. It ends with decision signals: what pattern across the 12 would back Bet A, what would back Bet B, and what would back neither (price, a business change, a missing capability that isn't robotics).

Critical failures (cap the score at 40)
  • Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks
  • Asks customers to choose between robotics and implementation as the reason they left
Case

Staff-level · v1.1 · synthetic · B2B SaaS, logistics, churn, exec bias, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit12 of 13 checks passed

The guide is thorough, respects all constraints, and provides actionable scripts and decision criteria. A PM could use it with only minor wording tweaks.

Output

Crate churn calls: interview guide

Purpose: Decide between Bet A (robotics integrations) and Bet B (faster implementation) by learning why customers actually leave. Neither theory is proven yet.

Why we're not already sure

  • Exit survey: "Missing features or integrations" bundles two things, isn't robotics-specific, was listed first, and has only 41 responses. People also tend to name the most recent or most socially easy reason rather than the first cause.
  • Robotics: 3 of roughly 60 churned accounts went to a robotics competitor. The 9 lost deals are real but are new-logo losses, not churn.
  • Implementation: The 2.4× churn gap after 90+ days is a correlation. Complex sites may be both slow to launch and likely to leave. We also don't know when customers decided to leave.
  • Ana's and Marcus's evidence: Each is anecdotal and selected. Ana hears from customers who will talk to the CEO.
  • The two theories may be linked. Robotics-heavy sites may be the slow implementations.

Learning goals

  1. When and why each account first considered leaving, relative to go-live.
  2. Whether automation was a cause, a rationalization, or irrelevant. The test is whether they took action (bought robots, evaluated replacements), not whether they mention it.
  3. Whether slow go-live caused the exit, and whether the delay was Crate's doing or the customer's (their data, staffing, customizations).
  4. What else drives exits: price, support, core product gaps, or events outside our control, such as a 3PL losing a shipper contract or being acquired.
  5. Whether the two causes overlap.

Changes to the plan

  • Call math. Two roles per account means 12 calls covers only 6 accounts. I'd use 8 accounts: 4 churned accounts with both roles (8 calls), and 4 at-risk accounts with the operations manager only (4 calls). At-risk executives are the people in renewal negotiations, which is where contamination risk is highest. Comparing what the signer and the daily user each say is itself valuable.
  • Recruit on criteria, not convenience. Among churned accounts, aim for 2 that left within 12 months of signing and 2 that left later. Include go-lives both over and under 90 days, at least 2 sites with robots or carts, and at least 1 that chose "price" in the survey. Pull this from CRM and Customer Success notes this week.
  • Ana doesn't run calls. Customers who know the CEO will be polite, and she has a stated hypothesis. Her relationships are valuable for introductions, and at most 1–2 of her contacts should be in the churned set. Someone with no stake in either bet runs the calls, with a second person taking notes. Ana and Marcus get recordings and summaries afterward.
  • At-risk accounts need Customer Success coordination. CS confirms no discount conversation is happening that week. The invitation states that this is research, separate from renewal. If a customer tries to negotiate, redirect and tell CS afterward. Weight at-risk answers less, since they describe intentions, not decisions.
  • Pre-register. Before call one, Ana and Marcus each read the last section and write what would change their mind.
  • Run a parallel data pull. Get tenure at churn, go-live days, and robotics on site for all roughly 60 churned accounts. If the calls and the data disagree, the data wins.

Guidance for whoever runs the calls

  • Consent first. The first thing you say after your name is a request to record. If they decline, take notes only and continue.
  • Off-limits: pricing, discounts, and roadmap. If they raise price, listen and ask what they compared us with and what they got, but no numbers or offers. If asked "will you build X?", say: "I can't speak to plans. I'm here to understand your experience." If they ask for a commercial conversation, point them to their account manager.
  • Do not say "robotics" or "implementation" until the prompted section. Record whether each topic came up unprompted or only when asked. These carry very different weight.
  • Alternate the order of the two prompted topics from call to call.
  • Ask for stories and dates, not opinions: "Tell me about the last time," "What happened next?" Favor actions ("Who did you call? What did you buy?") over stated reasons.
  • Don't defend Crate or fix problems. Answer criticism with "Say more about that."
  • Use silence. Wait a few seconds after an answer; the real reason often comes next.
  • Debrief within 24 hours on a one-page sheet: decision date, tenure, go-live days, robots on site, trigger event, unprompted causes, prompted causes, strongest quote, and anything that contradicts our hypotheses.

Guide: Operations manager (daily user), 45 min

TimeSectionQuestions
0:00–0:03OpenAsk permission to record. "I'm [name] from Crate's product team. I'm here to learn, not sell, and nothing you say affects your account. I can't discuss pricing or plans."
0:03–0:08ContextTell me about your site and your role. What does a normal day look like? What equipment and systems run alongside Crate?
0:08–0:20Start-up storyTake me back to when you started with Crate. What happened between signing and running real orders? What got delayed, and why? Who was involved on each side? When did it first feel like it was working?
0:20–0:32Doubt timelineWhen did you first wonder if Crate was the right system? Where were you and what happened? What did you do next, and who did you talk to? (Churned: what was the last straw? At-risk: have you looked at alternatives, and what prompted that?)
0:32–0:39Prompted topicsAutomation: Do you use, or plan to use, robots, autonomous carts, or pick-assist? What's in place, and when did it arrive? How did it work with Crate, and what did you do about it? Setup: Looking back, how did the time to get live affect your team or your view of Crate? (Skip anything already covered in depth.)
0:39–0:43CounterfactualWhat one thing would have kept you? (At-risk: what would need to change for you to be confident staying?) Has anything you've seen elsewhere worked better?
0:43–0:45CloseWhat haven't I asked that I should have? Who else should I talk to? Thank them.

Guide: VP Operations / COO (signer), 45 min

TimeSectionQuestions
0:00–0:03OpenSame as above. Ask permission to record first.
0:03–0:08ContextTell me about your business and your customers. How has the past year changed things for you?
0:08–0:18BuyingWhy did you choose Crate? What did you hope would be different after 6–12 months? What were you told about go-live, and what happened? How did you know whether it was working?
0:18–0:30DecisionWalk me through the decision to leave (or to reconsider). When did it first come up, and who raised it? What did you look at? Which alternatives did you consider? Who made the final call, and when? What would have changed the outcome?
0:30–0:38Prompted topicsAutomation: What's your automation plan over the next 2–3 years? What's happened so far? How did it figure in the decision? Setup: How did the time to go live figure in your thinking, if at all? (Skip anything already covered in depth.)
0:38–0:43Business contextWhat else changed around then: shipper contracts won or lost, sites opened or closed, leadership, budget, acquisitions? If you ranked everything that mattered, what's first? What do you use now, and why?
0:43–0:45CloseAnything I missed? Thank them.

What we'd need to hear

Eight accounts can't produce percentages. We're looking for a pattern strong enough to act on.

Back Bet A (robotics) if, in at least 3 of the 4 churned accounts: - Automation comes up unprompted, or with only light prompting, as a reason. - They acted on it: robots bought, piloted, or contracted before they decided to leave. - They were live and broadly content with Crate beforehand. - They evaluated or moved to something that connects to their robots, and say they'd likely have stayed if we had. - The same story appears in at least 2 of the 4 at-risk accounts.

Back Bet B (implementation) if, in at least 3 of the 4 churned accounts: - Their first serious doubt came during or shortly after a late go-live, typically within the first year. - They describe never reaching value, or losing their internal champion during the delay. - The delay was mainly Crate's doing (process, handoffs, configuration), not their own data or staffing. - They would say faster go-live would probably have kept them.

Neither bet if: - Exits are mainly driven by events such as a lost shipper contract, site closure, acquisition, or price. - Robotics is mentioned but no one acted on it. - Go-live was slow but they were happy afterward. - Reasons scatter with no pattern. - Core product gaps, such as reporting or inventory accuracy, show up repeatedly. That would point to a third bet we haven't framed.

Both: If the same accounts show robotics-heavy sites with long go-lives because integrations took months, that is a third reading. It favors Bet A if the cause is missing integrations, and Bet B if it's our process. Decide by asking what the customer says was the blocker.

Split results: If churned accounts split evenly, or contradict the data pull, say so. Recommend a targeted follow-up, such as segmenting churn by tenure, rather than forcing a choice.

Grades and run details

Decision model 85 · LLM judge 14 of 14 checks
Decision model checks
  • failUses the supplied evidence correctly13%
  • passAddresses the actual decision94%
  • partialRespects explicit constraints19%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims71%
  • passProduces the required deliverable91%
  • passTests both theories fairly100%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts96%
  • passDesigned to falsify82%
  • passAsks what people did, not what they think80%
  • passDoesn't lead the witness42%
  • partialMarks what to cut if the call runs over60%
Run
Run
#1
API response time
1.6 min
Submitted
30 Sept 2026

Needs more than a quick edit7 of 13 checks passed

A PM would need to shorten it, reconcile the account plan with the 12-call constraint, and remove or relabel unsupported portfolio claims before using it.

Output

Research Guide: Root Cause Analysis for H2 Squad Allocation

Prepared by: Staff Product Manager Target Audience: Executive Team (CEO, CPO, VP CS, Head of Product) Timeline: 3 weeks of interviews; executive decision in Week 4 Scope: 12 calls (45 minutes each)

---

1. Executive Summary & Recommended Changes to the Call Plan

Before executing this sprint, we must adjust three structural flaws in the proposed plan to protect data integrity:

  1. Resolve the Account-to-Call Ratio: 12 calls cannot cover two separate stakeholders across 12 accounts. We will target 6 accounts total (3 churned, 3 at-risk) and conduct 2 distinct calls per account: one with the Warehouse Operations Manager (day-to-day user) and one with the VP of Operations/COO (economic buyer). Evaluating both perspectives within the same operational context is essential to determine whether operational failure or strategic misalignment triggered the exit.
  2. Quarantine Active Renewals: 5 of the proposed 6 at-risk accounts are actively negotiating renewals and receiving discounts. Interviewing these buyers creates an immediate incentive for them to exaggerate product flaws to gain commercial leverage. We will swap 3 of these with at-risk accounts that are not currently negotiating pricing.
  3. CEO Interview Protocol: Executive presence introduces severe confirmation and deference bias—especially with personal contacts who already know Ana's passion for robotics. Recommendation: Ana should join her 4 target calls as an executive sponsor for the first 3 minutes, then hand off lead facilitation to a Product Manager, remaining on mute as an observer. If Ana leads, she must strictly follow the non-leading script provided.

---

2. Core Learning Goals

  1. Pinpoint the Churn Decision Window: Did the customer mentally churn during the onboarding/implementation phase (Bet B), or did they leave after reaching steady-state due to technological ceilings (Bet A)?
  2. Assess Real Robotics Demand vs. Narrative: Are customers actively deploying autonomous mobile robots (AMRs) and pick-assist hardware, or is "missing robotics" a convenient, forward-looking justification for leaving an underperforming platform?
  3. Quantify the Cost of Onboarding Drag: Does exceeding the 45-day SLA directly burn operational credibility and destroy ROI, or is delay merely a symptom of customer-side disorganization?
  4. Identify False Dichotomies (Bet Neither): Determine whether churn is driven by factors neither bet solves—such as baseline software unreliability, missing core 3PL billing/EDI features, or macro 3PL volume contraction.

---

3. Interviewer Guidance & Rules of Engagement

  • Recording Consent (Mandatory): State verbatim: "Before we begin, do you mind if I record this session purely for internal note-taking? None of this will be shared externally." If declined, proceed with manual notes.
  • The "No Roadmap, No Pricing" Wall: Customers will attempt to trade feedback for commitments. If asked about features or discounting, respond: "I'm on the product research side and have no visibility into commercials or delivery timelines. My sole focus today is understanding how your operations actually run."
  • Past Behavior Over Speculation: Never ask: "Would you use a robotics integration?" (Answer is always yes). Always ask: "What automation equipment do you have physically deployed on the floor today, and how does your team interact with it?"
  • Root-Cause Probing (The "Five Whys"): When a customer says "missing integrations," do not accept the label. Ask: "What specific warehouse task were you trying to execute that stalled? What was the manual workaround?"

---

4. Interview Scripts & Timings (45 Minutes Each)

Track 1: VP of Operations / COO (Economic Buyer)

#### Part 1: Context & Operational Profile (7 mins) * How has your facility footprint and throughput profile shifted over the last 18 months? * When you initially signed with Crate, what core business metric were you held accountable for improving?

#### Part 2: Implementation & Time-to-Value (12 mins) * Walk me back to your onboarding. What was the internal sentiment between signing the contract and processing your first live pallet? * Our contract targets a 45-day go-live; our median across customers is closer to 90. Where did the process stall, and what internal operational cost did that delay create? * At what point did your leadership team feel Crate was fully operational? Did you ever reach that state?

#### Part 3: Strategic Priorities, Automation, & Feature Gaps (16 mins) * Over the past year, what capital investments have you made on your warehouse floor (e.g., conveyor belts, automated guided vehicles, pick-assist carts, manual racking)? * If automation is present: Who manufactures it, what software controls it today, and what specific data must pass between it and your WMS? * If automation is absent: What is on your signed capital expenditure plan for the next 12 months? * When you evaluated the market (or competitors), what specific capability made it clear Crate was no longer the right long-term partner? * Think back to the moment you realized this contract wouldn't renew. When did that conversation happen, and who initiated it?

#### Part 4: Decision Trace & Wrap-up (10 mins) * If Crate had delivered on every promise made during the sales cycle within 45 days, would you still be with us today? Why or why not? * What is the single biggest operational bottleneck your business faces this quarter?

---

Track 2: Warehouse Operations Manager (Daily User)

#### Part 1: Floor Reality & Daily Routine (8 mins) * Walk me through the first two hours of your shift on the floor. Where do supervisors and pickers spend the majority of their time inside Crate? * What daily tasks force your team to leave Crate and use spreadsheets, paper, or secondary systems?

#### Part 2: Go-Live Experience & Transition (12 mins) * How disruptive was the transition to Crate for your floor staff? How long did it take for pick/pack rates to recover to baseline? * During the first three months, how often did you have to escalate systemic issues to leadership? What were those issues? * Did you feel the system was configured correctly for your physical warehouse layout, or did you have to mold your workflows to fit the software?

#### Part 3: Floor Automation & Daily Bottlenecks (15 mins) * How do orders physically move through your facility today? * If pickers are assisted by any mechanical or automated equipment, where does software failure slow them down? * Have you piloted or evaluated autonomous carts or robotics on your floor? * If yes: What was the hardest operational part of that pilot? * If no: What prevents you from automating (e.g., client SKU variability, facility constraints, budget)? * If you could fix one workflow in Crate that causes your floor workers the most frustration, what would it be?

#### Part 4: The Tipping Point (10 mins) * When did your team start discussing alternative systems or workarounds? Was there a specific breaking point or system failure? * If you had a magic wand to improve your team’s pick/pack speed by 20%, what would you change about your physical setup or your software?

---

5. Decision Rubric: What We Need to Hear

At the end of Week 3, interview findings will be mapped against this falsification framework:

OutcomeEvidence Required from CallsMetric Signal from Portfolio
:---:---:---
Back Bet A<br>(Robotics Integrations)• $\ge 4$ of the 6 accounts have active capital expenditures or signed contracts with AMR/robotics vendors (e.g., Locus, 6 River).<br>• Buyers explicitly state: "We loved Crate's daily usability and onboarding was fine, but our inability to ingest AMR pick-batches forced us to migrate."<br>• Daily users identify manual cart-pushing and picker transit time as their primary throughput ceiling.CRM shows churn clustered heavily in high-volume, automated accounts; robotics surfaced in competitive losses without prompting.
Back Bet B<br>(Implementation Rebuild)• Churn decisions were made inside the first 6 months, directly traceable to operational paralysis during go-live.<br>• Buyers report: "We burned through our operational budget running parallel systems for 100+ days; our team lost trust before we ever saw value."<br>• Robotics is dismissed as an aspirational 3-year vision, not a current operational requirement.First-year logo churn represents the majority of our 18% churn rate; churn correlates directly with onboarding days exceeding 45.
Back Neither<br>(Pivot to Bet C)• Automation is negligible among these mid-sized 3PLs; go-live delays are considered standard industry friction.<br>• Churn is driven by baseline platform stability (e.g., database timeouts, inventory sync errors), missing table-stakes 3PL multi-client billing, or standard EDI connectors to Shopify/Amazon.<br>• Customers cite price-to-value disconnects: "We left because Crate charges premium rates while basic core features remain broken."Exit survey "Missing features" was a proxy for basic warehouse workflows, not robotics; churn spread evenly across all customer tenure bands.

---

6. Next Steps

  1. Day 1-2: Approve adjusted account roster (3 non-renewing at-risk, 3 churned). CS to send intro emails.
  2. Day 3-17: Conduct 12 recorded interviews. Maintain a shared executive synthesis doc updated within 24 hours of each call.
  3. Day 21: Final Decision Meeting. Present thematic evidence and allocation recommendation to Ana and Marcus.

Claims the judge couldn’t find in the brief

  • Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year.
  • CRM shows churn clustered heavily in high-volume, automated accounts.
  • Robotics surfaced in competitive losses without prompting.
  • First-year logo churn represents the majority of the 18% churn rate.
  • Churn correlates directly with onboarding days exceeding 45.
  • Exit survey 'Missing features' was a proxy for basic warehouse workflows, not robotics.
  • Churn is spread evenly across all customer tenure bands.

Grades and run details

Decision model 77 · LLM judge 7 of 14 checks
Decision model checks
  • partialUses the supplied evidence correctly10%
  • passAddresses the actual decision91%
  • failRespects explicit constraints34%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims32%
  • passProduces the required deliverable48%
  • passTests both theories fairly82%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts99%
  • passDesigned to falsify77%
  • passAsks what people did, not what they think44%
  • partialDoesn't lead the witness38%
  • partialMarks what to cut if the call runs over50%
Run
Run
#1
API response time
37 s
Submitted
1 Oct 2026

Usable with a quick edit12 of 13 checks passed

The guide is thorough and evidence-based; a PM could add a note on which questions to prioritize if time is tight and use it immediately.

Output

Crate churn research: executive pre-read and interview guide

Decision: Allocate two squads next half to robotics integrations (A), implementation rebuild (B), or neither if the evidence does not justify either.

Research stance: Test both explanations, not confirm either executive’s account. These interviews can reveal mechanisms and decision timelines; they cannot estimate their prevalence across 340 customers or prove causality.

1. Learning goals

  1. Reconstruct the departure process: When did concerns emerge, when did leaving become likely, and when was the decision effectively made?
  2. Identify consequential problems: Separate reasons volunteered after leaving from problems that materially changed warehouse outcomes and the buying decision.
  3. Test Bet A: Was a specific robotics incompatibility a binding obstacle? Would addressing it have plausibly retained the account?
  4. Test Bet B: Did implementation delay prevent value, damage trust, or trigger departure? What caused the delay, and could Crate control it?
  5. Find alternatives and interactions: Price, service, reliability, business changes, or other missing capabilities may dominate. Robotics work may itself complicate implementation.
  6. Connect learning to an investable scope: What could two squads deliver next half that would address the demonstrated retention mechanism?

What the existing evidence does—and does not—say

  • The exit survey’s broad integration category, single-choice format, first-position placement, and sparse free text do not establish robotics as the cause.
  • The 2.4× churn association makes implementation worth investigating, but complexity, customer readiness, or other factors could drive both delay and churn.
  • Three competitor moves and nine lost deals indicate a robotics opportunity; acquisition losses are not evidence of a retention mechanism.
  • Unknown decision dates make chronology essential.

2. Changes to the call plan

Resolve the arithmetic: Twelve separate interviews cannot cover twelve accounts with two roles each. Use six accounts × two separate role interviews = twelve 45-minute calls.

Recommended mix: Four churned accounts and two at-risk accounts. This reduces reliance on renewal negotiations while retaining a prospective view. Findings will be depth-oriented, not representative.

Select accounts from the full CRM population, not executive contacts or survey answers alone:

  • Across churned accounts, include slow and faster implementations, first-year and later departures, and robotics-exposed and non-exposed customers.
  • Include at least one confirmed move to a robotics-enabled competitor and one slow implementation without an apparent robotics need.
  • For at-risk accounts, prefer one with active automation plans and one with implementation/value-realization problems. Include the non-renewal account if suitable.
  • Seek contrasting retained accounts through existing operational data, even though the call budget does not cover them.

Recruit both the daily warehouse operations manager and the contract-signing VP Operations/COO. If the original signer has left, recruit someone directly involved in the departure or renewal decision and document that substitution.

Ana should not lead calls with her personal contacts. Use a neutral interviewer without renewal responsibility. Ana can help recruit; observing requires explicit customer consent. Her existing conversations are useful leads, not independently verified findings.

CS should introduce research as separate from renewal discussions, then step out. Participation must not affect commercial treatment.

3. Interview guides: 45 minutes each

Use the relevant role column. In every section, distinguish direct experience from hearsay. Ask broad questions before naming either bet.

TimeDaily warehouse operations managerVP Operations / COO
0–3 min: permission and framingFirst ask: “May we record this conversation for internal research?” If declined, continue with notes. Then explain purpose and confidentiality limits.Same opening.
3–7 min: context and expected value“What did your warehouse handle, and what was your role with Crate? What was supposed to improve?”“What led you to buy Crate? What outcomes and deadlines mattered? How would you judge success?”
7–17 min: chronological reconstruction“Walk me from signing through onboarding, first operational use, and the point when problems became serious.” “Describe a particular shift or incident.” Establish dates, workarounds, operational impact, and who knew.“Walk me from purchase to the first concern, evaluating alternatives, and the decision.” “When did staying stop being the default?” Establish dates, decision-makers, triggers, and alternatives.
17–27 min: implementation and value“What had to happen before you could use Crate successfully? Where did work stall? Who owned each step?” “When did you first achieve useful results?” Probe migration, configuration, training, integrations, staffing, and rework after an open answer.“What go-live date did you expect, and what happened? What consequences did that have?” “What other factors accompanied the delay?” “If the same product had gone live in 45 days, what would still have put the relationship at risk?” Ask why.
27–36 min: workflow, automation, and unmet needs“Which workflows did Crate support poorly? Show or describe a recent example.” Then: “Were robots in use or planned? Which systems, what workflow, and what connection was required?” “What happened without it?”“What capabilities influenced staying or switching?” Then probe automation: vendor, deployment date, committed budget, required integration, and alternatives considered. “If that connection had existed, what else would have needed to change for you to stay?”
36–42 min: decision and competing explanations“Which problem mattered most in daily operations? What did you escalate, to whom, and when?” “What worked well?” “What have we missed?”“Which issues were necessary to the decision, and which were secondary?” “What would have had to be different for you to stay?” Ask about business changes, service, reliability, and other alternatives without forcing a category.
42–45 min: verify and closeSummarize the timeline and mechanism: “What have I misunderstood?” Request relevant artifacts and permission for a brief clarification follow-up.Same; verify whether operational problems actually influenced the commercial decision.

Status-specific wording

  • Churned: Ask what the replacement actually delivered, whether it went live, and whether the cited problem improved—not merely what the competitor promised.
  • At-risk: Ask what has already happened, current unresolved consequences, and what would cause escalation or departure. Do not imply a departure decision exists. Distinguish funded plans from aspirations.

4. Guidance for interviewers

  • Begin with recording permission; record only after consent. Explain that this is research, not a support, roadmap, or renewal discussion.
  • Do not discuss prices, discounts, or roadmap commitments. If raised, acknowledge without probing commercial terms: “I can’t discuss that here; your account team handles commercial conversations.” Record volunteered price concerns as an alternative explanation.
  • Avoid “Did missing robotics make you leave?” and “Would faster onboarding have saved you?” Start with events; use counterfactuals only afterward.
  • Ask for concrete examples, dates, artifacts, and consequences. “It was painful” is not yet a mechanism.
  • Do not defend Crate, disclose executive hypotheses, or offer fixes.
  • Interview roles separately before reconciling accounts. Treat disagreement as evidence, not an error to erase.
  • Use a note-taker and common template: timeline, problem, consequence, decision link, alternative explanation, evidence source, and confidence. Count accounts, not two interviewees as independent votes.

5. Evidence required for the decision

Back A: robotics integrations

Look for a repeated chain across independent accounts:

Committed robotics deployment → specific missing connection → material operational/business constraint → documented role in departure or credible near-term risk.

Stronger evidence includes decision-maker confirmation, pre-decision correspondence, and a replacement that solved the problem. Match the implicated vendors/workflows to current customer exposure. Before funding, establish that a feasible integration scope could address meaningful retention risk next half.

Mentions of “automation,” hypothetical plans, or competitor marketing are insufficient.

Back B: implementation rebuild

Look for:

Crate-controllable implementation bottleneck → delayed usable value or broken operational commitments → loss of confidence/avoidable cost → departure decision before value realization.

Seek repeated bottlenecks, role-level corroboration, and contrasts with faster implementations. Validate the mechanism against the full cohort, accounting where possible for complexity and customer readiness. Identify changes two squads could deliver; do not assume a wholesale rebuild is necessary.

Back neither—or declare the result inconclusive

Neither is justified if dominant causes lie elsewhere, robotics demand is narrow or speculative, delays are largely outside Crate’s control, or either intervention arrives too late to affect retention. If interviews remain contradictory or unsupported, report insufficient evidence, not “neither causes churn.”

Compare both bets on reachable ARR at risk, evidence strength, controllability, delivery feasibility, and time to impact—not mention counts.

6. Three-week execution

  • Week 1: Audit cohorts and timelines, recruit, pilot one role pair.
  • Week 2: Complete remaining interviews; synthesize after each pair without changing standards to favor an emerging winner.
  • Week 3: Verify discrepancies, triangulate with records, assess engineering scope, and prepare the decision memo with supporting evidence, counterexamples, uncertainties, and a recommendation.

Do not commit two squads merely because one story sounds more persuasive.

Grades and run details

Decision model 92 · LLM judge 13 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly54%
  • passAddresses the actual decision90%
  • partialRespects explicit constraints21%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims96%
  • passProduces the required deliverable89%
  • passTests both theories fairly100%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts95%
  • passDesigned to falsify63%
  • passAsks what people did, not what they think67%
  • passDoesn't lead the witness53%
  • partialMarks what to cut if the call runs over58%
Run
Run
#1
API response time
80 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 6

Uses the supplied evidence correctlyMixedWrongRight
Sonnet 5.5 · API

All factual claims about the current situation are directly supported by the brief or derived correctly from its numbers.

Gemini 3.8 Flash · API

The guide relies on several unsupported current-situation or portfolio claims, including CRM clustering, robotics surfacing without prompting, first-year churn majority, >45-day correlation, and exit-survey meaning.

GPT-6.1 Sol · API

All factual statements about the current situation are taken directly from the supplied context without invention.

Addresses the actual decisionRightMixedRight
Sonnet 5.5 · API

The output provides clear, actionable criteria for what would back each bet or neither after the calls, exactly as the brief requests.

Gemini 3.8 Flash · API

It does not commit to Bet A, Bet B, or neither; it defers the decision to the research, even though the supplied evidence is not framed as insufficient to make a call.

GPT-6.1 Sol · API

The output recommends more research before deciding, states the supplied evidence cannot yet support a call, and specifies exactly what evidence would back each bet or neither.

Respects explicit constraintsRightWrongRight
Sonnet 5.5 · API

The guide respects all constraints: it stays under 1,500 words, includes all requested sections, and provides concrete scripts and plans to enforce the off-limits topics and consent.

Gemini 3.8 Flash · API

It exceeds the requested 1,500-word limit and its proposed 6-account roster would not cover the 12 accounts specified in the supplied call plan.

GPT-6.1 Sol · API

The guide is under 1,500 words, includes all requested sections (learning goals, timed questions for two roles, guidance, plan changes, decision signals), and is addressed to the exec team.

Avoids unsupported claimsRightWrongRight
Sonnet 5.5 · API

Hypotheses and interpretations are clearly labelled as such, and no cause or forecast is presented as established fact.

Gemini 3.8 Flash · API

It presents several interpretations and portfolio signals as established facts or evidence, such as exit-survey meaning, churn distribution, and robotics mentions without prompting.

GPT-6.1 Sol · API

Interpretations are clearly labelled as such (e.g., the exit survey does not establish robotics as the cause), and no cause is presented as established fact.

Produces the required deliverableRightMixedRight
Sonnet 5.5 · API

The guide is complete with learning goals, timed questions for two roles, interviewer guidance, plan changes, and decision signals, all under 1,500 words and ready for the exec team.

Gemini 3.8 Flash · API

Although it is a usable interview guide, it is too long and changes the account scope in a way that does not satisfy the supplied 12-call plan.

GPT-6.1 Sol · API

The output is a complete guide with learning goals, timed questions for both roles, interviewer guidance, plan changes, and decision signals; an exec team could act on it.

Marks what to cut if the call runs overRightWrongWrong
Sonnet 5.5 · API

Both 45-minute guides have timed sections that sum correctly, a realistic number of questions, and instructions to skip covered topics if time is short.

Gemini 3.8 Flash · API

The timings add up to 45 minutes, but the number of questions is too dense for a realistic 45-minute call and there are no must-ask markers or cut-down instructions.

GPT-6.1 Sol · API

The guide includes timed sections that add up to 45 minutes but does not say what to cut if time runs short, nor does it mark must-ask questions.

All got right 7

Identifies material uncertaintyRightRightRight
Sonnet 5.5 · API

The output explicitly names the key unknowns (weak survey, correlation vs causation, unknown decision timing) and says how the calls and a parallel data pull would resolve them.

Gemini 3.8 Flash · API

It names material unknowns such as timing of churn decisions, real robotics deployment, onboarding drag versus customer-side disorganization, and third causes, and proposes calls to resolve them.

GPT-6.1 Sol · API

It names key unknowns (decision dates, whether robotics was a binding obstacle, whether delays were Crate-controllable) and says how the interview evidence would resolve them.

Tests both theories fairlyRightRightRight
Sonnet 5.5 · API

Both theories get questions that could disprove them (e.g., no action on robotics, happiness after slow go-live), and open-ended timeline questions leave room for a third cause.

Gemini 3.8 Flash · API

It gives both robotics and implementation disconfirming questions and explicitly leaves room for neither, including stability, billing/EDI, and macro causes.

GPT-6.1 Sol · API

Questions could disprove robotics (e.g., 'Were robots in use? What happened without it?') and implementation (e.g., 'Where did work stall? If go-live had been 45 days, what would still have put the relationship at risk?'), and open questions leave room for a third cause.

Different questions for user and signerRightRightRight
Sonnet 5.5 · API

The guide has distinct tracks: daily user questions focus on day-to-day experience and workarounds, signer questions on the decision process and alternatives.

Gemini 3.8 Flash · API

It provides separate tracks for the daily warehouse operations manager and the VP/COO signer, with role-appropriate questions.

GPT-6.1 Sol · API

The table provides distinct tracks: daily user gets questions about specific shifts, workarounds, and escalations; the signer gets questions about the purchase decision, alternatives, and go-live expectations.

Protects the calls and the accountsRightRightRight
Sonnet 5.5 · API

It includes a clear off-limits rule with scripts for pricing/roadmap, starts with consent, and addresses the CEO's bias by having her not lead calls and limiting her contacts.

Gemini 3.8 Flash · API

It includes recording consent, a no-pricing/no-roadmap rule with a script, and a concrete protocol for Ana's calls.

GPT-6.1 Sol · API

It forbids discussing pricing/discounts/roadmap, provides a script to deflect, starts with recording consent, and has a concrete plan for Ana (she does not lead, observes with consent, her prior conversations are treated as leads).

Designed to falsifyRightRightRight
Sonnet 5.5 · API

Every learning goal has a disconfirming question, such as asking what they did about robots or whether they ever reached value after a slow go-live.

Gemini 3.8 Flash · API

Each learning goal has questions that could falsify the relevant hypothesis, such as actual automation deployment, go-live timeline, and alternative causes.

GPT-6.1 Sol · API

Each learning goal has disconfirming questions (e.g., 'What happened without the robot connection?' for Bet A, 'If the same product had gone live in 45 days, what would still have put the relationship at risk?' for Bet B).

Asks what people did, not what they thinkRightRightRight
Sonnet 5.5 · API

Core questions ask for specific recent actions and events ('What happened next?', 'Who did you call?'), not opinions or predictions.

Gemini 3.8 Flash · API

The core questions ask for recent concrete behavior, including deployed equipment, go-live experience, decision timing, and workarounds.

GPT-6.1 Sol · API

Core questions ask for a chronological walk-through, specific incidents, dates, and what they did, not opinions or predictions.

Doesn't lead the witnessRightRightRight
Sonnet 5.5 · API

Questions are neutral and open-ended; the prompted topics are introduced without pitching or hinting at the desired answer.

Gemini 3.8 Flash · API

The questions are mostly neutral and avoid asking customers to choose between the bets or pitching the ideas.

GPT-6.1 Sol · API

Questions are open and neutral; robotics and implementation are not named until after broad problem questions, and no answer is hinted at.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review