Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Crate. In four weeks, the exec team decides which of two bets gets our two squads next half, and both bets rest on a theory of why customers are leaving. We have three weeks for 12 customer calls of 45 minutes. Write the guide for them. It should include: the learning goals; the questions, with timings, for the two kinds of people we'll talk to; guidance for whoever runs the calls; any changes you'd make to the plan below; and what we'd need to hear to back each bet, or neither. Keep it under 1,500 words. The exec team will read it before the calls start.

What the model was given8 items: About Crate, The decision, CEO, Ana Ferreira, CPO, Marcus Hill, Exit survey (41 churned customers, last 12 months), Implementation and CRM data, The call plan, Notes from Customer Success and Legal
About CrateWarehouse management software for mid-size third-party logistics companies (3PLs). 340 customers, $22M ARR. Gross logo churn rose from 11% to 18% over the last year.
The decisionBet A: integrations with warehouse robotics (autonomous carts, pick-assist robots). Bet B: rebuild implementation so customers go live faster. Two squads for the next half go to one of them.
CEO, Ana Ferreira“We're losing warehouses because they're automating and we don't talk to their robots. Every lost customer I've spoken to mentions it.”
CPO, Marcus Hill“We lose them in the first year because implementation takes forever. They're gone before they ever see the value.”
Exit survey (41 churned customers, last 12 months)A single-choice question, 'Why are you leaving?'. Missing features or integrations: 46%. Too hard to implement: 22%. Price: 17%. Other: 15% (free-text box, left blank by all but two). 'Missing features or integrations' was the first option in the list.
Implementation and CRM dataMedian time from signing to go-live: 94 days; the contract promises 45. Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year. 3 churned customers moved to a competitor with robotics integrations; robotics was cited in 9 lost new-logo deals. We don't know when in the year churned customers decided to leave.
The call plan12 calls: 6 with customers who churned in the last 9 months, 6 with at-risk customers (health score red). On each account we want to talk to the warehouse operations manager who used Crate every day and, separately, the VP of Operations or COO who signed the contract. Ana wants to run 4 of the calls with former customers she knows personally.
Notes from Customer Success and Legal5 of the 6 at-risk accounts are in renewal talks, and Customer Success is offering them discounts. Interviewers must not discuss pricing, discounts or what's on our roadmap. Every call must start by asking permission to record.
What a strong answer doesThe answer key the graders mark against

A guide that tests both theories fairly and leaves room for a third: the goals ask what actually drove each decision to leave or doubt, when it happened, and who made it, with questions that could disprove robotics (did they automate, did robots come up before they decided, what did they switch to and why) and implementation (how long go-live took, what happened in those months, whether value arrived) alike. It treats the exit survey as weak evidence (single choice, integrations listed first, 'Other' left blank) and the 2.4× figure as a correlation to probe, not a cause. The core is a timeline of the account: the story from signing to the moment they decided to leave or started doubting, with sequential probes. It has separate tracks for the daily user (what went wrong day to day) and the signer (how the decision was made, what alternatives they weighed), fits 45 minutes with timings, and opens with consent to record. It guards the calls: no pitching, pricing, discounts or roadmap promises, with a script for when an at-risk customer asks, and it addresses Ana's bias: she has a public view and personal relationships, so she should observe or take notes rather than lead, or at minimum use the same script and have her calls read separately. It ends with decision signals: what pattern across the 12 would back Bet A, what would back Bet B, and what would back neither (price, a business change, a missing capability that isn't robotics).

Critical failures (cap the score at 40)
  • Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks
  • Asks customers to choose between robotics and implementation as the reason they left
Case

Staff-level · v1.1 · synthetic · B2B SaaS, logistics, churn, exec bias, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit7 of 13 checks passed

A PM would need to shorten it, reconcile the account plan with the 12-call constraint, and remove or relabel unsupported portfolio claims before using it.

Output

Research Guide: Root Cause Analysis for H2 Squad Allocation

Prepared by: Staff Product Manager Target Audience: Executive Team (CEO, CPO, VP CS, Head of Product) Timeline: 3 weeks of interviews; executive decision in Week 4 Scope: 12 calls (45 minutes each)

---

1. Executive Summary & Recommended Changes to the Call Plan

Before executing this sprint, we must adjust three structural flaws in the proposed plan to protect data integrity:

  1. Resolve the Account-to-Call Ratio: 12 calls cannot cover two separate stakeholders across 12 accounts. We will target 6 accounts total (3 churned, 3 at-risk) and conduct 2 distinct calls per account: one with the Warehouse Operations Manager (day-to-day user) and one with the VP of Operations/COO (economic buyer). Evaluating both perspectives within the same operational context is essential to determine whether operational failure or strategic misalignment triggered the exit.
  2. Quarantine Active Renewals: 5 of the proposed 6 at-risk accounts are actively negotiating renewals and receiving discounts. Interviewing these buyers creates an immediate incentive for them to exaggerate product flaws to gain commercial leverage. We will swap 3 of these with at-risk accounts that are not currently negotiating pricing.
  3. CEO Interview Protocol: Executive presence introduces severe confirmation and deference bias—especially with personal contacts who already know Ana's passion for robotics. Recommendation: Ana should join her 4 target calls as an executive sponsor for the first 3 minutes, then hand off lead facilitation to a Product Manager, remaining on mute as an observer. If Ana leads, she must strictly follow the non-leading script provided.

---

2. Core Learning Goals

  1. Pinpoint the Churn Decision Window: Did the customer mentally churn during the onboarding/implementation phase (Bet B), or did they leave after reaching steady-state due to technological ceilings (Bet A)?
  2. Assess Real Robotics Demand vs. Narrative: Are customers actively deploying autonomous mobile robots (AMRs) and pick-assist hardware, or is "missing robotics" a convenient, forward-looking justification for leaving an underperforming platform?
  3. Quantify the Cost of Onboarding Drag: Does exceeding the 45-day SLA directly burn operational credibility and destroy ROI, or is delay merely a symptom of customer-side disorganization?
  4. Identify False Dichotomies (Bet Neither): Determine whether churn is driven by factors neither bet solves—such as baseline software unreliability, missing core 3PL billing/EDI features, or macro 3PL volume contraction.

---

3. Interviewer Guidance & Rules of Engagement

  • Recording Consent (Mandatory): State verbatim: "Before we begin, do you mind if I record this session purely for internal note-taking? None of this will be shared externally." If declined, proceed with manual notes.
  • The "No Roadmap, No Pricing" Wall: Customers will attempt to trade feedback for commitments. If asked about features or discounting, respond: "I'm on the product research side and have no visibility into commercials or delivery timelines. My sole focus today is understanding how your operations actually run."
  • Past Behavior Over Speculation: Never ask: "Would you use a robotics integration?" (Answer is always yes). Always ask: "What automation equipment do you have physically deployed on the floor today, and how does your team interact with it?"
  • Root-Cause Probing (The "Five Whys"): When a customer says "missing integrations," do not accept the label. Ask: "What specific warehouse task were you trying to execute that stalled? What was the manual workaround?"

---

4. Interview Scripts & Timings (45 Minutes Each)

Track 1: VP of Operations / COO (Economic Buyer)

#### Part 1: Context & Operational Profile (7 mins) * How has your facility footprint and throughput profile shifted over the last 18 months? * When you initially signed with Crate, what core business metric were you held accountable for improving?

#### Part 2: Implementation & Time-to-Value (12 mins) * Walk me back to your onboarding. What was the internal sentiment between signing the contract and processing your first live pallet? * Our contract targets a 45-day go-live; our median across customers is closer to 90. Where did the process stall, and what internal operational cost did that delay create? * At what point did your leadership team feel Crate was fully operational? Did you ever reach that state?

#### Part 3: Strategic Priorities, Automation, & Feature Gaps (16 mins) * Over the past year, what capital investments have you made on your warehouse floor (e.g., conveyor belts, automated guided vehicles, pick-assist carts, manual racking)? * If automation is present: Who manufactures it, what software controls it today, and what specific data must pass between it and your WMS? * If automation is absent: What is on your signed capital expenditure plan for the next 12 months? * When you evaluated the market (or competitors), what specific capability made it clear Crate was no longer the right long-term partner? * Think back to the moment you realized this contract wouldn't renew. When did that conversation happen, and who initiated it?

#### Part 4: Decision Trace & Wrap-up (10 mins) * If Crate had delivered on every promise made during the sales cycle within 45 days, would you still be with us today? Why or why not? * What is the single biggest operational bottleneck your business faces this quarter?

---

Track 2: Warehouse Operations Manager (Daily User)

#### Part 1: Floor Reality & Daily Routine (8 mins) * Walk me through the first two hours of your shift on the floor. Where do supervisors and pickers spend the majority of their time inside Crate? * What daily tasks force your team to leave Crate and use spreadsheets, paper, or secondary systems?

#### Part 2: Go-Live Experience & Transition (12 mins) * How disruptive was the transition to Crate for your floor staff? How long did it take for pick/pack rates to recover to baseline? * During the first three months, how often did you have to escalate systemic issues to leadership? What were those issues? * Did you feel the system was configured correctly for your physical warehouse layout, or did you have to mold your workflows to fit the software?

#### Part 3: Floor Automation & Daily Bottlenecks (15 mins) * How do orders physically move through your facility today? * If pickers are assisted by any mechanical or automated equipment, where does software failure slow them down? * Have you piloted or evaluated autonomous carts or robotics on your floor? * If yes: What was the hardest operational part of that pilot? * If no: What prevents you from automating (e.g., client SKU variability, facility constraints, budget)? * If you could fix one workflow in Crate that causes your floor workers the most frustration, what would it be?

#### Part 4: The Tipping Point (10 mins) * When did your team start discussing alternative systems or workarounds? Was there a specific breaking point or system failure? * If you had a magic wand to improve your team’s pick/pack speed by 20%, what would you change about your physical setup or your software?

---

5. Decision Rubric: What We Need to Hear

At the end of Week 3, interview findings will be mapped against this falsification framework:

OutcomeEvidence Required from CallsMetric Signal from Portfolio
:---:---:---
Back Bet A<br>(Robotics Integrations)• $\ge 4$ of the 6 accounts have active capital expenditures or signed contracts with AMR/robotics vendors (e.g., Locus, 6 River).<br>• Buyers explicitly state: "We loved Crate's daily usability and onboarding was fine, but our inability to ingest AMR pick-batches forced us to migrate."<br>• Daily users identify manual cart-pushing and picker transit time as their primary throughput ceiling.CRM shows churn clustered heavily in high-volume, automated accounts; robotics surfaced in competitive losses without prompting.
Back Bet B<br>(Implementation Rebuild)• Churn decisions were made inside the first 6 months, directly traceable to operational paralysis during go-live.<br>• Buyers report: "We burned through our operational budget running parallel systems for 100+ days; our team lost trust before we ever saw value."<br>• Robotics is dismissed as an aspirational 3-year vision, not a current operational requirement.First-year logo churn represents the majority of our 18% churn rate; churn correlates directly with onboarding days exceeding 45.
Back Neither<br>(Pivot to Bet C)• Automation is negligible among these mid-sized 3PLs; go-live delays are considered standard industry friction.<br>• Churn is driven by baseline platform stability (e.g., database timeouts, inventory sync errors), missing table-stakes 3PL multi-client billing, or standard EDI connectors to Shopify/Amazon.<br>• Customers cite price-to-value disconnects: "We left because Crate charges premium rates while basic core features remain broken."Exit survey "Missing features" was a proxy for basic warehouse workflows, not robotics; churn spread evenly across all customer tenure bands.

---

6. Next Steps

  1. Day 1-2: Approve adjusted account roster (3 non-renewing at-risk, 3 churned). CS to send intro emails.
  2. Day 3-17: Conduct 12 recorded interviews. Maintain a shared executive synthesis doc updated within 24 hours of each call.
  3. Day 21: Final Decision Meeting. Present thematic evidence and allocation recommendation to Ana and Marcus.

Claims the judge couldn’t find in the brief

  • Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year.
  • CRM shows churn clustered heavily in high-volume, automated accounts.
  • Robotics surfaced in competitive losses without prompting.
  • First-year logo churn represents the majority of the 18% churn rate.
  • Churn correlates directly with onboarding days exceeding 45.
  • Exit survey 'Missing features' was a proxy for basic warehouse workflows, not robotics.
  • Churn is spread evenly across all customer tenure bands.

Grades and run details

Decision model 77 · LLM judge 7 of 14 checks
Decision model checks
  • partialUses the supplied evidence correctly10%
  • passAddresses the actual decision91%
  • failRespects explicit constraints34%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims32%
  • passProduces the required deliverable48%
  • passTests both theories fairly82%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts99%
  • passDesigned to falsify77%
  • passAsks what people did, not what they think44%
  • partialDoesn't lead the witness38%
  • partialMarks what to cut if the call runs over50%
Run
Run
#1
API response time
37 s
Submitted
1 Oct 2026
Critical failure: Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks. Score capped at 40.

Needs more than a quick edit3 of 13 checks passed

A PM would need to fix the prohibited either/or question, add a pricing/roadmap handling script, remove unsupported claims, and clarify the revised call coverage before using it. The reasoning and evidence also need rechecking.

Output

Customer Research Guide: H2 Strategic Bets (Robotics vs. Implementation)

To: Executive Team From: Staff Product Manager, Crate Date: Today Subject: 3-Week Customer Research Plan for H2 Resource Allocation (12 Calls)

---

Executive Summary & Plan Adjustments

In four weeks, the executive team will decide where to allocate both engineering squads for the next half. Ana believes we are losing customers because we lack warehouse robotics integrations (Bet A). Marcus believes we are losing them because our implementation cycle is broken and slow (Bet B).

The exit survey heavily favors Ana’s view (46% cite missing features/integrations), but UX research 101 teaches us that exit surveys are structurally flawed. Listing "missing features" first creates primacy bias, and the survey failed to capture temporal dynamics—specifically, when the decision to leave was made. Meanwhile, our internal data tells a damning story: our median time to go-live is 94 days (vs. 45 promised), and customers crossing 90 days churn at $2.4\times$ the rate.

We have 3 weeks to conduct 12 interviews (45 minutes each) to uncover the truth.

Critical Plan Adjustments (Read Before We Start)

  1. Split the Interviewees by Persona, Not Account: The original plan called for interviewing both the Ops Manager and the VP/COO on the same accounts. In 45 minutes, splitting time between a daily user and a high-level signer dilutes the insight. Instead, we will interview Warehouse Ops Managers to understand day-to-day friction (Implementation, UI, operational bottlenecks) and VPs of Ops / COOs to understand strategic drivers (ROI, contract expectations, robotics, and when the churn decision was actually made).
  2. Revise Ana’s Participation: Ana wants to run 4 calls with former customers she knows personally. We must advise against this. Founder-led interviews with churned customers carry massive politeness bias; customers will pull punches about implementation failures or invent reasons to let Ana down gently. Ana should observe these calls anonymously or listen to the recordings, but an unbiased PM or User Researcher must conduct them.
  3. Guardrails for At-Risk Accounts: Five of our six at-risk accounts are in active renewal talks with CS discounts on the table. Interviewers are strictly prohibited from discussing pricing, commercial concessions, or our product roadmap. We are strictly here to diagnose reality.

---

Learning Goals

  1. The Churn Timeline: Determine the exact month the customer mentally decided to leave. Did they decide during a painful 90+ day onboarding (supporting Bet B), or later when evaluating modernizing warehouse automation (supporting Bet A)?
  2. The "Missing Features" Audit: Deconstruct the 46% exit survey metric. When respondents checked "missing features/integrations," were they missing critical WMS functionality (e.g., wave planning, inventory accuracy) or specifically robotics middleware?
  3. Implementation Reality Check: Measure the operational impact of our 94-day median go-live time. Did the delay create cash-flow/trust erosion that doomed the account before day one, regardless of subsequent features?
  4. Robotics Urgency: For customers who mention robotics, are they actively deploying autonomous carts today, or is it a 3-year roadmap item they use as a stick to beat us with during negotiations?

---

Interview Guide: Two Personas

Persona 1: The Daily User (Warehouse Operations Manager)

Focus: Implementation friction, daily utility, operational roadblocks.

  • 00:00–03:00 | Intro & Permission:
  • "Thanks for taking the time. I’m [Name] from Crate. We’re doing research to figure out where we failed our customers and how we can improve. I am strictly here to learn about your experience—we won't be talking about pricing, renewals, or selling you anything today. Do I have your permission to record this call for internal notes?"
  • 03:00–12:00 | Day-to-Day Workflow & Reality:
  • "Take me back to your first 30 days using Crate. What did the day-to-day onboarding look like?"
  • "How long did it actually take for your floor team to feel comfortable using Crate without holding their breath?"
  • "What was the most painful workaround your team had to invent during your first three months?"
  • 12:00–25:00 | The Implementation Experience (Bet B Probe):
  • "Our data shows it takes about 90+ days for most warehouses to go live on Crate. What was your experience? How did those delays affect your team's stress levels and your relationship with our implementation team?"
  • "Looking back, did the slow start permanently damage your trust in the software, or did you bounce back once things were running?"
  • 25:00–35:00 | Equipment & Floor Integration (Bet A Probe):
  • "What automated hardware or mobile carts are running on your floor today? How do they talk to Crate, or do they?"
  • "When a pick-assist cart or AGV moves through your warehouse, where does the process break down between the robot and Crate?"
  • 35:00–42:00 | The Breaking Point:
  • "Think about the moment you realized you wanted to leave Crate (or started looking at alternatives). What happened on the warehouse floor that week?"
  • 42:00–45:00 | Wrap-Up:
  • "If you could wave a magic wand and change one thing about your onboarding or your day-one experience with Crate, what would it be?"

---

Persona 2: The Decision Maker (VP of Operations / COO)

Focus: Strategic expectations, ROI timelines, contract promises, and robotics strategy.

  • 00:00–03:00 | Intro & Permission:
  • "Thanks for joining. I’m [Name] from Crate. We’re reviewing why customers leave us to make radical changes to our roadmap and processes. I'm not here to pitch you or talk commercial terms—just to understand your true experience. May I record this call for note-taking purposes?"
  • 03:00–10:00 | The Purchase & Promise:
  • "When you signed with Crate, what was the primary business problem you needed solved in your first 6 months?"
  • "What timeline were you given for go-live during the sales cycle, and how did reality compare?"
  • 10:00–22:00 | Implementation & Time-to-Value (Bet B Probe):
  • "For a mid-size 3PL, cash flow and client onboarding speed are everything. How did Crate's onboarding timeline impact your ability to onboard your own end-clients?"
  • "At what point in your first year did you officially consider the implementation 'failed' or deeply troubled?"
  • 22:00–34:00 | The Robotics Mandate (Bet A Probe):
  • "Let’s talk about warehouse automation. How central are autonomous carts or robotics to your strategy today versus three years from now? Are you actively running them?"
  • "When you evaluated staying vs. leaving Crate, how heavy was the absence of robotics integrations in that final calculus? Was it the straw that broke the camel's back, or a nice-to-have?"
  • 34:00–42:00 | Uncovering the Timeline:
  • "Looking back at the lifecycle of your account, when did you mentally check out and decide you wouldn't renew? Was it during implementation, or months later?"
  • 42:00–45:00 | Wrap-Up:
  • "If Crate could fix only one thing—either our onboarding speed or our hardware/robotics connectivity—which one would have saved your account?"

---

Decision Criteria: What We Need to Hear

To make our final recommendation to Ana and Marcus, we will evaluate the 12 calls against the following thresholds:

Backing Bet A (Robotics Integrations)

We back Bet A only if: 1. VPs/COOs explicitly state that the lack of robotics was the primary, active catalyst for churning (not just a grievance listed in a post-hoc survey). 2. Churned customers are actively utilizing or piloting autonomous hardware today that could not function with Crate, resulting in operational deadlocks. 3. The exit survey's 46% "missing features" bucket is proven via interviews to be overwhelmingly about hardware/robotics connectivity rather than core WMS gaps (e.g., billing, inventory tracking).

Backing Bet B (Implementation Rebuild)

We back Bet B only if: 1. Interviews reveal that the 90+ day onboarding process created an insurmountable trust deficit; customers felt abandoned and never recovered faith in the product, rendering subsequent features irrelevant. 2. Daily users report that the complexity of setup and configuration was the primary driver of operational frustration, overshadowing missing integrations. 3. Churned customers admit they mentally checked out during or immediately after the botched implementation window, long before robotics ever entered the conversation.

Backing Neither (The Pivot Option)

We back Neither if: * The qualitative data shows a split narrative where both bets are symptoms of a deeper root cause—such as core WMS unreliability, pricing/value misalignment (17% in exit survey), or poor account management—meaning that building robots or speeding up a broken implementation will simply churn customers faster down a different path.

Claims the judge couldn’t find in the brief

  • For a mid-size 3PL, cash flow and client onboarding speed are everything.

Grades and run details

Decision model 58 · LLM judge 3 of 14 checks
Decision model checks
  • partialUses the supplied evidence correctly23%
  • passAddresses the actual decision93%
  • partialRespects explicit constraints21%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims46%
  • partialProduces the required deliverable26%
  • partialTests both theories fairly47%
  • passDifferent questions for user and signer98%
  • partialProtects the calls and the accounts84%
  • partialDesigned to falsify40%
  • partialAsks what people did, not what they think52%
  • failDoesn't lead the witness30%
  • partialMarks what to cut if the call runs over45%
Run
Run
#1
Time to output
11 s
Submitted
1 Oct 2026

Needs more than a quick edit11 of 13 checks passed

The guide lacks separate question tracks for the daily user and the decision-maker, and it doesn't prioritize questions for time management. A PM would need to add role-specific questions and mark must-asks, which is more than a quick edit.

Output

Crate customer research guide: why are warehouses leaving?

Purpose: In three weeks, learn whether churn is more plausibly driven by missing robotics integrations (Bet A), slow implementation (Bet B), or something else. These calls are for understanding customers’ experiences—not selling, negotiating renewals, or validating either executive’s theory.

Learning goals

For each account, establish:

  • What happened, and when? Separate implementation, go-live, first value, automation needs, and the point when leaving became likely or decided.
  • What mattered most? Distinguish a decisive problem from a feature request, contributing frustration, or reason offered after the fact.
  • What evidence supports the explanation? Look for actual workflow impact, workarounds, alternatives evaluated, and actions taken—not just opinions or hypotheticals.
  • What might change the outcome? Explore credible counterfactuals without promising anything.
  • What else explains churn? Surface causes beyond A and B.

Recommended plan changes

The proposed 12 calls cannot cover both the daily user and contract signer on all 12 accounts: that would require 24 interviews. Keep the 12-call cap, but interview two roles separately at each of six accounts: three churned accounts and three red-health accounts, with one warehouse operations manager and one VP of Operations/COO per account. This gives paired perspectives and preserves both customer situations, but is a small, directional sample—not a prevalence estimate. Select accounts for varied implementation times and automation situations where possible; don’t choose only accounts with known robotics issues.

Do not have Ana conduct four calls with former customers she knows personally. Her involvement risks courtesy bias and leading the conversation. Use an independent interviewer; ideally don’t include personal contacts in this small core sample. If one is included, disclose the relationship, and have Ana neither attend nor receive attributable notes.

Because five at-risk accounts are in renewal talks and receiving discount offers, use an interviewer outside the account/renewal team. Tell participants their answers won’t affect service or renewal discussions. Don’t share interview content with the account team in a way that could be used in negotiation. Follow Legal’s recording requirement below.

45-minute guide: churned customers

TimeQuestions
0–3Start by asking: “May I record this conversation?” If no, take notes instead. Explain the purpose, that there are no sales or renewal implications, and that we won’t discuss pricing, discounts, or roadmap.
3–7“What was your role in choosing and using Crate?” “What did your warehouse need the system to help you do?”
7–18“Take me through your experience with Crate, from signing through the decision to leave.” “When did you first consider switching? What happened just before that?” “Who was involved, and when was the decision effectively made?”
18–27“How did implementation and go-live go?” “Where, if anywhere, did it take longer or require more effort than expected?” “What was the impact on your operation?” “When did you first get value from Crate?”
27–35“Did automation or robotics play a role in your warehouse during this period?” If yes: “What equipment or workflow?” “How did it work with Crate?” “What did you do when it didn’t?” “When did this become a problem relative to the decision to leave?”
35–41“What alternatives did you consider, and what drove the choice?” “What, if anything, could realistically have changed your decision?” “What was the most important reason—not just one frustration?”
41–45“What haven’t I asked that matters?” “Is there anything we should correct about our understanding?” Thank them.

45-minute guide: red-health customers

Use the same opening and close. Do not imply that Crate can fix an issue or influence renewal terms.

TimeQuestions
0–3Start by asking: “May I record this conversation?” If no, take notes. Explain that this is research, not a sales or renewal conversation; participation won’t affect service or renewal discussions.
3–7“What is your role in using or overseeing Crate?” “What does a successful warehouse operation look like for you?”
7–17“How has your experience with Crate unfolded so far?” “When have things gone better or worse than expected?” “What is working well today?”
17–27“Where are you in implementation and go-live?” “What has taken the most effort or time?” “What value, if any, have you realized so far?” “What is the operational impact of anything that remains difficult?”
27–35“How, if at all, does automation or robotics fit into your operation?” If relevant: “What equipment or workflow?” “How does it work with Crate today?” “What workarounds do you use?” “When did this become important?”
35–41“Have you considered alternatives to Crate? What prompted that, if so?” “What would need to be true for Crate to work well for your operation?” “What are the most important unresolved issues?”
41–45“What haven’t I asked that matters?” “Is there anything we should correct?” Thank them.

Guidance for interviewers

  • Ask the questions as written, but follow the participant’s timeline. Use neutral probes: “What happened next?” “Can you give me an example?” “How do you know?”
  • Don’t offer Ana’s or Marcus’s theories, mention the exit-survey results, or ask leading questions such as “Would robots have saved you?” Ask about events before asking for interpretations.
  • Separate when an issue occurred from when the participant says it mattered. Probe for specific actions and evidence: evaluations, workarounds, escalations, or a switch decision.
  • Don’t discuss pricing, discounts, renewal negotiations, or roadmap. If asked about roadmap, say you can’t speak to future plans. Don’t make promises or relay an interview answer as a renewal signal.
  • Interview roles separately. Record role and account context in notes, but distinguish direct experience from second-hand views. Afterward, write a short timeline, leading cause(s), evidence, and uncertainties for each account.
  • Treat recordings and notes as confidential; use only with consent and Legal-approved handling. Don’t present anecdotes as representative of all customers.

What would support each bet?

Support Bet A (robotics integrations): In multiple independent accounts, a specific robotics workflow or integration gap caused meaningful operational friction, and that issue arose before the decision to leave or seriously consider leaving. Stronger evidence: an actual robot deployment or concrete plan, a costly workaround, and a competitor selected or evaluated for this capability. A feature mention alone—or robotics cited in a new-logo deal—doesn’t establish that it drove existing-customer churn.

Support Bet B (faster implementation): In multiple accounts, implementation delays or effort materially postponed go-live or first value, caused operational consequences, and preceded the decision to leave or consider leaving. Stronger evidence: specific preventable blockers, an unmet expectation, and a credible account that earlier go-live would have changed the outcome. The 94-day median, 45-day promise, and 2.4× association make this worth testing, but do not prove causation.

Support neither / remain uncertain: The call evidence points more strongly to another cause; A or B appears only after the decision; accounts describe issues without meaningful consequences; or evidence is mixed and no cause consistently precedes churn. Don’t force a winner from 12 interviews. Compare patterns across accounts, not individual interview counts, and triangulate with CRM timelines and other churn evidence. The exit survey’s single-choice format, first-listed feature option, and sparse free text limit what it can establish.

Grades and run details

Decision model 92 · LLM judge 11 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly68%
  • passAddresses the actual decision92%
  • passRespects explicit constraints27%
  • passIdentifies material uncertainty94%
  • passAvoids unsupported claims89%
  • passProduces the required deliverable75%
  • passTests both theories fairly99%
  • partialDifferent questions for user and signer71%
  • passProtects the calls and the accounts76%
  • passDesigned to falsify77%
  • passAsks what people did, not what they think68%
  • passDoesn't lead the witness65%
  • partialMarks what to cut if the call runs over50%
Run
Run
#1
API response time
29 s
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 11

Uses the supplied evidence correctlyWrongWrongRight
Gemini 3.8 Flash · API

The guide relies on several unsupported current-situation or portfolio claims, including CRM clustering, robotics surfacing without prompting, first-year churn majority, >45-day correlation, and exit-survey meaning.

Gemini 3.5 Flash-Lite · Gemini

It uses most supplied numbers correctly, but invents or overstates current-situation claims such as “cash flow and client onboarding speed are everything” and “90+ days for most warehouses” when the context only gives a 94-day median.

GPT-6 Luna · API

Every statement about the current situation is taken directly from the supplied context, with no invented facts.

Addresses the actual decisionMixedRightRight
Gemini 3.8 Flash · API

It does not commit to Bet A, Bet B, or neither; it defers the decision to the research, even though the supplied evidence is not framed as insufficient to make a call.

Gemini 3.5 Flash-Lite · Gemini

It does not choose a bet now, but the supplied evidence cannot support a final call; it clearly says the research will settle the decision and gives explicit thresholds for Bet A, Bet B, or neither.

GPT-6 Luna · API

The guide clearly states what evidence from the calls would back Bet A, Bet B, or neither, which is the decision framework the exec team needs.

Respects explicit constraintsWrongWrongRight
Gemini 3.8 Flash · API

It exceeds the requested 1,500-word limit and its proposed 6-account roster would not cover the 12 accounts specified in the supplied call plan.

Gemini 3.5 Flash-Lite · Gemini

It is under 1,500 words and includes the required sections, but it does not enforce the no-pricing/no-roadmap rule in practice because it lacks a script for when customers raise those topics and even mentions “roadmap” in the executive intro.

GPT-6 Luna · API

The guide respects all constraints: it includes learning goals, timed questions for both customer types, interviewer guidance, plan changes, and decision signals; it stays under 1,500 words; it enforces no pricing/roadmap discussion and recording consent.

Avoids unsupported claimsWrongWrongRight
Gemini 3.8 Flash · API

It presents several interpretations and portfolio signals as established facts or evidence, such as exit-survey meaning, churn distribution, and robotics mentions without prompting.

Gemini 3.5 Flash-Lite · Gemini

It presents several interpretations as established facts, including that the exit survey “heavily favors” Ana’s view, that founder-led interviews carry “massive” politeness bias, and that cash flow/onboarding speed are “everything” for mid-size 3PLs.

GPT-6 Luna · API

Interpretations like the exit survey's limitations and the 2.4× association are clearly labeled as not proving causation, not presented as established fact.

Produces the required deliverableMixedWrongRight
Gemini 3.8 Flash · API

Although it is a usable interview guide, it is too long and changes the account scope in a way that does not satisfy the supplied 12-call plan.

Gemini 3.5 Flash-Lite · Gemini

The guide is mostly present, but it is not fully usable because it changes the plan to split personas without specifying how the 12 calls and 6 accounts will be covered, and it omits a required handling script for pricing/roadmap questions.

GPT-6 Luna · API

The output is a complete call guide with all requested sections, written for the exec team, and well under 1,500 words.

Tests both theories fairlyRightWrongRight
Gemini 3.8 Flash · API

It gives both robotics and implementation disconfirming questions and explicitly leaves room for neither, including stability, billing/EDI, and macro causes.

Gemini 3.5 Flash-Lite · Gemini

It gives both theories probes, but the daily-user track is weighted toward implementation, the signer track asks about robotics only after implementation, and the wrap-up asks which bet would have saved the account, so it does not test both theories fairly.

GPT-6 Luna · API

Both the robotics and implementation theories get questions that could disprove them, open timeline questions come first, and the guide explicitly leaves room for a third cause.

Different questions for user and signerRightRightWrong
Gemini 3.8 Flash · API

It provides separate tracks for the daily warehouse operations manager and the VP/COO signer, with role-appropriate questions.

Gemini 3.5 Flash-Lite · Gemini

It has distinct tracks: the daily user is asked about onboarding, workarounds, floor equipment, and the breaking point, while the signer is asked about purchase promises, time-to-value, robotics strategy, and renewal decision timing.

GPT-6 Luna · API

The guide does not provide distinct question tracks for the daily warehouse operations manager and the executive who signed; it only differentiates churned vs at-risk customers, with the same questions for both roles within each track.

Protects the calls and the accountsRightWrongRight
Gemini 3.8 Flash · API

It includes recording consent, a no-pricing/no-roadmap rule with a script, and a concrete protocol for Ana's calls.

Gemini 3.5 Flash-Lite · Gemini

It includes consent to record and a no-pricing/no-roadmap rule, but it provides no script for handling pricing, discount, or roadmap questions and the executive intro mentions roadmap changes.

GPT-6 Luna · API

It gives a clear rule and script for pricing/roadmap questions, starts with recording consent, and addresses Ana's bias by recommending she not lead calls and, if a contact is included, that she not attend or receive attributable notes.

Designed to falsifyRightWrongRight
Gemini 3.8 Flash · API

Each learning goal has questions that could falsify the relevant hypothesis, such as actual automation deployment, go-live timeline, and alternative causes.

Gemini 3.5 Flash-Lite · Gemini

Not every learning goal has a neutral disconfirming question; for example, the implementation goal asks whether delays “permanently damaged trust,” and the robotics goal asks whether robotics was “the straw that broke the camel's back,” rather than first eliciting the decision story.

GPT-6 Luna · API

Every learning goal has questions whose honest answers could disprove the hypothesis, such as asking when automation became a problem relative to the decision to leave.

Asks what people did, not what they thinkRightWrongRight
Gemini 3.8 Flash · API

The core questions ask for recent concrete behavior, including deployed equipment, go-live experience, decision timing, and workarounds.

Gemini 3.5 Flash-Lite · Gemini

Some questions ask about recent actions, but the core includes many opinions, hypotheticals, and leading prompts such as “magic wand,” “how heavy was the absence of robotics,” and “which one would have saved your account.”

GPT-6 Luna · API

Core questions ask for specific past events and actions (e.g., 'When did you first consider switching? What happened just before that?'), with opinions and predictions clearly secondary.

Doesn't lead the witnessRightWrongRight
Gemini 3.8 Flash · API

The questions are mostly neutral and avoid asking customers to choose between the bets or pitching the ideas.

Gemini 3.5 Flash-Lite · Gemini

Several questions are leading or biased, including “Our data shows it takes about 90+ days,” “did the slow start permanently damage your trust,” “cash flow and client onboarding speed are everything,” and the final either/or bet question.

GPT-6 Luna · API

Questions are neutral and open; the guide avoids mentioning either executive's theory until after the participant's own timeline, and prohibits leading probes.

All got wrong 1

Marks what to cut if the call runs overWrongWrongWrong
Gemini 3.8 Flash · API

The timings add up to 45 minutes, but the number of questions is too dense for a realistic 45-minute call and there are no must-ask markers or cut-down instructions.

Gemini 3.5 Flash-Lite · Gemini

Timings add to 45 minutes, but the number of questions is too high for 45 minutes, must-ask questions are not marked, and there is no guidance on what to cut if time runs short.

GPT-6 Luna · API

Although timed sections add up to 45 minutes, the guide does not mark must-ask questions or say what to cut if time runs short, so an interviewer could run out of time without covering the most critical probes.

All got right 1

Identifies material uncertaintyRightRightRight
Gemini 3.8 Flash · API

It names material unknowns such as timing of churn decisions, real robotics deployment, onboarding drag versus customer-side disorganization, and third causes, and proposes calls to resolve them.

Gemini 3.5 Flash-Lite · Gemini

It names the key unknowns—when customers decided to leave, whether “missing features” means robotics or core WMS, and whether slow implementation caused trust loss—and says the calls and decision thresholds would resolve them.

GPT-6 Luna · API

It names key unknowns (exit survey weakness, correlation vs causation, timing of churn decision) and says how the calls will resolve them through patterns and timelines.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review