Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Crate. In four weeks, the exec team decides which of two bets gets our two squads next half, and both bets rest on a theory of why customers are leaving. We have three weeks for 12 customer calls of 45 minutes. Write the guide for them. It should include: the learning goals; the questions, with timings, for the two kinds of people we'll talk to; guidance for whoever runs the calls; any changes you'd make to the plan below; and what we'd need to hear to back each bet, or neither. Keep it under 1,500 words. The exec team will read it before the calls start.

What the model was given8 items: About Crate, The decision, CEO, Ana Ferreira, CPO, Marcus Hill, Exit survey (41 churned customers, last 12 months), Implementation and CRM data, The call plan, Notes from Customer Success and Legal
About CrateWarehouse management software for mid-size third-party logistics companies (3PLs). 340 customers, $22M ARR. Gross logo churn rose from 11% to 18% over the last year.
The decisionBet A: integrations with warehouse robotics (autonomous carts, pick-assist robots). Bet B: rebuild implementation so customers go live faster. Two squads for the next half go to one of them.
CEO, Ana Ferreira“We're losing warehouses because they're automating and we don't talk to their robots. Every lost customer I've spoken to mentions it.”
CPO, Marcus Hill“We lose them in the first year because implementation takes forever. They're gone before they ever see the value.”
Exit survey (41 churned customers, last 12 months)A single-choice question, 'Why are you leaving?'. Missing features or integrations: 46%. Too hard to implement: 22%. Price: 17%. Other: 15% (free-text box, left blank by all but two). 'Missing features or integrations' was the first option in the list.
Implementation and CRM dataMedian time from signing to go-live: 94 days; the contract promises 45. Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year. 3 churned customers moved to a competitor with robotics integrations; robotics was cited in 9 lost new-logo deals. We don't know when in the year churned customers decided to leave.
The call plan12 calls: 6 with customers who churned in the last 9 months, 6 with at-risk customers (health score red). On each account we want to talk to the warehouse operations manager who used Crate every day and, separately, the VP of Operations or COO who signed the contract. Ana wants to run 4 of the calls with former customers she knows personally.
Notes from Customer Success and Legal5 of the 6 at-risk accounts are in renewal talks, and Customer Success is offering them discounts. Interviewers must not discuss pricing, discounts or what's on our roadmap. Every call must start by asking permission to record.
What a strong answer doesThe answer key the graders mark against

A guide that tests both theories fairly and leaves room for a third: the goals ask what actually drove each decision to leave or doubt, when it happened, and who made it, with questions that could disprove robotics (did they automate, did robots come up before they decided, what did they switch to and why) and implementation (how long go-live took, what happened in those months, whether value arrived) alike. It treats the exit survey as weak evidence (single choice, integrations listed first, 'Other' left blank) and the 2.4× figure as a correlation to probe, not a cause. The core is a timeline of the account: the story from signing to the moment they decided to leave or started doubting, with sequential probes. It has separate tracks for the daily user (what went wrong day to day) and the signer (how the decision was made, what alternatives they weighed), fits 45 minutes with timings, and opens with consent to record. It guards the calls: no pitching, pricing, discounts or roadmap promises, with a script for when an at-risk customer asks, and it addresses Ana's bias: she has a public view and personal relationships, so she should observe or take notes rather than lead, or at minimum use the same script and have her calls read separately. It ends with decision signals: what pattern across the 12 would back Bet A, what would back Bet B, and what would back neither (price, a business change, a missing capability that isn't robotics).

Critical failures (cap the score at 40)
  • Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks
  • Asks customers to choose between robotics and implementation as the reason they left
Case

Staff-level · v1.1 · synthetic · B2B SaaS, logistics, churn, exec bias, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit12 of 13 checks passed

The guide is thorough and evidence-based; a PM could add a note on which questions to prioritize if time is tight and use it immediately.

Output

Crate churn research: executive pre-read and interview guide

Decision: Allocate two squads next half to robotics integrations (A), implementation rebuild (B), or neither if the evidence does not justify either.

Research stance: Test both explanations, not confirm either executive’s account. These interviews can reveal mechanisms and decision timelines; they cannot estimate their prevalence across 340 customers or prove causality.

1. Learning goals

  1. Reconstruct the departure process: When did concerns emerge, when did leaving become likely, and when was the decision effectively made?
  2. Identify consequential problems: Separate reasons volunteered after leaving from problems that materially changed warehouse outcomes and the buying decision.
  3. Test Bet A: Was a specific robotics incompatibility a binding obstacle? Would addressing it have plausibly retained the account?
  4. Test Bet B: Did implementation delay prevent value, damage trust, or trigger departure? What caused the delay, and could Crate control it?
  5. Find alternatives and interactions: Price, service, reliability, business changes, or other missing capabilities may dominate. Robotics work may itself complicate implementation.
  6. Connect learning to an investable scope: What could two squads deliver next half that would address the demonstrated retention mechanism?

What the existing evidence does—and does not—say

  • The exit survey’s broad integration category, single-choice format, first-position placement, and sparse free text do not establish robotics as the cause.
  • The 2.4× churn association makes implementation worth investigating, but complexity, customer readiness, or other factors could drive both delay and churn.
  • Three competitor moves and nine lost deals indicate a robotics opportunity; acquisition losses are not evidence of a retention mechanism.
  • Unknown decision dates make chronology essential.

2. Changes to the call plan

Resolve the arithmetic: Twelve separate interviews cannot cover twelve accounts with two roles each. Use six accounts × two separate role interviews = twelve 45-minute calls.

Recommended mix: Four churned accounts and two at-risk accounts. This reduces reliance on renewal negotiations while retaining a prospective view. Findings will be depth-oriented, not representative.

Select accounts from the full CRM population, not executive contacts or survey answers alone:

  • Across churned accounts, include slow and faster implementations, first-year and later departures, and robotics-exposed and non-exposed customers.
  • Include at least one confirmed move to a robotics-enabled competitor and one slow implementation without an apparent robotics need.
  • For at-risk accounts, prefer one with active automation plans and one with implementation/value-realization problems. Include the non-renewal account if suitable.
  • Seek contrasting retained accounts through existing operational data, even though the call budget does not cover them.

Recruit both the daily warehouse operations manager and the contract-signing VP Operations/COO. If the original signer has left, recruit someone directly involved in the departure or renewal decision and document that substitution.

Ana should not lead calls with her personal contacts. Use a neutral interviewer without renewal responsibility. Ana can help recruit; observing requires explicit customer consent. Her existing conversations are useful leads, not independently verified findings.

CS should introduce research as separate from renewal discussions, then step out. Participation must not affect commercial treatment.

3. Interview guides: 45 minutes each

Use the relevant role column. In every section, distinguish direct experience from hearsay. Ask broad questions before naming either bet.

TimeDaily warehouse operations managerVP Operations / COO
0–3 min: permission and framingFirst ask: “May we record this conversation for internal research?” If declined, continue with notes. Then explain purpose and confidentiality limits.Same opening.
3–7 min: context and expected value“What did your warehouse handle, and what was your role with Crate? What was supposed to improve?”“What led you to buy Crate? What outcomes and deadlines mattered? How would you judge success?”
7–17 min: chronological reconstruction“Walk me from signing through onboarding, first operational use, and the point when problems became serious.” “Describe a particular shift or incident.” Establish dates, workarounds, operational impact, and who knew.“Walk me from purchase to the first concern, evaluating alternatives, and the decision.” “When did staying stop being the default?” Establish dates, decision-makers, triggers, and alternatives.
17–27 min: implementation and value“What had to happen before you could use Crate successfully? Where did work stall? Who owned each step?” “When did you first achieve useful results?” Probe migration, configuration, training, integrations, staffing, and rework after an open answer.“What go-live date did you expect, and what happened? What consequences did that have?” “What other factors accompanied the delay?” “If the same product had gone live in 45 days, what would still have put the relationship at risk?” Ask why.
27–36 min: workflow, automation, and unmet needs“Which workflows did Crate support poorly? Show or describe a recent example.” Then: “Were robots in use or planned? Which systems, what workflow, and what connection was required?” “What happened without it?”“What capabilities influenced staying or switching?” Then probe automation: vendor, deployment date, committed budget, required integration, and alternatives considered. “If that connection had existed, what else would have needed to change for you to stay?”
36–42 min: decision and competing explanations“Which problem mattered most in daily operations? What did you escalate, to whom, and when?” “What worked well?” “What have we missed?”“Which issues were necessary to the decision, and which were secondary?” “What would have had to be different for you to stay?” Ask about business changes, service, reliability, and other alternatives without forcing a category.
42–45 min: verify and closeSummarize the timeline and mechanism: “What have I misunderstood?” Request relevant artifacts and permission for a brief clarification follow-up.Same; verify whether operational problems actually influenced the commercial decision.

Status-specific wording

  • Churned: Ask what the replacement actually delivered, whether it went live, and whether the cited problem improved—not merely what the competitor promised.
  • At-risk: Ask what has already happened, current unresolved consequences, and what would cause escalation or departure. Do not imply a departure decision exists. Distinguish funded plans from aspirations.

4. Guidance for interviewers

  • Begin with recording permission; record only after consent. Explain that this is research, not a support, roadmap, or renewal discussion.
  • Do not discuss prices, discounts, or roadmap commitments. If raised, acknowledge without probing commercial terms: “I can’t discuss that here; your account team handles commercial conversations.” Record volunteered price concerns as an alternative explanation.
  • Avoid “Did missing robotics make you leave?” and “Would faster onboarding have saved you?” Start with events; use counterfactuals only afterward.
  • Ask for concrete examples, dates, artifacts, and consequences. “It was painful” is not yet a mechanism.
  • Do not defend Crate, disclose executive hypotheses, or offer fixes.
  • Interview roles separately before reconciling accounts. Treat disagreement as evidence, not an error to erase.
  • Use a note-taker and common template: timeline, problem, consequence, decision link, alternative explanation, evidence source, and confidence. Count accounts, not two interviewees as independent votes.

5. Evidence required for the decision

Back A: robotics integrations

Look for a repeated chain across independent accounts:

Committed robotics deployment → specific missing connection → material operational/business constraint → documented role in departure or credible near-term risk.

Stronger evidence includes decision-maker confirmation, pre-decision correspondence, and a replacement that solved the problem. Match the implicated vendors/workflows to current customer exposure. Before funding, establish that a feasible integration scope could address meaningful retention risk next half.

Mentions of “automation,” hypothetical plans, or competitor marketing are insufficient.

Back B: implementation rebuild

Look for:

Crate-controllable implementation bottleneck → delayed usable value or broken operational commitments → loss of confidence/avoidable cost → departure decision before value realization.

Seek repeated bottlenecks, role-level corroboration, and contrasts with faster implementations. Validate the mechanism against the full cohort, accounting where possible for complexity and customer readiness. Identify changes two squads could deliver; do not assume a wholesale rebuild is necessary.

Back neither—or declare the result inconclusive

Neither is justified if dominant causes lie elsewhere, robotics demand is narrow or speculative, delays are largely outside Crate’s control, or either intervention arrives too late to affect retention. If interviews remain contradictory or unsupported, report insufficient evidence, not “neither causes churn.”

Compare both bets on reachable ARR at risk, evidence strength, controllability, delivery feasibility, and time to impact—not mention counts.

6. Three-week execution

  • Week 1: Audit cohorts and timelines, recruit, pilot one role pair.
  • Week 2: Complete remaining interviews; synthesize after each pair without changing standards to favor an emerging winner.
  • Week 3: Verify discrepancies, triangulate with records, assess engineering scope, and prepare the decision memo with supporting evidence, counterexamples, uncertainties, and a recommendation.

Do not commit two squads merely because one story sounds more persuasive.

Grades and run details

Decision model 92 · LLM judge 13 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly54%
  • passAddresses the actual decision90%
  • partialRespects explicit constraints21%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims96%
  • passProduces the required deliverable89%
  • passTests both theories fairly100%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts95%
  • passDesigned to falsify63%
  • passAsks what people did, not what they think67%
  • passDoesn't lead the witness53%
  • partialMarks what to cut if the call runs over58%
Run
Run
#1
API response time
80 s
Submitted
30 Sept 2026
Critical failure: Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks. Score capped at 40.

Needs more than a quick edit3 of 13 checks passed

A PM would need to fix the prohibited either/or question, add a pricing/roadmap handling script, remove unsupported claims, and clarify the revised call coverage before using it. The reasoning and evidence also need rechecking.

Output

Customer Research Guide: H2 Strategic Bets (Robotics vs. Implementation)

To: Executive Team From: Staff Product Manager, Crate Date: Today Subject: 3-Week Customer Research Plan for H2 Resource Allocation (12 Calls)

---

Executive Summary & Plan Adjustments

In four weeks, the executive team will decide where to allocate both engineering squads for the next half. Ana believes we are losing customers because we lack warehouse robotics integrations (Bet A). Marcus believes we are losing them because our implementation cycle is broken and slow (Bet B).

The exit survey heavily favors Ana’s view (46% cite missing features/integrations), but UX research 101 teaches us that exit surveys are structurally flawed. Listing "missing features" first creates primacy bias, and the survey failed to capture temporal dynamics—specifically, when the decision to leave was made. Meanwhile, our internal data tells a damning story: our median time to go-live is 94 days (vs. 45 promised), and customers crossing 90 days churn at $2.4\times$ the rate.

We have 3 weeks to conduct 12 interviews (45 minutes each) to uncover the truth.

Critical Plan Adjustments (Read Before We Start)

  1. Split the Interviewees by Persona, Not Account: The original plan called for interviewing both the Ops Manager and the VP/COO on the same accounts. In 45 minutes, splitting time between a daily user and a high-level signer dilutes the insight. Instead, we will interview Warehouse Ops Managers to understand day-to-day friction (Implementation, UI, operational bottlenecks) and VPs of Ops / COOs to understand strategic drivers (ROI, contract expectations, robotics, and when the churn decision was actually made).
  2. Revise Ana’s Participation: Ana wants to run 4 calls with former customers she knows personally. We must advise against this. Founder-led interviews with churned customers carry massive politeness bias; customers will pull punches about implementation failures or invent reasons to let Ana down gently. Ana should observe these calls anonymously or listen to the recordings, but an unbiased PM or User Researcher must conduct them.
  3. Guardrails for At-Risk Accounts: Five of our six at-risk accounts are in active renewal talks with CS discounts on the table. Interviewers are strictly prohibited from discussing pricing, commercial concessions, or our product roadmap. We are strictly here to diagnose reality.

---

Learning Goals

  1. The Churn Timeline: Determine the exact month the customer mentally decided to leave. Did they decide during a painful 90+ day onboarding (supporting Bet B), or later when evaluating modernizing warehouse automation (supporting Bet A)?
  2. The "Missing Features" Audit: Deconstruct the 46% exit survey metric. When respondents checked "missing features/integrations," were they missing critical WMS functionality (e.g., wave planning, inventory accuracy) or specifically robotics middleware?
  3. Implementation Reality Check: Measure the operational impact of our 94-day median go-live time. Did the delay create cash-flow/trust erosion that doomed the account before day one, regardless of subsequent features?
  4. Robotics Urgency: For customers who mention robotics, are they actively deploying autonomous carts today, or is it a 3-year roadmap item they use as a stick to beat us with during negotiations?

---

Interview Guide: Two Personas

Persona 1: The Daily User (Warehouse Operations Manager)

Focus: Implementation friction, daily utility, operational roadblocks.

  • 00:00–03:00 | Intro & Permission:
  • "Thanks for taking the time. I’m [Name] from Crate. We’re doing research to figure out where we failed our customers and how we can improve. I am strictly here to learn about your experience—we won't be talking about pricing, renewals, or selling you anything today. Do I have your permission to record this call for internal notes?"
  • 03:00–12:00 | Day-to-Day Workflow & Reality:
  • "Take me back to your first 30 days using Crate. What did the day-to-day onboarding look like?"
  • "How long did it actually take for your floor team to feel comfortable using Crate without holding their breath?"
  • "What was the most painful workaround your team had to invent during your first three months?"
  • 12:00–25:00 | The Implementation Experience (Bet B Probe):
  • "Our data shows it takes about 90+ days for most warehouses to go live on Crate. What was your experience? How did those delays affect your team's stress levels and your relationship with our implementation team?"
  • "Looking back, did the slow start permanently damage your trust in the software, or did you bounce back once things were running?"
  • 25:00–35:00 | Equipment & Floor Integration (Bet A Probe):
  • "What automated hardware or mobile carts are running on your floor today? How do they talk to Crate, or do they?"
  • "When a pick-assist cart or AGV moves through your warehouse, where does the process break down between the robot and Crate?"
  • 35:00–42:00 | The Breaking Point:
  • "Think about the moment you realized you wanted to leave Crate (or started looking at alternatives). What happened on the warehouse floor that week?"
  • 42:00–45:00 | Wrap-Up:
  • "If you could wave a magic wand and change one thing about your onboarding or your day-one experience with Crate, what would it be?"

---

Persona 2: The Decision Maker (VP of Operations / COO)

Focus: Strategic expectations, ROI timelines, contract promises, and robotics strategy.

  • 00:00–03:00 | Intro & Permission:
  • "Thanks for joining. I’m [Name] from Crate. We’re reviewing why customers leave us to make radical changes to our roadmap and processes. I'm not here to pitch you or talk commercial terms—just to understand your true experience. May I record this call for note-taking purposes?"
  • 03:00–10:00 | The Purchase & Promise:
  • "When you signed with Crate, what was the primary business problem you needed solved in your first 6 months?"
  • "What timeline were you given for go-live during the sales cycle, and how did reality compare?"
  • 10:00–22:00 | Implementation & Time-to-Value (Bet B Probe):
  • "For a mid-size 3PL, cash flow and client onboarding speed are everything. How did Crate's onboarding timeline impact your ability to onboard your own end-clients?"
  • "At what point in your first year did you officially consider the implementation 'failed' or deeply troubled?"
  • 22:00–34:00 | The Robotics Mandate (Bet A Probe):
  • "Let’s talk about warehouse automation. How central are autonomous carts or robotics to your strategy today versus three years from now? Are you actively running them?"
  • "When you evaluated staying vs. leaving Crate, how heavy was the absence of robotics integrations in that final calculus? Was it the straw that broke the camel's back, or a nice-to-have?"
  • 34:00–42:00 | Uncovering the Timeline:
  • "Looking back at the lifecycle of your account, when did you mentally check out and decide you wouldn't renew? Was it during implementation, or months later?"
  • 42:00–45:00 | Wrap-Up:
  • "If Crate could fix only one thing—either our onboarding speed or our hardware/robotics connectivity—which one would have saved your account?"

---

Decision Criteria: What We Need to Hear

To make our final recommendation to Ana and Marcus, we will evaluate the 12 calls against the following thresholds:

Backing Bet A (Robotics Integrations)

We back Bet A only if: 1. VPs/COOs explicitly state that the lack of robotics was the primary, active catalyst for churning (not just a grievance listed in a post-hoc survey). 2. Churned customers are actively utilizing or piloting autonomous hardware today that could not function with Crate, resulting in operational deadlocks. 3. The exit survey's 46% "missing features" bucket is proven via interviews to be overwhelmingly about hardware/robotics connectivity rather than core WMS gaps (e.g., billing, inventory tracking).

Backing Bet B (Implementation Rebuild)

We back Bet B only if: 1. Interviews reveal that the 90+ day onboarding process created an insurmountable trust deficit; customers felt abandoned and never recovered faith in the product, rendering subsequent features irrelevant. 2. Daily users report that the complexity of setup and configuration was the primary driver of operational frustration, overshadowing missing integrations. 3. Churned customers admit they mentally checked out during or immediately after the botched implementation window, long before robotics ever entered the conversation.

Backing Neither (The Pivot Option)

We back Neither if: * The qualitative data shows a split narrative where both bets are symptoms of a deeper root cause—such as core WMS unreliability, pricing/value misalignment (17% in exit survey), or poor account management—meaning that building robots or speeding up a broken implementation will simply churn customers faster down a different path.

Claims the judge couldn’t find in the brief

  • For a mid-size 3PL, cash flow and client onboarding speed are everything.

Grades and run details

Decision model 58 · LLM judge 3 of 14 checks
Decision model checks
  • partialUses the supplied evidence correctly23%
  • passAddresses the actual decision93%
  • partialRespects explicit constraints21%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims46%
  • partialProduces the required deliverable26%
  • partialTests both theories fairly47%
  • passDifferent questions for user and signer98%
  • partialProtects the calls and the accounts84%
  • partialDesigned to falsify40%
  • partialAsks what people did, not what they think52%
  • failDoesn't lead the witness30%
  • partialMarks what to cut if the call runs over45%
Run
Run
#1
Time to output
11 s
Submitted
1 Oct 2026

Needs more than a quick edit8 of 13 checks passed

The guide is strong and actionable, but it is over the word limit and contains a few unsupported factual claims that need correction before use.

Output

Churn Discovery Calls: Discussion Guide

For: exec team (read before calls start), interviewers, note-takers Timeline: week 0 for data pull and recruiting, weeks 1–3 for calls, week 4 for synthesis and recommendation

Why we're doing this

Both bets rest on a theory we haven't tested. The data we have can't settle it:

  • The exit survey doesn't separate the theories. "Missing features or integrations" (46%) was listed first, it was a single choice, and it combines robotics with every other feature gap. The free-text box was almost always left blank.
  • The robotics evidence is thin for churn. 3 churned customers moved to a competitor with robotics integrations. That is 3 of roughly 60 churned accounts. The 9 lost new-logo deals are about acquisition, not churn.
  • The implementation evidence is a correlation. Customers who went live after 90 days churned at 2.4× the rate. But complex customers may both take longer to go live and churn more for other reasons.
  • We don't know when or why the leave decision formed. That is the question the calls exist to answer.

Twelve calls can show us how customers decide to leave. They can't show us how many leave for each reason. The CRM pull below fills that gap, and the calls and the data should be read together.

1. Learning goals

  1. When did the leave decision form, and what triggered it? In the first months or later? During implementation, after go-live, or around an automation project?
  2. Does implementation delay cause churn, or just come with it? If delay mattered, was the cause on our side (fixable by Bet B) or on theirs (data readiness, their staffing)?
  3. Is automation a real driver? Are customers actually automating? Did Crate's lack of robot integration block them, or did it come up afterwards as a reason?
  4. Do signers and daily users tell the same story? Executives may cite strategy (automation). Operations managers may cite experience (implementation, usability).
  5. Are the reasons something else entirely? For example: price, support, acquisition, consolidation, or a non-robotics feature gap.

2. Changes to the plan

1. Pull CRM data in week 0, before the first call. For every account that churned in the last 12 months, pull: - tenure at churn - days from signing to go-live - whether they run or were deploying automation (ask CS and account managers) - where they went

Also answer one question: did the rise from 11% to 18% churn come from first-year customers? If it did, that supports Bet B. If the increase is among long-tenured accounts, that points toward Bet A. This may be the most decisive single fact we get, and it costs a day.

2. Replace at-risk accounts that are in renewal talks. Five of the six are negotiating discounts. They have a reason to exaggerate complaints, and a call with them could easily drift into pricing. Use red-health accounts that are not in active renewal. If there aren't enough of those, add churned accounts instead.

3. Fix the call count. The plan says 12 calls, but two people on each of 12 accounts is 24 calls. Within the 12, I propose 8 accounts:

GroupAccountsWho we interviewCalls
Churned, paired4Signer and ops manager, separately8
Churned, signer only2Signer2
At-risk, not in renewal2Signer2
  • Split the churned accounts so that 3 left in their first year and 3 left later.
  • Include at least one account that moved to a competitor with robotics integrations.
  • Signers get priority because they made the decision.
  • The paired accounts let us compare what the signer says with what the daily user saw.
  • If we can fund 24 calls, pair every account.

4. Ana should not interview her own contacts. Anyone will soften or reshape their story when the CEO asks, especially a CEO they know and whose view they may already have heard. Her contacts also aren't a random sample of the people who left.

My suggestion: - Her contacts can enter the pool if they fit the sampling groups, but a neutral interviewer runs those calls. - Ana and Marcus both listen to the recordings. - If Ana wants to speak with former customers personally, those calls are valuable for the relationship. We run them in addition to the 12 and don't count them as research.

The same rule applies to Marcus: neither exec interviews.

5. Offer a thank-you gift to former customers, such as a gift card. It must not be a credit or discount on Crate.

3. Discussion guide: VP Ops / COO (signer), 45 minutes

0–3 min · Opening - Ask permission to record. If they say no, don't record and take notes only. - Say: "This isn't a sales call. I can't discuss pricing or our product plans. I'm here to understand your experience, good and bad."

3–8 min · Context - "Tell me about your business and how it's changed over the last two years." - "What were your biggest operational priorities this past year?"

8–22 min · The decision story (the core of the call) - Churned: "Take me back to the first moment you started to wonder whether Crate was right for you. What was going on?" - At-risk: "Tell me about the last time you seriously questioned renewing." - Build a timeline with follow-ups: - "When was that, relative to signing?" - "What happened next?" - "Who else was involved?" - "What alternatives did you look at?" - "When was the decision final?" - "What finally tipped it?"

22–30 min · Expectations and early months - "When you signed, what did you expect Crate to do for you, and by when?" - "How did going live compare to what you expected?" - "When, if ever, did you first see the value you signed up for?"

30–37 min · Operational direction - "What investments in the warehouse have you made or planned: equipment, systems, staffing?" - "What did you need your WMS to do as part of those?" - If they mention automation: "How did Crate fit, or not? What did you do about it?"

37–42 min · Prompted check - Hand over a card of possible reasons: implementation time, robotics/automation integration, other missing features, ease of use, support, price, business change. - Rotate the order on every call. The survey showed the first option gets picked more. - Ask: "Pick the top three for you." Then: "Which one, if it had been different, would have changed the outcome?"

42–45 min · Close - "What should I have asked?" - "Who else should we talk to?" - Thank them.

4. Discussion guide: Warehouse operations manager, 45 minutes

0–3 min · Opening - Same as the signer call: recording permission and ground rules.

3–10 min · Daily work - "Walk me through a typical day using Crate." - "What other tools or spreadsheets did you use alongside it?"

10–22 min · Implementation story - "Tell me what happened from signing to go-live." - "What took longest?" - "What were you waiting on, and who was it waiting on?" - "When did it start working the way you needed?" - Get specific episodes, not a verdict. Always ask whether a delay came from us, from them, or from a third party.

22–32 min · Gaps and workarounds - "Tell me about the last time Crate got in your way." - "What equipment or automation runs on your floor? How does information move between it and Crate?" - "What manual steps did that create?"

32–37 min · The decision - "Were you part of the decision to leave or renew? What would you have told leadership?"

37–42 min · Prompted check - Same card as the signer call, with a new random order.

42–45 min · Close - Same as the signer call.

5. Guidance for interviewers

Stay neutral - Don't say "robotics," "automation," or "implementation" before the prompted check, unless the interviewee says it first. What people raise unprompted is our strongest evidence. - Ask about specific past events ("the last time…", "walk me through…"), not opinions or hypotheticals. Treat "we'd have stayed if you'd had X" as weak evidence. - Use their words, not ours. Don't correct them or defend Crate. - In paired accounts, use different interviewers where possible. Never tell the second person what the first one said.

Stay within the rules from Legal and CS - Pricing or discounts: "That's outside what I can cover today. I'll let your account team know you'd like to discuss it." Then tell CS. - Roadmap: "I can't speak to product plans, but I'll make sure what you've said reaches the people who make them." Never promise a feature. - If there's no recording consent, don't record. The note-taker writes verbatim quotes where possible.

Process - A note-taker joins every call. - Within 24 hours, fill in the debrief template: - primary reason, in their words - the decision timeline, with the month the doubt started - robotics raised unprompted? (Y/N) - implementation raised unprompted? (Y/N) - was the delay caused by us, them, or both? - top three from the card - the best quote - Hold a 30-minute team review of the debriefs every Friday.

6. What would back each bet

Each account gets classified by its primary cause, meaning the thing that started or tipped the decision. When signer and ops manager disagree, record both views; don't average them.

Back Bet A (robotics) if: - In at least 3 of the 6 churned accounts, automation was raised unprompted and came before the leave decision. - The customer had a live or funded automation project that Crate blocked or made harder. - Ops managers describe concrete workarounds between robots and Crate. - The CRM pull shows churn concentrated in longer-tenured accounts, with meaningfully more automating churners than the 3 we already know about. - Warning sign: robotics only comes up at the card, or appears in the timeline after the decision was already made. That suggests it's a justification, not a cause.

Back Bet B (implementation) if: - In at least 3 of the 6 churned accounts, doubt formed during or soon after a slow implementation, often phrased as "never saw the value" or "lost confidence." - Ops managers put the main delays on our side: our process, staffing, data migration, or configuration. Those are what a rebuild would fix. - The CRM pull shows the rise in churn comes mostly from first-year customers. - Warning sign: delays were mostly on the customer's side (their data wasn't ready, they didn't have staff). A faster process on our end would then help less than the 2.4× figure suggests.

Back neither if: - Primary causes are spread across several reasons with no clear lead, such as price, support, consolidation or acquisition, or a feature gap unrelated to robotics. - Or signers and ops managers consistently tell different stories that neither bet addresses.

In that case, I'd recommend a smaller fix aimed at whichever cause leads, plus a better-designed exit survey, rather than committing two squads on a weak signal.

If both show up: check whether automating customers are also the ones with the slowest implementations. If they are, those customers' problem is really integration work during setup, and the recommendation should address that directly instead of forcing a choice between A and B.

In week 4, I'll bring the exec team a one-page summary that combines the call findings with the CRM figures, a recommendation, and our confidence in it.

Claims the judge couldn’t find in the brief

  • The survey showed the first option gets picked more.

Grades and run details

Decision model 88 · LLM judge 8 of 14 checks
Decision model checks
  • partialUses the supplied evidence correctly8%
  • passAddresses the actual decision95%
  • partialRespects explicit constraints22%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims72%
  • passProduces the required deliverable83%
  • passTests both theories fairly99%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts99%
  • passDesigned to falsify73%
  • passAsks what people did, not what they think90%
  • passDoesn't lead the witness48%
  • partialMarks what to cut if the call runs over45%
Run
Run
#1
Time to output
1.7 min
Submitted
30 Sept 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 9

Uses the supplied evidence correctlyRightWrongWrong
GPT-6.1 Sol · API

All factual statements about the current situation are taken directly from the supplied context without invention.

Gemini 3.5 Flash-Lite · Gemini

It uses most supplied numbers correctly, but invents or overstates current-situation claims such as “cash flow and client onboarding speed are everything” and “90+ days for most warehouses” when the context only gives a 94-day median.

Opus 5.5 · Claude

It invents or overstates a few current-situation facts, including 'roughly 60 churned accounts' and 'the survey showed the first option gets picked more'.

Respects explicit constraintsRightWrongWrong
GPT-6.1 Sol · API

The guide is under 1,500 words, includes all requested sections (learning goals, timed questions for two roles, guidance, plan changes, decision signals), and is addressed to the exec team.

Gemini 3.5 Flash-Lite · Gemini

It is under 1,500 words and includes the required sections, but it does not enforce the no-pricing/no-roadmap rule in practice because it lacks a script for when customers raise those topics and even mentions “roadmap” in the executive intro.

Opus 5.5 · Claude

It exceeds the 1,500-word limit and proposes replacing the at-risk cohort in a way that changes the supplied plan beyond the brief's constraints.

Avoids unsupported claimsRightWrongMixed
GPT-6.1 Sol · API

Interpretations are clearly labelled as such (e.g., the exit survey does not establish robotics as the cause), and no cause is presented as established fact.

Gemini 3.5 Flash-Lite · Gemini

It presents several interpretations as established facts, including that the exit survey “heavily favors” Ana’s view, that founder-led interviews carry “massive” politeness bias, and that cash flow/onboarding speed are “everything” for mid-size 3PLs.

Opus 5.5 · Claude

It presents the first-option bias and the 'roughly 60 churned accounts' figure as established facts when the supplied evidence does not support them.

Produces the required deliverableRightWrongMixed
GPT-6.1 Sol · API

The output is a complete guide with learning goals, timed questions for both roles, interviewer guidance, plan changes, and decision signals; an exec team could act on it.

Gemini 3.5 Flash-Lite · Gemini

The guide is mostly present, but it is not fully usable because it changes the plan to split personas without specifying how the 12 calls and 6 accounts will be covered, and it omits a required handling script for pricing/roadmap questions.

Opus 5.5 · Claude

The guide is complete and usable, but it is over the requested length.

Tests both theories fairlyRightWrongRight
GPT-6.1 Sol · API

Questions could disprove robotics (e.g., 'Were robots in use? What happened without it?') and implementation (e.g., 'Where did work stall? If go-live had been 45 days, what would still have put the relationship at risk?'), and open questions leave room for a third cause.

Gemini 3.5 Flash-Lite · Gemini

It gives both theories probes, but the daily-user track is weighted toward implementation, the signer track asks about robotics only after implementation, and the wrap-up asks which bet would have saved the account, so it does not test both theories fairly.

Opus 5.5 · Claude

It gives both robotics and implementation equal chances to be disproved and leaves room for a third cause.

Protects the calls and the accountsRightWrongRight
GPT-6.1 Sol · API

It forbids discussing pricing/discounts/roadmap, provides a script to deflect, starts with recording consent, and has a concrete plan for Ana (she does not lead, observes with consent, her prior conversations are treated as leads).

Gemini 3.5 Flash-Lite · Gemini

It includes consent to record and a no-pricing/no-roadmap rule, but it provides no script for handling pricing, discount, or roadmap questions and the executive intro mentions roadmap changes.

Opus 5.5 · Claude

It includes recording consent, scripts for pricing/roadmap, and a concrete plan to keep Ana from leading calls with her contacts.

Designed to falsifyRightWrongRight
GPT-6.1 Sol · API

Each learning goal has disconfirming questions (e.g., 'What happened without the robot connection?' for Bet A, 'If the same product had gone live in 45 days, what would still have put the relationship at risk?' for Bet B).

Gemini 3.5 Flash-Lite · Gemini

Not every learning goal has a neutral disconfirming question; for example, the implementation goal asks whether delays “permanently damaged trust,” and the robotics goal asks whether robotics was “the straw that broke the camel's back,” rather than first eliciting the decision story.

Opus 5.5 · Claude

Each learning goal has questions that could falsify the corresponding hypothesis, such as timing, cause of delay, and unprompted automation mentions.

Asks what people did, not what they thinkRightWrongRight
GPT-6.1 Sol · API

Core questions ask for a chronological walk-through, specific incidents, dates, and what they did, not opinions or predictions.

Gemini 3.5 Flash-Lite · Gemini

Some questions ask about recent actions, but the core includes many opinions, hypotheticals, and leading prompts such as “magic wand,” “how heavy was the absence of robotics,” and “which one would have saved your account.”

Opus 5.5 · Claude

The core questions ask for specific past events, timelines, and actions rather than opinions or hypotheticals.

Doesn't lead the witnessRightWrongRight
GPT-6.1 Sol · API

Questions are open and neutral; robotics and implementation are not named until after broad problem questions, and no answer is hinted at.

Gemini 3.5 Flash-Lite · Gemini

Several questions are leading or biased, including “Our data shows it takes about 90+ days,” “did the slow start permanently damage your trust,” “cash flow and client onboarding speed are everything,” and the final either/or bet question.

Opus 5.5 · Claude

It explicitly avoids naming robotics or implementation before the prompted check and uses neutral, open questions.

All got wrong 1

Marks what to cut if the call runs overWrongWrongWrong
GPT-6.1 Sol · API

The guide includes timed sections that add up to 45 minutes but does not say what to cut if time runs short, nor does it mark must-ask questions.

Gemini 3.5 Flash-Lite · Gemini

Timings add to 45 minutes, but the number of questions is too high for 45 minutes, must-ask questions are not marked, and there is no guidance on what to cut if time runs short.

Opus 5.5 · Claude

It provides timings but does not mark must-ask questions or say what to cut if the call runs over.

All got right 3

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

The output recommends more research before deciding, states the supplied evidence cannot yet support a call, and specifies exactly what evidence would back each bet or neither.

Gemini 3.5 Flash-Lite · Gemini

It does not choose a bet now, but the supplied evidence cannot support a final call; it clearly says the research will settle the decision and gives explicit thresholds for Bet A, Bet B, or neither.

Opus 5.5 · Claude

It gives the exec team a clear decision rule for Bet A, Bet B, or neither, with thresholds and warning signs.

Identifies material uncertaintyRightRightRight
GPT-6.1 Sol · API

It names key unknowns (decision dates, whether robotics was a binding obstacle, whether delays were Crate-controllable) and says how the interview evidence would resolve them.

Gemini 3.5 Flash-Lite · Gemini

It names the key unknowns—when customers decided to leave, whether “missing features” means robotics or core WMS, and whether slow implementation caused trust loss—and says the calls and decision thresholds would resolve them.

Opus 5.5 · Claude

It names the key unknowns—timing of the leave decision, whether delay causes churn, and whether automation is a real driver—and says how calls and CRM pulls would resolve them.

Different questions for user and signerRightRightRight
GPT-6.1 Sol · API

The table provides distinct tracks: daily user gets questions about specific shifts, workarounds, and escalations; the signer gets questions about the purchase decision, alternatives, and go-live expectations.

Gemini 3.5 Flash-Lite · Gemini

It has distinct tracks: the daily user is asked about onboarding, workarounds, floor equipment, and the breaking point, while the signer is asked about purchase promises, time-to-value, robotics strategy, and renewal decision timing.

Opus 5.5 · Claude

It provides distinct 45-minute tracks for the signer and the daily operations manager.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review