Tasks / Discover

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

Measures the modelTask v1.0 · 2 casesDifficulty

What AI gets right here, and what you’ll still have to catch

From 14 graded outputs by 7 models. 36% were usable with at most a quick edit.

Reliably right

  1. Identifies material uncertainty96% pass
    Names specific unknowns (cost of lateness, ease of chasing, automation fit, willingness to pay) and says mixed findings mean targeted follow-up.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Coaches first-time interviewers96% pass
    Provides specific, actionable instructions: alternate roles, practice call, avoid pitching, allow silence, separate observations from interpretations.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  3. Addresses the actual decision93% pass
    Provides clear go/stop conditions tied to specific call outcomes, framed for the team.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?

Where it slips

  1. Marks what to cut if the call runs over29% pass
    Timings add up to 30 minutes but no must-ask questions are marked and no guidance on what to cut if time runs short.
    GPT-6.1 Sol · API · Would freelancers pay to stop chasing invoices?
  2. Respects explicit constraints61% pass
    The guide is clearly over 900 words, violating the explicit length constraint.
    Opus 5.5 · Claude · Would freelancers pay to stop chasing invoices?
  3. Uses the supplied evidence correctly70% pass
    It uses most supplied numbers correctly, but invents or overstates current-situation details such as an 'Autochase support ticket' population and treats recruiting-plan inferences as facts.
    Sonnet 5.5 · API · Would freelancers pay to stop chasing invoices?

Case viewer

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're a Staff PM at Crate. In four weeks, the exec team decides which of two bets gets our two squads next half, and both bets rest on a theory of why customers are leaving. We have three weeks for 12 customer calls of 45 minutes. Write the guide for them. It should include: the learning goals; the questions, with timings, for the two kinds of people we'll talk to; guidance for whoever runs the calls; any changes you'd make to the plan below; and what we'd need to hear to back each bet, or neither. Keep it under 1,500 words. The exec team will read it before the calls start.

What the model was given8 items: About Crate, The decision, CEO, Ana Ferreira, CPO, Marcus Hill, Exit survey (41 churned customers, last 12 months), Implementation and CRM data, The call plan, Notes from Customer Success and Legal
About CrateWarehouse management software for mid-size third-party logistics companies (3PLs). 340 customers, $22M ARR. Gross logo churn rose from 11% to 18% over the last year.
The decisionBet A: integrations with warehouse robotics (autonomous carts, pick-assist robots). Bet B: rebuild implementation so customers go live faster. Two squads for the next half go to one of them.
CEO, Ana Ferreira“We're losing warehouses because they're automating and we don't talk to their robots. Every lost customer I've spoken to mentions it.”
CPO, Marcus Hill“We lose them in the first year because implementation takes forever. They're gone before they ever see the value.”
Exit survey (41 churned customers, last 12 months)A single-choice question, 'Why are you leaving?'. Missing features or integrations: 46%. Too hard to implement: 22%. Price: 17%. Other: 15% (free-text box, left blank by all but two). 'Missing features or integrations' was the first option in the list.
Implementation and CRM dataMedian time from signing to go-live: 94 days; the contract promises 45. Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year. 3 churned customers moved to a competitor with robotics integrations; robotics was cited in 9 lost new-logo deals. We don't know when in the year churned customers decided to leave.
The call plan12 calls: 6 with customers who churned in the last 9 months, 6 with at-risk customers (health score red). On each account we want to talk to the warehouse operations manager who used Crate every day and, separately, the VP of Operations or COO who signed the contract. Ana wants to run 4 of the calls with former customers she knows personally.
Notes from Customer Success and Legal5 of the 6 at-risk accounts are in renewal talks, and Customer Success is offering them discounts. Interviewers must not discuss pricing, discounts or what's on our roadmap. Every call must start by asking permission to record.
What a strong answer doesThe answer key the graders mark against

A guide that tests both theories fairly and leaves room for a third: the goals ask what actually drove each decision to leave or doubt, when it happened, and who made it, with questions that could disprove robotics (did they automate, did robots come up before they decided, what did they switch to and why) and implementation (how long go-live took, what happened in those months, whether value arrived) alike. It treats the exit survey as weak evidence (single choice, integrations listed first, 'Other' left blank) and the 2.4× figure as a correlation to probe, not a cause. The core is a timeline of the account: the story from signing to the moment they decided to leave or started doubting, with sequential probes. It has separate tracks for the daily user (what went wrong day to day) and the signer (how the decision was made, what alternatives they weighed), fits 45 minutes with timings, and opens with consent to record. It guards the calls: no pitching, pricing, discounts or roadmap promises, with a script for when an at-risk customer asks, and it addresses Ana's bias: she has a public view and personal relationships, so she should observe or take notes rather than lead, or at minimum use the same script and have her calls read separately. It ends with decision signals: what pattern across the 12 would back Bet A, what would back Bet B, and what would back neither (price, a business change, a missing capability that isn't robotics).

Critical failures (cap the score at 40)
  • Includes questions that promise roadmap items or discuss pricing or discounts with customers in renewal talks
  • Asks customers to choose between robotics and implementation as the reason they left
Case

Staff-level · v1.1 · synthetic · B2B SaaS, logistics, churn, exec bias, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit8 of 13 checks passed

The guide is strong and actionable, but it is over the word limit and contains a few unsupported factual claims that need correction before use.

Output

Churn Discovery Calls: Discussion Guide

For: exec team (read before calls start), interviewers, note-takers Timeline: week 0 for data pull and recruiting, weeks 1–3 for calls, week 4 for synthesis and recommendation

Why we're doing this

Both bets rest on a theory we haven't tested. The data we have can't settle it:

  • The exit survey doesn't separate the theories. "Missing features or integrations" (46%) was listed first, it was a single choice, and it combines robotics with every other feature gap. The free-text box was almost always left blank.
  • The robotics evidence is thin for churn. 3 churned customers moved to a competitor with robotics integrations. That is 3 of roughly 60 churned accounts. The 9 lost new-logo deals are about acquisition, not churn.
  • The implementation evidence is a correlation. Customers who went live after 90 days churned at 2.4× the rate. But complex customers may both take longer to go live and churn more for other reasons.
  • We don't know when or why the leave decision formed. That is the question the calls exist to answer.

Twelve calls can show us how customers decide to leave. They can't show us how many leave for each reason. The CRM pull below fills that gap, and the calls and the data should be read together.

1. Learning goals

  1. When did the leave decision form, and what triggered it? In the first months or later? During implementation, after go-live, or around an automation project?
  2. Does implementation delay cause churn, or just come with it? If delay mattered, was the cause on our side (fixable by Bet B) or on theirs (data readiness, their staffing)?
  3. Is automation a real driver? Are customers actually automating? Did Crate's lack of robot integration block them, or did it come up afterwards as a reason?
  4. Do signers and daily users tell the same story? Executives may cite strategy (automation). Operations managers may cite experience (implementation, usability).
  5. Are the reasons something else entirely? For example: price, support, acquisition, consolidation, or a non-robotics feature gap.

2. Changes to the plan

1. Pull CRM data in week 0, before the first call. For every account that churned in the last 12 months, pull: - tenure at churn - days from signing to go-live - whether they run or were deploying automation (ask CS and account managers) - where they went

Also answer one question: did the rise from 11% to 18% churn come from first-year customers? If it did, that supports Bet B. If the increase is among long-tenured accounts, that points toward Bet A. This may be the most decisive single fact we get, and it costs a day.

2. Replace at-risk accounts that are in renewal talks. Five of the six are negotiating discounts. They have a reason to exaggerate complaints, and a call with them could easily drift into pricing. Use red-health accounts that are not in active renewal. If there aren't enough of those, add churned accounts instead.

3. Fix the call count. The plan says 12 calls, but two people on each of 12 accounts is 24 calls. Within the 12, I propose 8 accounts:

GroupAccountsWho we interviewCalls
Churned, paired4Signer and ops manager, separately8
Churned, signer only2Signer2
At-risk, not in renewal2Signer2
  • Split the churned accounts so that 3 left in their first year and 3 left later.
  • Include at least one account that moved to a competitor with robotics integrations.
  • Signers get priority because they made the decision.
  • The paired accounts let us compare what the signer says with what the daily user saw.
  • If we can fund 24 calls, pair every account.

4. Ana should not interview her own contacts. Anyone will soften or reshape their story when the CEO asks, especially a CEO they know and whose view they may already have heard. Her contacts also aren't a random sample of the people who left.

My suggestion: - Her contacts can enter the pool if they fit the sampling groups, but a neutral interviewer runs those calls. - Ana and Marcus both listen to the recordings. - If Ana wants to speak with former customers personally, those calls are valuable for the relationship. We run them in addition to the 12 and don't count them as research.

The same rule applies to Marcus: neither exec interviews.

5. Offer a thank-you gift to former customers, such as a gift card. It must not be a credit or discount on Crate.

3. Discussion guide: VP Ops / COO (signer), 45 minutes

0–3 min · Opening - Ask permission to record. If they say no, don't record and take notes only. - Say: "This isn't a sales call. I can't discuss pricing or our product plans. I'm here to understand your experience, good and bad."

3–8 min · Context - "Tell me about your business and how it's changed over the last two years." - "What were your biggest operational priorities this past year?"

8–22 min · The decision story (the core of the call) - Churned: "Take me back to the first moment you started to wonder whether Crate was right for you. What was going on?" - At-risk: "Tell me about the last time you seriously questioned renewing." - Build a timeline with follow-ups: - "When was that, relative to signing?" - "What happened next?" - "Who else was involved?" - "What alternatives did you look at?" - "When was the decision final?" - "What finally tipped it?"

22–30 min · Expectations and early months - "When you signed, what did you expect Crate to do for you, and by when?" - "How did going live compare to what you expected?" - "When, if ever, did you first see the value you signed up for?"

30–37 min · Operational direction - "What investments in the warehouse have you made or planned: equipment, systems, staffing?" - "What did you need your WMS to do as part of those?" - If they mention automation: "How did Crate fit, or not? What did you do about it?"

37–42 min · Prompted check - Hand over a card of possible reasons: implementation time, robotics/automation integration, other missing features, ease of use, support, price, business change. - Rotate the order on every call. The survey showed the first option gets picked more. - Ask: "Pick the top three for you." Then: "Which one, if it had been different, would have changed the outcome?"

42–45 min · Close - "What should I have asked?" - "Who else should we talk to?" - Thank them.

4. Discussion guide: Warehouse operations manager, 45 minutes

0–3 min · Opening - Same as the signer call: recording permission and ground rules.

3–10 min · Daily work - "Walk me through a typical day using Crate." - "What other tools or spreadsheets did you use alongside it?"

10–22 min · Implementation story - "Tell me what happened from signing to go-live." - "What took longest?" - "What were you waiting on, and who was it waiting on?" - "When did it start working the way you needed?" - Get specific episodes, not a verdict. Always ask whether a delay came from us, from them, or from a third party.

22–32 min · Gaps and workarounds - "Tell me about the last time Crate got in your way." - "What equipment or automation runs on your floor? How does information move between it and Crate?" - "What manual steps did that create?"

32–37 min · The decision - "Were you part of the decision to leave or renew? What would you have told leadership?"

37–42 min · Prompted check - Same card as the signer call, with a new random order.

42–45 min · Close - Same as the signer call.

5. Guidance for interviewers

Stay neutral - Don't say "robotics," "automation," or "implementation" before the prompted check, unless the interviewee says it first. What people raise unprompted is our strongest evidence. - Ask about specific past events ("the last time…", "walk me through…"), not opinions or hypotheticals. Treat "we'd have stayed if you'd had X" as weak evidence. - Use their words, not ours. Don't correct them or defend Crate. - In paired accounts, use different interviewers where possible. Never tell the second person what the first one said.

Stay within the rules from Legal and CS - Pricing or discounts: "That's outside what I can cover today. I'll let your account team know you'd like to discuss it." Then tell CS. - Roadmap: "I can't speak to product plans, but I'll make sure what you've said reaches the people who make them." Never promise a feature. - If there's no recording consent, don't record. The note-taker writes verbatim quotes where possible.

Process - A note-taker joins every call. - Within 24 hours, fill in the debrief template: - primary reason, in their words - the decision timeline, with the month the doubt started - robotics raised unprompted? (Y/N) - implementation raised unprompted? (Y/N) - was the delay caused by us, them, or both? - top three from the card - the best quote - Hold a 30-minute team review of the debriefs every Friday.

6. What would back each bet

Each account gets classified by its primary cause, meaning the thing that started or tipped the decision. When signer and ops manager disagree, record both views; don't average them.

Back Bet A (robotics) if: - In at least 3 of the 6 churned accounts, automation was raised unprompted and came before the leave decision. - The customer had a live or funded automation project that Crate blocked or made harder. - Ops managers describe concrete workarounds between robots and Crate. - The CRM pull shows churn concentrated in longer-tenured accounts, with meaningfully more automating churners than the 3 we already know about. - Warning sign: robotics only comes up at the card, or appears in the timeline after the decision was already made. That suggests it's a justification, not a cause.

Back Bet B (implementation) if: - In at least 3 of the 6 churned accounts, doubt formed during or soon after a slow implementation, often phrased as "never saw the value" or "lost confidence." - Ops managers put the main delays on our side: our process, staffing, data migration, or configuration. Those are what a rebuild would fix. - The CRM pull shows the rise in churn comes mostly from first-year customers. - Warning sign: delays were mostly on the customer's side (their data wasn't ready, they didn't have staff). A faster process on our end would then help less than the 2.4× figure suggests.

Back neither if: - Primary causes are spread across several reasons with no clear lead, such as price, support, consolidation or acquisition, or a feature gap unrelated to robotics. - Or signers and ops managers consistently tell different stories that neither bet addresses.

In that case, I'd recommend a smaller fix aimed at whichever cause leads, plus a better-designed exit survey, rather than committing two squads on a weak signal.

If both show up: check whether automating customers are also the ones with the slowest implementations. If they are, those customers' problem is really integration work during setup, and the recommendation should address that directly instead of forcing a choice between A and B.

In week 4, I'll bring the exec team a one-page summary that combines the call findings with the CRM figures, a recommendation, and our confidence in it.

Claims the judge couldn’t find in the brief

  • The survey showed the first option gets picked more.

Grades and run details

Decision model 88 · LLM judge 8 of 14 checks
Decision model checks
  • partialUses the supplied evidence correctly8%
  • passAddresses the actual decision95%
  • partialRespects explicit constraints22%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims72%
  • passProduces the required deliverable83%
  • passTests both theories fairly99%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts99%
  • passDesigned to falsify73%
  • passAsks what people did, not what they think90%
  • passDoesn't lead the witness48%
  • partialMarks what to cut if the call runs over45%
Run
Run
#1
Time to output
1.7 min
Submitted
30 Sept 2026

Needs more than a quick edit9 of 13 checks passed

A PM could use most of it after light edits, but the unsupported data-source claim and missing must-ask/cut guidance require re-checking and revision before running the calls.

Output

Crate churn research: decision guide

Decision: In four weeks, allocate two squads next half to robotics integrations (A), implementation redesign (B), or conclude that neither retention thesis is sufficiently supported.

Research window: Three weeks; twelve 45-minute interviews. These calls will test explanations and reconstruct decisions—not estimate how common each churn cause is across 340 customers.

1. Learning goals

  1. Reconstruct the departure sequence: What happened between signing, implementation, first value, operational problems, searching for alternatives, and deciding to leave?
  2. Identify the decisive mechanism: Was robotics incompatibility or implementation friction necessary to the departure, merely contributory, or a justification offered afterward?
  3. Find actionable scope: Which integrations or implementation changes could plausibly have prevented the problem? Could two squads address them within a half?
  4. Look for competing explanations and counterexamples: Other missing capabilities, service failures, customer-side constraints, or business changes may explain both delays and churn.
  5. Separate operational pain from buying decisions: Did daily users and economic buyers experience the same problem and agree on why the relationship failed?

Starting evidence is suggestive, not conclusive. The survey’s broad, first-listed integration option does not establish robotics demand. The 2.4× churn association does not establish that delays caused churn. Competitor destinations do not establish purchase motives; nine lost prospects are acquisition evidence, not retention evidence.

2. Changes to the call plan

Resolve the interview arithmetic. Twelve accounts with two separate interviews each would require 24 calls. Within the limit, recruit six accounts, interviewing both roles separately: twelve calls total.

Use four churned accounts and two at-risk accounts. This prioritizes actual decisions and reduces exposure to active renewal negotiations while retaining prospective evidence.

Select accounts purposively from the eligible pool, not through executive relationships:

  • Include late and on-time go-lives, plus a failed/never-live implementation if available.
  • Include known robotics signals and accounts without them.
  • Include first-year and longer-tenure departures.
  • Seek counterexamples: late implementation without abandonment; robotics discussion without an actual deployment or switch.

Do not attempt a fully balanced matrix with six accounts. Document selection, refusals, substitutions, and gaps. Check whether the nine-month churn window excludes relevant cases from the annual churn rise; expand to twelve months if needed.

For at-risk accounts, prioritize the one outside renewal talks. CS must confirm that participation—or refusal—will not affect commercial treatment. If a renewal makes independent research impractical, substitute another eligible account.

Ana should not lead interviews with former customers she knows. Use a neutral researcher or PM who owns neither bet. Ana can help recruit through a standard invitation and review consented recordings afterward. Avoid executive attendance that could suppress criticism.

Before calls, prepare a factual account timeline from CRM, implementation logs, support records, and available usage data. Keep interpretations separate. Examine cohort definitions and potential confounders behind the 2.4× figure.

3. Interview guides

Both guides total 45 minutes. Ask open questions first; introduce the two hypotheses only after the participant’s unaided account.

Guide A: Warehouse operations manager

TimeQuestions
0–4 minFirst words, before recording: “May we record this conversation for internal research? Saying no is completely fine; we can take notes instead.” Wait for permission before starting recording. Explain that this is research, not a sales or renewal conversation; participation is optional. “What did you personally own, and during which period?”
4–10 min“What was happening in the warehouse when you chose Crate? What job did you expect it to improve? How would you have recognized success?”
10–20 min“Walk me through signing to the first real production use.” Probe planned versus actual milestones, dependencies, workarounds, ownership, and first useful outcome. “Tell me about a specific day when progress stalled. What happened next?” If never live, trace the last completed milestone and stopping point.
20–30 minChurned: “When did you first think Crate might not work for you? Describe the incident. What happened between that and leaving?” At-risk: “How is Crate working today? Describe the most recent serious problem. Has anyone discussed changing systems? What has actually happened so far?” For both: “What did this cost operationally—time, throughput, errors, or customer commitments? Who saw it?”
30–39 min“What changes in equipment or workflow occurred during this period?” Then probe robotics neutrally: “Were robots evaluated or deployed? Which systems, for what workflow, and when? What specifically could Crate not do? What workaround did you try?” Separately: “Once implementation ended—or stalled—what problems remained?” Do not assume either issue existed.
39–45 min“Which issue mattered most, and what makes you say that? If only that issue had been resolved then, what would still have made Crate unsuitable?” Ask for optional, redacted supporting artifacts and names of decision participants. Summarize the timeline and invite corrections: “What important explanation have I missed?”

Guide B: VP Operations or COO

TimeQuestions
0–4 minUse the same recording-permission opening and research boundaries. “What was your role in selection, implementation oversight, and the decision to stay or leave?”
4–10 min“What business outcome justified choosing Crate? What deadline or event made that outcome important? What expectations were set?”
10–22 minChurned: “Walk me from the first concern to the decision to leave. When was that decision effectively made, rather than formally communicated? Who influenced it? What alternatives did you evaluate?” At-risk: “How are you evaluating whether Crate is working? Has a change been proposed or authorized? What evidence and next steps are involved?” Separate firsthand knowledge from reports by others.
22–31 min“What specific event most changed your confidence? What did you do afterward? What attempts were made to recover the relationship?” Churned: “What did you choose instead, and what requirement made it preferable? Is it operating successfully yet?” At-risk: “What requirements would any alternative have to meet?”
31–39 minTest both explanations, varying their order between interviews. “What role, if any, did implementation timing play? When did it affect your judgment?” “What role, if any, did warehouse automation play? Was there an approved project, named equipment, deployment date, or demonstrated integration requirement?” “Which problems would have remained if each had been solved?”
39–45 min“What would Crate have needed to change, by what date, to alter your decision?” Treat this counterfactual as weaker evidence than actions already taken. “What evidence might contradict your explanation?” Request optional artifacts, recap, and invite corrections.

4. Interviewer operating rules

  • Be curious, not persuasive. Do not quote Ana, Marcus, survey percentages, or the churn correlation.
  • Probe concrete events: “When?” “Who?” “What happened next?” “How do you know?” Distinguish recollection, inference, and observed action.
  • Do not discuss pricing, discounts, or roadmap commitments. If raised: “I can’t discuss commercial terms or future plans here; your account team can handle those separately.” Do not promise escalation outcomes.
  • Interview roles separately; do not reveal the other person’s answers. Disagreement is evidence to investigate, not average away.
  • Use a neutral note-taker. After each call, record chronology, spontaneous versus prompted reasons, supporting evidence, contradictions, and confidence.

5. Evidence required for the decision

Back A if multiple independent accounts—including actual departures—show committed robotics adoption, a specific blocking Crate incompatibility before the departure decision, failed workarounds, and credible evidence that resolving it would have preserved the relationship. Shared, feasible integration needs must emerge. Generic “missing features,” aspirational automation, or competitor capabilities alone do not qualify.

Back B if multiple accounts—including actual departures—show preventable implementation failures preceding loss of confidence, concrete operational consequences, and a credible path from faster successful activation to retention. Identify repeatable Crate-owned bottlenecks. Delays driven mainly by customer readiness, complexity, or already-failing relationships weaken this case.

Back neither if accounts remain contradictory, both theories depend on hypothetical claims, another mechanism dominates, or neither intervention fits the capacity and time available. If both qualify, compare affected retention exposure, intervention feasibility, and remaining uncertainty—not interview vote totals.

Run two pilot calls in week one, finish most interviews in week two, and reserve week three for remaining calls and synthesis. Produce one evidence sheet per account; paired interviews count as one account, not two confirmations. The final readout should combine those cases with cohort data, explicitly separate causality from correlation, and state what remains unproven before recommending either allocation.

Claims the judge couldn’t find in the brief

  • Support records and usage data exist and can be used to prepare account timelines.

Grades and run details

Decision model 96 · LLM judge 9 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly40%
  • passAddresses the actual decision92%
  • passRespects explicit constraints31%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims94%
  • passProduces the required deliverable86%
  • passTests both theories fairly100%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts99%
  • passDesigned to falsify78%
  • passAsks what people did, not what they think74%
  • passDoesn't lead the witness61%
  • partialMarks what to cut if the call runs over52%
Run
Run
#1
Time to output
60 s
Submitted
1 Oct 2026

Needs more than a quick edit7 of 13 checks passed

A PM would need to shorten it, reconcile the account plan with the 12-call constraint, and remove or relabel unsupported portfolio claims before using it.

Output

Research Guide: Root Cause Analysis for H2 Squad Allocation

Prepared by: Staff Product Manager Target Audience: Executive Team (CEO, CPO, VP CS, Head of Product) Timeline: 3 weeks of interviews; executive decision in Week 4 Scope: 12 calls (45 minutes each)

---

1. Executive Summary & Recommended Changes to the Call Plan

Before executing this sprint, we must adjust three structural flaws in the proposed plan to protect data integrity:

  1. Resolve the Account-to-Call Ratio: 12 calls cannot cover two separate stakeholders across 12 accounts. We will target 6 accounts total (3 churned, 3 at-risk) and conduct 2 distinct calls per account: one with the Warehouse Operations Manager (day-to-day user) and one with the VP of Operations/COO (economic buyer). Evaluating both perspectives within the same operational context is essential to determine whether operational failure or strategic misalignment triggered the exit.
  2. Quarantine Active Renewals: 5 of the proposed 6 at-risk accounts are actively negotiating renewals and receiving discounts. Interviewing these buyers creates an immediate incentive for them to exaggerate product flaws to gain commercial leverage. We will swap 3 of these with at-risk accounts that are not currently negotiating pricing.
  3. CEO Interview Protocol: Executive presence introduces severe confirmation and deference bias—especially with personal contacts who already know Ana's passion for robotics. Recommendation: Ana should join her 4 target calls as an executive sponsor for the first 3 minutes, then hand off lead facilitation to a Product Manager, remaining on mute as an observer. If Ana leads, she must strictly follow the non-leading script provided.

---

2. Core Learning Goals

  1. Pinpoint the Churn Decision Window: Did the customer mentally churn during the onboarding/implementation phase (Bet B), or did they leave after reaching steady-state due to technological ceilings (Bet A)?
  2. Assess Real Robotics Demand vs. Narrative: Are customers actively deploying autonomous mobile robots (AMRs) and pick-assist hardware, or is "missing robotics" a convenient, forward-looking justification for leaving an underperforming platform?
  3. Quantify the Cost of Onboarding Drag: Does exceeding the 45-day SLA directly burn operational credibility and destroy ROI, or is delay merely a symptom of customer-side disorganization?
  4. Identify False Dichotomies (Bet Neither): Determine whether churn is driven by factors neither bet solves—such as baseline software unreliability, missing core 3PL billing/EDI features, or macro 3PL volume contraction.

---

3. Interviewer Guidance & Rules of Engagement

  • Recording Consent (Mandatory): State verbatim: "Before we begin, do you mind if I record this session purely for internal note-taking? None of this will be shared externally." If declined, proceed with manual notes.
  • The "No Roadmap, No Pricing" Wall: Customers will attempt to trade feedback for commitments. If asked about features or discounting, respond: "I'm on the product research side and have no visibility into commercials or delivery timelines. My sole focus today is understanding how your operations actually run."
  • Past Behavior Over Speculation: Never ask: "Would you use a robotics integration?" (Answer is always yes). Always ask: "What automation equipment do you have physically deployed on the floor today, and how does your team interact with it?"
  • Root-Cause Probing (The "Five Whys"): When a customer says "missing integrations," do not accept the label. Ask: "What specific warehouse task were you trying to execute that stalled? What was the manual workaround?"

---

4. Interview Scripts & Timings (45 Minutes Each)

Track 1: VP of Operations / COO (Economic Buyer)

#### Part 1: Context & Operational Profile (7 mins) * How has your facility footprint and throughput profile shifted over the last 18 months? * When you initially signed with Crate, what core business metric were you held accountable for improving?

#### Part 2: Implementation & Time-to-Value (12 mins) * Walk me back to your onboarding. What was the internal sentiment between signing the contract and processing your first live pallet? * Our contract targets a 45-day go-live; our median across customers is closer to 90. Where did the process stall, and what internal operational cost did that delay create? * At what point did your leadership team feel Crate was fully operational? Did you ever reach that state?

#### Part 3: Strategic Priorities, Automation, & Feature Gaps (16 mins) * Over the past year, what capital investments have you made on your warehouse floor (e.g., conveyor belts, automated guided vehicles, pick-assist carts, manual racking)? * If automation is present: Who manufactures it, what software controls it today, and what specific data must pass between it and your WMS? * If automation is absent: What is on your signed capital expenditure plan for the next 12 months? * When you evaluated the market (or competitors), what specific capability made it clear Crate was no longer the right long-term partner? * Think back to the moment you realized this contract wouldn't renew. When did that conversation happen, and who initiated it?

#### Part 4: Decision Trace & Wrap-up (10 mins) * If Crate had delivered on every promise made during the sales cycle within 45 days, would you still be with us today? Why or why not? * What is the single biggest operational bottleneck your business faces this quarter?

---

Track 2: Warehouse Operations Manager (Daily User)

#### Part 1: Floor Reality & Daily Routine (8 mins) * Walk me through the first two hours of your shift on the floor. Where do supervisors and pickers spend the majority of their time inside Crate? * What daily tasks force your team to leave Crate and use spreadsheets, paper, or secondary systems?

#### Part 2: Go-Live Experience & Transition (12 mins) * How disruptive was the transition to Crate for your floor staff? How long did it take for pick/pack rates to recover to baseline? * During the first three months, how often did you have to escalate systemic issues to leadership? What were those issues? * Did you feel the system was configured correctly for your physical warehouse layout, or did you have to mold your workflows to fit the software?

#### Part 3: Floor Automation & Daily Bottlenecks (15 mins) * How do orders physically move through your facility today? * If pickers are assisted by any mechanical or automated equipment, where does software failure slow them down? * Have you piloted or evaluated autonomous carts or robotics on your floor? * If yes: What was the hardest operational part of that pilot? * If no: What prevents you from automating (e.g., client SKU variability, facility constraints, budget)? * If you could fix one workflow in Crate that causes your floor workers the most frustration, what would it be?

#### Part 4: The Tipping Point (10 mins) * When did your team start discussing alternative systems or workarounds? Was there a specific breaking point or system failure? * If you had a magic wand to improve your team’s pick/pack speed by 20%, what would you change about your physical setup or your software?

---

5. Decision Rubric: What We Need to Hear

At the end of Week 3, interview findings will be mapped against this falsification framework:

OutcomeEvidence Required from CallsMetric Signal from Portfolio
:---:---:---
Back Bet A<br>(Robotics Integrations)• $\ge 4$ of the 6 accounts have active capital expenditures or signed contracts with AMR/robotics vendors (e.g., Locus, 6 River).<br>• Buyers explicitly state: "We loved Crate's daily usability and onboarding was fine, but our inability to ingest AMR pick-batches forced us to migrate."<br>• Daily users identify manual cart-pushing and picker transit time as their primary throughput ceiling.CRM shows churn clustered heavily in high-volume, automated accounts; robotics surfaced in competitive losses without prompting.
Back Bet B<br>(Implementation Rebuild)• Churn decisions were made inside the first 6 months, directly traceable to operational paralysis during go-live.<br>• Buyers report: "We burned through our operational budget running parallel systems for 100+ days; our team lost trust before we ever saw value."<br>• Robotics is dismissed as an aspirational 3-year vision, not a current operational requirement.First-year logo churn represents the majority of our 18% churn rate; churn correlates directly with onboarding days exceeding 45.
Back Neither<br>(Pivot to Bet C)• Automation is negligible among these mid-sized 3PLs; go-live delays are considered standard industry friction.<br>• Churn is driven by baseline platform stability (e.g., database timeouts, inventory sync errors), missing table-stakes 3PL multi-client billing, or standard EDI connectors to Shopify/Amazon.<br>• Customers cite price-to-value disconnects: "We left because Crate charges premium rates while basic core features remain broken."Exit survey "Missing features" was a proxy for basic warehouse workflows, not robotics; churn spread evenly across all customer tenure bands.

---

6. Next Steps

  1. Day 1-2: Approve adjusted account roster (3 non-renewing at-risk, 3 churned). CS to send intro emails.
  2. Day 3-17: Conduct 12 recorded interviews. Maintain a shared executive synthesis doc updated within 24 hours of each call.
  3. Day 21: Final Decision Meeting. Present thematic evidence and allocation recommendation to Ana and Marcus.

Claims the judge couldn’t find in the brief

  • Customers who went live after 90 days churned at 2.4× the rate of the rest in their first year.
  • CRM shows churn clustered heavily in high-volume, automated accounts.
  • Robotics surfaced in competitive losses without prompting.
  • First-year logo churn represents the majority of the 18% churn rate.
  • Churn correlates directly with onboarding days exceeding 45.
  • Exit survey 'Missing features' was a proxy for basic warehouse workflows, not robotics.
  • Churn is spread evenly across all customer tenure bands.

Grades and run details

Decision model 77 · LLM judge 7 of 14 checks
Decision model checks
  • partialUses the supplied evidence correctly10%
  • passAddresses the actual decision91%
  • failRespects explicit constraints34%
  • passIdentifies material uncertainty97%
  • partialAvoids unsupported claims32%
  • passProduces the required deliverable48%
  • passTests both theories fairly82%
  • passDifferent questions for user and signer100%
  • passProtects the calls and the accounts99%
  • passDesigned to falsify77%
  • passAsks what people did, not what they think44%
  • partialDoesn't lead the witness38%
  • partialMarks what to cut if the call runs over50%
Run
Run
#1
API response time
37 s
Submitted
1 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyWrongMixedWrong
Opus 5.5 · Claude

It invents or overstates a few current-situation facts, including 'roughly 60 churned accounts' and 'the survey showed the first option gets picked more'.

GPT-6 Astra · ChatGPT

It invents current data sources by saying to prepare timelines from support records and usage data, which are not supplied.

Gemini 3.8 Flash · API

The guide relies on several unsupported current-situation or portfolio claims, including CRM clustering, robotics surfacing without prompting, first-year churn majority, >45-day correlation, and exit-survey meaning.

Addresses the actual decisionRightMixedMixed
Opus 5.5 · Claude

It gives the exec team a clear decision rule for Bet A, Bet B, or neither, with thresholds and warning signs.

GPT-6 Astra · ChatGPT

It does not commit to one answer early; it only states the decision frame and later gives evidence thresholds, leaving the exec team without a clear recommendation.

Gemini 3.8 Flash · API

It does not commit to Bet A, Bet B, or neither; it defers the decision to the research, even though the supplied evidence is not framed as insufficient to make a call.

Respects explicit constraintsWrongMixedWrong
Opus 5.5 · Claude

It exceeds the 1,500-word limit and proposes replacing the at-risk cohort in a way that changes the supplied plan beyond the brief's constraints.

GPT-6 Astra · ChatGPT

It respects the no-pricing/roadmap and consent constraints, but it does not mark must-ask questions or say what to cut if time runs short, and it introduces unsupported data sources.

Gemini 3.8 Flash · API

It exceeds the requested 1,500-word limit and its proposed 6-account roster would not cover the 12 accounts specified in the supplied call plan.

Avoids unsupported claimsMixedRightWrong
Opus 5.5 · Claude

It presents the first-option bias and the 'roughly 60 churned accounts' figure as established facts when the supplied evidence does not support them.

GPT-6 Astra · ChatGPT

It labels exec theories and survey/correlation evidence as suggestive rather than established, though it still has a minor unsupported data-source claim.

Gemini 3.8 Flash · API

It presents several interpretations and portfolio signals as established facts or evidence, such as exit-survey meaning, churn distribution, and robotics mentions without prompting.

Produces the required deliverableMixedRightMixed
Opus 5.5 · Claude

The guide is complete and usable, but it is over the requested length.

GPT-6 Astra · ChatGPT

It provides learning goals, timed two-role questions, interviewer guidance, plan changes, and decision signals within the length, and the exec team could act on it with light edits.

Gemini 3.8 Flash · API

Although it is a usable interview guide, it is too long and changes the account scope in a way that does not satisfy the supplied 12-call plan.

All got wrong 1

Marks what to cut if the call runs overWrongWrongWrong
Opus 5.5 · Claude

It provides timings but does not mark must-ask questions or say what to cut if the call runs over.

GPT-6 Astra · ChatGPT

Timings add to 45 minutes, but the guide does not mark must-ask questions or specify what to cut if time runs short.

Gemini 3.8 Flash · API

The timings add up to 45 minutes, but the number of questions is too dense for a realistic 45-minute call and there are no must-ask markers or cut-down instructions.

All got right 7

Identifies material uncertaintyRightRightRight
Opus 5.5 · Claude

It names the key unknowns—timing of the leave decision, whether delay causes churn, and whether automation is a real driver—and says how calls and CRM pulls would resolve them.

GPT-6 Astra · ChatGPT

It names key unknowns such as survey bias, correlation versus causation, timing of churn decisions, confounders, and renewal exposure, and says how calls and data checks would resolve them.

Gemini 3.8 Flash · API

It names material unknowns such as timing of churn decisions, real robotics deployment, onboarding drag versus customer-side disorganization, and third causes, and proposes calls to resolve them.

Tests both theories fairlyRightRightRight
Opus 5.5 · Claude

It gives both robotics and implementation equal chances to be disproved and leaves room for a third cause.

GPT-6 Astra · ChatGPT

It asks open timeline questions before introducing either theory and gives both robotics and implementation questions that could disprove each, while leaving room for third causes.

Gemini 3.8 Flash · API

It gives both robotics and implementation disconfirming questions and explicitly leaves room for neither, including stability, billing/EDI, and macro causes.

Different questions for user and signerRightRightRight
Opus 5.5 · Claude

It provides distinct 45-minute tracks for the signer and the daily operations manager.

GPT-6 Astra · ChatGPT

It has separate timed tracks for the daily warehouse operations manager and the VP/COO signer, matched to operational experience versus decision process.

Gemini 3.8 Flash · API

It provides separate tracks for the daily warehouse operations manager and the VP/COO signer, with role-appropriate questions.

Protects the calls and the accountsRightRightRight
Opus 5.5 · Claude

It includes recording consent, scripts for pricing/roadmap, and a concrete plan to keep Ana from leading calls with her contacts.

GPT-6 Astra · ChatGPT

It starts with recording consent, bans pricing/discounts/roadmap, gives a script if raised, and concretely removes Ana from leading calls with people she knows.

Gemini 3.8 Flash · API

It includes recording consent, a no-pricing/no-roadmap rule with a script, and a concrete protocol for Ana's calls.

Designed to falsifyRightRightRight
Opus 5.5 · Claude

Each learning goal has questions that could falsify the corresponding hypothesis, such as timing, cause of delay, and unprompted automation mentions.

GPT-6 Astra · ChatGPT

Each learning goal has disconfirming probes, including counterexamples, timing of decisions, failed workarounds, and questions about what would still have caused churn.

Gemini 3.8 Flash · API

Each learning goal has questions that could falsify the relevant hypothesis, such as actual automation deployment, go-live timeline, and alternative causes.

Asks what people did, not what they thinkRightRightRight
Opus 5.5 · Claude

The core questions ask for specific past events, timelines, and actions rather than opinions or hypotheticals.

GPT-6 Astra · ChatGPT

The core questions ask for specific events, timelines, actions, alternatives evaluated, and what happened next, with hypotheticals clearly secondary.

Gemini 3.8 Flash · API

The core questions ask for recent concrete behavior, including deployed equipment, go-live experience, decision timing, and workarounds.

Doesn't lead the witnessRightRightRight
Opus 5.5 · Claude

It explicitly avoids naming robotics or implementation before the prompted check and uses neutral, open questions.

GPT-6 Astra · ChatGPT

Questions are open and neutral, hypotheses are introduced only after unaided accounts, and the guide does not pitch either bet or ask customers to choose.

Gemini 3.8 Flash · API

The questions are mostly neutral and avoid asking customers to choose between the bets or pitching the ideas.

Results

Every setup we’ve tested on this task, across all cases and repeats, graded on the current checklist. Calibrated: the graders match our PM on 98% of checks.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.192.62None
2Sonnet 5.5withAPI90.284.62None
3GPT-6 AstrawithChatGPT96.078.32None
4Opus 5.5withClaude92.163.22None
5Gemini 3.8 FlashwithAPI80.167.32None
6GPT-6 LunawithAPI92.073.921 capped
7Gemini 3.5 Flash-LitewithGemini53.826.122 capped

About the task

The PM job

Preparing for customer calls.

Why it matters

Leading questions produce the answers you wanted. A good guide is designed to prove you wrong.

What good looks like

  • Past-behaviour questions
  • No leading prompts
  • Clear learning goals
  • Could disprove the team's hypothesis

Deliberately not measured

    Capability tested

    Question design

    The failure we’re looking for

    Hypothetical, leading questions

    Grading

    Decision model and LLM judge, calibrated against a blind PM review