Needs more than a quick edit9 of 13 checks passed
A PM could use most of it after light edits, but the unsupported data-source claim and missing must-ask/cut guidance require re-checking and revision before running the calls.
Output
Crate churn research: decision guide
Decision: In four weeks, allocate two squads next half to robotics integrations (A), implementation redesign (B), or conclude that neither retention thesis is sufficiently supported.
Research window: Three weeks; twelve 45-minute interviews. These calls will test explanations and reconstruct decisions—not estimate how common each churn cause is across 340 customers.
1. Learning goals
- Reconstruct the departure sequence: What happened between signing, implementation, first value, operational problems, searching for alternatives, and deciding to leave?
- Identify the decisive mechanism: Was robotics incompatibility or implementation friction necessary to the departure, merely contributory, or a justification offered afterward?
- Find actionable scope: Which integrations or implementation changes could plausibly have prevented the problem? Could two squads address them within a half?
- Look for competing explanations and counterexamples: Other missing capabilities, service failures, customer-side constraints, or business changes may explain both delays and churn.
- Separate operational pain from buying decisions: Did daily users and economic buyers experience the same problem and agree on why the relationship failed?
Starting evidence is suggestive, not conclusive. The survey’s broad, first-listed integration option does not establish robotics demand. The 2.4× churn association does not establish that delays caused churn. Competitor destinations do not establish purchase motives; nine lost prospects are acquisition evidence, not retention evidence.
2. Changes to the call plan
Resolve the interview arithmetic. Twelve accounts with two separate interviews each would require 24 calls. Within the limit, recruit six accounts, interviewing both roles separately: twelve calls total.
Use four churned accounts and two at-risk accounts. This prioritizes actual decisions and reduces exposure to active renewal negotiations while retaining prospective evidence.
Select accounts purposively from the eligible pool, not through executive relationships:
- Include late and on-time go-lives, plus a failed/never-live implementation if available.
- Include known robotics signals and accounts without them.
- Include first-year and longer-tenure departures.
- Seek counterexamples: late implementation without abandonment; robotics discussion without an actual deployment or switch.
Do not attempt a fully balanced matrix with six accounts. Document selection, refusals, substitutions, and gaps. Check whether the nine-month churn window excludes relevant cases from the annual churn rise; expand to twelve months if needed.
For at-risk accounts, prioritize the one outside renewal talks. CS must confirm that participation—or refusal—will not affect commercial treatment. If a renewal makes independent research impractical, substitute another eligible account.
Ana should not lead interviews with former customers she knows. Use a neutral researcher or PM who owns neither bet. Ana can help recruit through a standard invitation and review consented recordings afterward. Avoid executive attendance that could suppress criticism.
Before calls, prepare a factual account timeline from CRM, implementation logs, support records, and available usage data. Keep interpretations separate. Examine cohort definitions and potential confounders behind the 2.4× figure.
3. Interview guides
Both guides total 45 minutes. Ask open questions first; introduce the two hypotheses only after the participant’s unaided account.
Guide A: Warehouse operations manager
| Time | Questions |
|---|---|
| 0–4 min | First words, before recording: “May we record this conversation for internal research? Saying no is completely fine; we can take notes instead.” Wait for permission before starting recording. Explain that this is research, not a sales or renewal conversation; participation is optional. “What did you personally own, and during which period?” |
| 4–10 min | “What was happening in the warehouse when you chose Crate? What job did you expect it to improve? How would you have recognized success?” |
| 10–20 min | “Walk me through signing to the first real production use.” Probe planned versus actual milestones, dependencies, workarounds, ownership, and first useful outcome. “Tell me about a specific day when progress stalled. What happened next?” If never live, trace the last completed milestone and stopping point. |
| 20–30 min | Churned: “When did you first think Crate might not work for you? Describe the incident. What happened between that and leaving?” At-risk: “How is Crate working today? Describe the most recent serious problem. Has anyone discussed changing systems? What has actually happened so far?” For both: “What did this cost operationally—time, throughput, errors, or customer commitments? Who saw it?” |
| 30–39 min | “What changes in equipment or workflow occurred during this period?” Then probe robotics neutrally: “Were robots evaluated or deployed? Which systems, for what workflow, and when? What specifically could Crate not do? What workaround did you try?” Separately: “Once implementation ended—or stalled—what problems remained?” Do not assume either issue existed. |
| 39–45 min | “Which issue mattered most, and what makes you say that? If only that issue had been resolved then, what would still have made Crate unsuitable?” Ask for optional, redacted supporting artifacts and names of decision participants. Summarize the timeline and invite corrections: “What important explanation have I missed?” |
Guide B: VP Operations or COO
| Time | Questions |
|---|---|
| 0–4 min | Use the same recording-permission opening and research boundaries. “What was your role in selection, implementation oversight, and the decision to stay or leave?” |
| 4–10 min | “What business outcome justified choosing Crate? What deadline or event made that outcome important? What expectations were set?” |
| 10–22 min | Churned: “Walk me from the first concern to the decision to leave. When was that decision effectively made, rather than formally communicated? Who influenced it? What alternatives did you evaluate?” At-risk: “How are you evaluating whether Crate is working? Has a change been proposed or authorized? What evidence and next steps are involved?” Separate firsthand knowledge from reports by others. |
| 22–31 min | “What specific event most changed your confidence? What did you do afterward? What attempts were made to recover the relationship?” Churned: “What did you choose instead, and what requirement made it preferable? Is it operating successfully yet?” At-risk: “What requirements would any alternative have to meet?” |
| 31–39 min | Test both explanations, varying their order between interviews. “What role, if any, did implementation timing play? When did it affect your judgment?” “What role, if any, did warehouse automation play? Was there an approved project, named equipment, deployment date, or demonstrated integration requirement?” “Which problems would have remained if each had been solved?” |
| 39–45 min | “What would Crate have needed to change, by what date, to alter your decision?” Treat this counterfactual as weaker evidence than actions already taken. “What evidence might contradict your explanation?” Request optional artifacts, recap, and invite corrections. |
4. Interviewer operating rules
- Be curious, not persuasive. Do not quote Ana, Marcus, survey percentages, or the churn correlation.
- Probe concrete events: “When?” “Who?” “What happened next?” “How do you know?” Distinguish recollection, inference, and observed action.
- Do not discuss pricing, discounts, or roadmap commitments. If raised: “I can’t discuss commercial terms or future plans here; your account team can handle those separately.” Do not promise escalation outcomes.
- Interview roles separately; do not reveal the other person’s answers. Disagreement is evidence to investigate, not average away.
- Use a neutral note-taker. After each call, record chronology, spontaneous versus prompted reasons, supporting evidence, contradictions, and confidence.
5. Evidence required for the decision
Back A if multiple independent accounts—including actual departures—show committed robotics adoption, a specific blocking Crate incompatibility before the departure decision, failed workarounds, and credible evidence that resolving it would have preserved the relationship. Shared, feasible integration needs must emerge. Generic “missing features,” aspirational automation, or competitor capabilities alone do not qualify.
Back B if multiple accounts—including actual departures—show preventable implementation failures preceding loss of confidence, concrete operational consequences, and a credible path from faster successful activation to retention. Identify repeatable Crate-owned bottlenecks. Delays driven mainly by customer readiness, complexity, or already-failing relationships weaken this case.
Back neither if accounts remain contradictory, both theories depend on hypothetical claims, another mechanism dominates, or neither intervention fits the capacity and time available. If both qualify, compare affected retention exposure, intervention feasibility, and remaining uncertainty—not interview vote totals.
Run two pilot calls in week one, finish most interviews in week two, and reserve week three for remaining calls and synthesis. Produce one evidence sheet per account; paired interviews count as one account, not two confirmations. The final readout should combine those cases with cohort data, explicitly separate causality from correlation, and state what remains unproven before recommending either allocation.
Claims the judge couldn’t find in the brief
- Support records and usage data exist and can be used to prepare account timelines.
Grades and run details
Decision model 96 · LLM judge 9 of 14 checks
Decision model checks
- passUses the supplied evidence correctly40%
- passAddresses the actual decision92%
- passRespects explicit constraints31%
- passIdentifies material uncertainty99%
- passAvoids unsupported claims94%
- passProduces the required deliverable86%
- passTests both theories fairly100%
- passDifferent questions for user and signer100%
- passProtects the calls and the accounts99%
- passDesigned to falsify78%
- passAsks what people did, not what they think74%
- passDoesn't lead the witness61%
- partialMarks what to cut if the call runs over52%
Run
- Run
- #1
- Time to output
- 60 s
- Submitted
- 1 Oct 2026