Usable with a quick edit11 of 13 checks passed
The guide is thorough and evidence-based; a PM could add a note on which questions to prioritize if time is tight and use it immediately.
Output
Crate churn research: executive pre-read and interview guide
Decision: Allocate two squads next half to robotics integrations (A), implementation rebuild (B), or neither if the evidence does not justify either.
Research stance: Test both explanations, not confirm either executive’s account. These interviews can reveal mechanisms and decision timelines; they cannot estimate their prevalence across 340 customers or prove causality.
1. Learning goals
- Reconstruct the departure process: When did concerns emerge, when did leaving become likely, and when was the decision effectively made?
- Identify consequential problems: Separate reasons volunteered after leaving from problems that materially changed warehouse outcomes and the buying decision.
- Test Bet A: Was a specific robotics incompatibility a binding obstacle? Would addressing it have plausibly retained the account?
- Test Bet B: Did implementation delay prevent value, damage trust, or trigger departure? What caused the delay, and could Crate control it?
- Find alternatives and interactions: Price, service, reliability, business changes, or other missing capabilities may dominate. Robotics work may itself complicate implementation.
- Connect learning to an investable scope: What could two squads deliver next half that would address the demonstrated retention mechanism?
What the existing evidence does—and does not—say
- The exit survey’s broad integration category, single-choice format, first-position placement, and sparse free text do not establish robotics as the cause.
- The 2.4× churn association makes implementation worth investigating, but complexity, customer readiness, or other factors could drive both delay and churn.
- Three competitor moves and nine lost deals indicate a robotics opportunity; acquisition losses are not evidence of a retention mechanism.
- Unknown decision dates make chronology essential.
2. Changes to the call plan
Resolve the arithmetic: Twelve separate interviews cannot cover twelve accounts with two roles each. Use six accounts × two separate role interviews = twelve 45-minute calls.
Recommended mix: Four churned accounts and two at-risk accounts. This reduces reliance on renewal negotiations while retaining a prospective view. Findings will be depth-oriented, not representative.
Select accounts from the full CRM population, not executive contacts or survey answers alone:
- Across churned accounts, include slow and faster implementations, first-year and later departures, and robotics-exposed and non-exposed customers.
- Include at least one confirmed move to a robotics-enabled competitor and one slow implementation without an apparent robotics need.
- For at-risk accounts, prefer one with active automation plans and one with implementation/value-realization problems. Include the non-renewal account if suitable.
- Seek contrasting retained accounts through existing operational data, even though the call budget does not cover them.
Recruit both the daily warehouse operations manager and the contract-signing VP Operations/COO. If the original signer has left, recruit someone directly involved in the departure or renewal decision and document that substitution.
Ana should not lead calls with her personal contacts. Use a neutral interviewer without renewal responsibility. Ana can help recruit; observing requires explicit customer consent. Her existing conversations are useful leads, not independently verified findings.
CS should introduce research as separate from renewal discussions, then step out. Participation must not affect commercial treatment.
3. Interview guides: 45 minutes each
Use the relevant role column. In every section, distinguish direct experience from hearsay. Ask broad questions before naming either bet.
| Time | Daily warehouse operations manager | VP Operations / COO |
|---|---|---|
| 0–3 min: permission and framing | First ask: “May we record this conversation for internal research?” If declined, continue with notes. Then explain purpose and confidentiality limits. | Same opening. |
| 3–7 min: context and expected value | “What did your warehouse handle, and what was your role with Crate? What was supposed to improve?” | “What led you to buy Crate? What outcomes and deadlines mattered? How would you judge success?” |
| 7–17 min: chronological reconstruction | “Walk me from signing through onboarding, first operational use, and the point when problems became serious.” “Describe a particular shift or incident.” Establish dates, workarounds, operational impact, and who knew. | “Walk me from purchase to the first concern, evaluating alternatives, and the decision.” “When did staying stop being the default?” Establish dates, decision-makers, triggers, and alternatives. |
| 17–27 min: implementation and value | “What had to happen before you could use Crate successfully? Where did work stall? Who owned each step?” “When did you first achieve useful results?” Probe migration, configuration, training, integrations, staffing, and rework after an open answer. | “What go-live date did you expect, and what happened? What consequences did that have?” “What other factors accompanied the delay?” “If the same product had gone live in 45 days, what would still have put the relationship at risk?” Ask why. |
| 27–36 min: workflow, automation, and unmet needs | “Which workflows did Crate support poorly? Show or describe a recent example.” Then: “Were robots in use or planned? Which systems, what workflow, and what connection was required?” “What happened without it?” | “What capabilities influenced staying or switching?” Then probe automation: vendor, deployment date, committed budget, required integration, and alternatives considered. “If that connection had existed, what else would have needed to change for you to stay?” |
| 36–42 min: decision and competing explanations | “Which problem mattered most in daily operations? What did you escalate, to whom, and when?” “What worked well?” “What have we missed?” | “Which issues were necessary to the decision, and which were secondary?” “What would have had to be different for you to stay?” Ask about business changes, service, reliability, and other alternatives without forcing a category. |
| 42–45 min: verify and close | Summarize the timeline and mechanism: “What have I misunderstood?” Request relevant artifacts and permission for a brief clarification follow-up. | Same; verify whether operational problems actually influenced the commercial decision. |
Status-specific wording
- Churned: Ask what the replacement actually delivered, whether it went live, and whether the cited problem improved—not merely what the competitor promised.
- At-risk: Ask what has already happened, current unresolved consequences, and what would cause escalation or departure. Do not imply a departure decision exists. Distinguish funded plans from aspirations.
4. Guidance for interviewers
- Begin with recording permission; record only after consent. Explain that this is research, not a support, roadmap, or renewal discussion.
- Do not discuss prices, discounts, or roadmap commitments. If raised, acknowledge without probing commercial terms: “I can’t discuss that here; your account team handles commercial conversations.” Record volunteered price concerns as an alternative explanation.
- Avoid “Did missing robotics make you leave?” and “Would faster onboarding have saved you?” Start with events; use counterfactuals only afterward.
- Ask for concrete examples, dates, artifacts, and consequences. “It was painful” is not yet a mechanism.
- Do not defend Crate, disclose executive hypotheses, or offer fixes.
- Interview roles separately before reconciling accounts. Treat disagreement as evidence, not an error to erase.
- Use a note-taker and common template: timeline, problem, consequence, decision link, alternative explanation, evidence source, and confidence. Count accounts, not two interviewees as independent votes.
5. Evidence required for the decision
Back A: robotics integrations
Look for a repeated chain across independent accounts:
Committed robotics deployment → specific missing connection → material operational/business constraint → documented role in departure or credible near-term risk.
Stronger evidence includes decision-maker confirmation, pre-decision correspondence, and a replacement that solved the problem. Match the implicated vendors/workflows to current customer exposure. Before funding, establish that a feasible integration scope could address meaningful retention risk next half.
Mentions of “automation,” hypothetical plans, or competitor marketing are insufficient.
Back B: implementation rebuild
Look for:
Crate-controllable implementation bottleneck → delayed usable value or broken operational commitments → loss of confidence/avoidable cost → departure decision before value realization.
Seek repeated bottlenecks, role-level corroboration, and contrasts with faster implementations. Validate the mechanism against the full cohort, accounting where possible for complexity and customer readiness. Identify changes two squads could deliver; do not assume a wholesale rebuild is necessary.
Back neither—or declare the result inconclusive
Neither is justified if dominant causes lie elsewhere, robotics demand is narrow or speculative, delays are largely outside Crate’s control, or either intervention arrives too late to affect retention. If interviews remain contradictory or unsupported, report insufficient evidence, not “neither causes churn.”
Compare both bets on reachable ARR at risk, evidence strength, controllability, delivery feasibility, and time to impact—not mention counts.
6. Three-week execution
- Week 1: Audit cohorts and timelines, recruit, pilot one role pair.
- Week 2: Complete remaining interviews; synthesize after each pair without changing standards to favor an emerging winner.
- Week 3: Verify discrepancies, triangulate with records, assess engineering scope, and prepare the decision memo with supporting evidence, counterexamples, uncertainties, and a recommendation.
Do not commit two squads merely because one story sounds more persuasive.
Check by check
Got wrong · 1
- Fits the callThe guide includes timed sections that add up to 45 minutes but does not say what to cut if time runs short, nor does it mark must-ask questions.
Mixed · 1
- Addresses the actual decisionThe output recommends more research before deciding, states the supplied evidence cannot yet support a call, and specifies exactly what evidence would back each bet or neither.The two graders disagreed on this one.
Got right · 11
- Uses the supplied evidence correctlyAll factual statements about the current situation are taken directly from the supplied context without invention.
- Respects explicit constraintsThe guide is under 1,500 words, includes all requested sections (learning goals, timed questions for two roles, guidance, plan changes, decision signals), and is addressed to the exec team.
- Identifies material uncertaintyIt names key unknowns (decision dates, whether robotics was a binding obstacle, whether delays were Crate-controllable) and says how the interview evidence would resolve them.
- Avoids unsupported claimsInterpretations are clearly labelled as such (e.g., the exit survey does not establish robotics as the cause), and no cause is presented as established fact.
- Produces the required deliverableThe output is a complete guide with learning goals, timed questions for both roles, interviewer guidance, plan changes, and decision signals; an exec team could act on it.
- Tests both theories fairlyQuestions could disprove robotics (e.g., 'Were robots in use? What happened without it?') and implementation (e.g., 'Where did work stall? If go-live had been 45 days, what would still have put the relationship at risk?'), and open questions leave room for a third cause.
- Different questions for user and signerThe table provides distinct tracks: daily user gets questions about specific shifts, workarounds, and escalations; the signer gets questions about the purchase decision, alternatives, and go-live expectations.
- Protects the calls and the accountsIt forbids discussing pricing/discounts/roadmap, provides a script to deflect, starts with recording consent, and has a concrete plan for Ana (she does not lead, observes with consent, her prior conversations are treated as leads).
- Designed to falsifyEach learning goal has disconfirming questions (e.g., 'What happened without the robot connection?' for Bet A, 'If the same product had gone live in 45 days, what would still have put the relationship at risk?' for Bet B).
- Asks what people did, not what they thinkCore questions ask for a chronological walk-through, specific incidents, dates, and what they did, not opinions or predictions.
- Doesn't lead the witnessQuestions are open and neutral; robotics and implementation are not named until after broad problem questions, and no answer is hinted at.
Grades and run details
Decision model 88 · LLM judge 13 of 14 checks
Decision model checks
- passUses the supplied evidence correctly62%
- failAddresses the actual decision55%
- passRespects explicit constraints20%
- passIdentifies material uncertainty99%
- passAvoids unsupported claims95%
- passProduces the required deliverable86%
- passTests both theories fairly100%
- passDifferent questions for user and signer100%
- passProtects the calls and the accounts95%
- passDesigned to falsify64%
- passAsks what people did, not what they think71%
- passDoesn't lead the witness58%
- partialFits the call68%
Run
- Run
- #1
- API response time
- 80 s
- Submitted
- 30 Sept 2026