Usable with a quick edit11 of 12 checks passed
The output is a complete, actionable plan that a PM could implement with minor wording tweaks; no evidence re-checking or reasoning redo is needed.
Output
Growth PM interview loop
Assumptions: The two-hour cap is each person's total across all candidates, so I've designed for three finalists. Early rounds are cheap and don't use the panel's time. Only Amara's time is spent before the finalist stage.
Why the change
- The 5-day, 10-page deck cost us 14 of 40 candidates, mostly people with caregiving duties or demanding jobs. It also tests deck-making, not the job.
- The last hire failed on influencing Sales, and nothing in the old loop tested that.
- "Great presence" and "very smart" in the notes are not evidence, and a show-of-hands vote rewards the loudest voice.
The loop
| # | Round | Who | Tests | Time per candidate |
|---|---|---|---|---|
| 1 | Screen + written question | Recruiter, scored by Amara | Basics, comp, and a 300-word written answer: "A conversion experiment you ran: baseline, result, what you'd change." | 30 min call, 20 min writing |
| 2 | Ownership deep-dive | Amara | Track record on revenue outcomes | 20 min |
| 3 | Funnel case | Data analyst | Conversion judgment | 30 min |
| 4 | Ship-it scoping | Head of Eng + senior designer | Weekly shipping with a small team | 30 min |
| 5 | Influence role-play | Head of Sales | Changing others' behavior without authority | 30 min |
| 6 | Business and working style | CEO | Judgment, first-hire fit | 20 min |
| 7 | Debrief | All interviewers | Decision | 20 min total |
Rounds 3–6 can be done in one sitting, or split across two days on video. Total candidate time is about 3 hours.
Round details - Round 1: Amara spends 5 minutes on each of about six written answers (30 min total) and picks three finalists. - Round 3: We send a one-page anonymized trial funnel 24 hours ahead, with a note to spend no more than 30 minutes on it. In the room the candidate says where they'd look first, picks two experiments, and designs one, including metric, guardrail, and what a realistic sample size allows at our volume. - Round 4: We give a half-formed experiment idea. The candidate must cut it to something shippable within a week by three engineers and a designer. The interviewers push back on scope. - Round 5: The Head of Sales plays a skeptical rep or manager who thinks trial follow-up is a distraction. The candidate has to get agreement on a small pilot. The last 5 minutes cover a real past example. - Round 6: A structured conversation covering how the candidate would think about expansion revenue for small accountancy firms, what being the first growth hire means, and a time they disagreed and then committed. This replaces "would I grab a beer."
Scorecard
Each criterion is rated 1–4, with no midpoint. Every interviewer writes specific observed behaviors, not impressions. "Presence" and "smart" don't count as evidence.
1. Influence without authority (weight 30%; Sales, plus Amara and references) - Strong: Starts from what Sales cares about (quota, time, commissions). Proposes a small, reversible pilot. Offers something in return. Has a past example where a team they didn't manage changed behavior, with details. - Weak: Says "align stakeholders." Escalates to the boss. Presents the change as a mandate. Blames other teams in past stories.
2. Conversion judgment (25%; Analyst) - Strong: Segments before theorizing. Notes that our volume limits what tests can detect. Ties each experiment to a hypothesis and a guardrail. Asks about accountancy seasonality and who actually decides to pay. - Weak: Lists generic tactics (checklists, gamification). Has no view on sample size. Doesn't ask about the data.
3. Shipping cadence (20%; Eng + Design) - Strong: Cuts scope readily to the smallest test. Gives real numbers on past experiments per week or month. Treats engineers and designers as partners and takes their objections seriously. - Weak: Proposes large projects with distant payoffs. Writes specs without hypotheses. Defends scope instead of trading it off.
4. Ownership of revenue outcomes (15%; Amara) - Strong: Gives baselines, deltas, and their specific contribution. Talks candidly about experiments that failed. - Weak: Vague credit for team results. Describes outputs (launches, decks) instead of outcomes.
5. Judgment and fit for a 60-person company (10%; CEO) - Strong: Is energized by ambiguity. Builds their own tools and doesn't wait for resources. Disagrees directly and then commits. - Weak: Needs a playbook or a team. Can't prioritize without data. Becomes defensive when challenged.
Note for the Head of Sales: Score what the candidate did in the role-play, not whether you enjoyed being challenged.
Decision process
- Independent scoring. Each interviewer submits ratings and evidence before the debrief and can't see others' scores.
- Debrief (20 min). Amara reads out the scores, lowest first, on each criterion. Disagreements get discussed against the evidence, not general impressions.
- Must-haves. A score of 2 or below on influence or shipping means no offer. The last miss was an influence failure, and weak shipping defeats the weekly-experiment job.
- Ranking. Among candidates who pass, the highest weighted score wins. Close calls go to Amara.
- Decision owner. Amara decides, and there is no vote. The CEO can block only by citing specific observed evidence tied to the scorecard.
- References. The recruiter checks references for the top candidate, using Amara's script. At least one reference must be a cross-functional partner such as someone from sales or marketing, asking, "Did they change how your team worked? How?"
Timeline and time budget
- Week 1: Recruiter screens and written questions. Amara picks finalists by the end of the week.
- Week 2: Amara's deep-dives, then rounds 3–6 for all three finalists.
- Week 3: Debrief on Monday, references on Tuesday and Wednesday, offer by Friday.
| Person | Time |
|---|---|
| Amara | 30 min (written answers) + 60 min (3 deep-dives) + 20 min (debrief) = 110 min |
| Analyst, Sales, Eng, Design | 90 min (3 interviews) + 20 min (debrief) = 110 min each |
| CEO | 60 min + 20 min = 80 min |
If more than three candidates reach round 2, Amara's time is the first limit. In that case, tighten the written screen before asking for more of her time.
Grades and run details
Decision model 92 · LLM judge 12 of 13 checks
Decision model checks
- passUses the supplied evidence correctly30%
- passAddresses the actual decision78%
- passRespects explicit constraints28%
- partialIdentifies material uncertainty22%
- partialAvoids unsupported claims47%
- passProduces the required deliverable90%
- passTests what the last hire failed at100%
- passFixes the take-home's cost to candidates100%
- passFits the people and the time79%
- passJudges on evidence, not presence100%
- passDefines good for this role first92%
- passKeeps each judgement independent99%
Run
- Run
- #1
- API response time
- 53 s
- Submitted
- 5 Oct 2026