Usable with a quick edit11 of 12 checks passed
The output is complete and usable; the only missing piece is a brief mention of the recruiter's withdrawal data, which is a quick edit.
Output
Recommended loop
Replace the unpaid strategy deck, six-person presentation and “beer test” with structured evidence of doing this job. The previous hires’ notes cannot explain why one succeeded and one failed; cross-functional influence must be a hiring requirement, not a personality impression.
Rounds
| Round | Who / duration | What it tests and how |
|---|---|---|
| 1. Structured screen | Recruiter, 20 minutes | Confirm scope, motivation, compensation and logistics. Ask for one conversion or expansion initiative: personal contribution, measured result and partner teams involved. Use consistent questions; accept relevant experience without requiring a Growth PM title. Advance at most two finalists under the interviewer-time budget. |
| 2. Ownership deep dive | Amara, 35 minutes | Examine one growth initiative and one unsuccessful attempt to change another team’s behavior. Probe baseline, decision-making, individual contribution, results, uncertainty and lessons. Test senior ownership, growth judgment and influence through specific actions—not polished storytelling. |
| 3. Live growth work sample | Head of Engineering + data analyst, 40 minutes | Give everyone the same short, synthetic packet: a funnel beginning at Copperline’s 9% trial conversion, cohort/segment data, expansion opportunities, customer observations and delivery constraints. Allow eight minutes of quiet reading within the session. Ask candidates to diagnose, prioritize conversion versus expansion opportunities, and scope the first weekly experiment with three engineers and a designer. Introduce one standardized complication. Test analysis, prioritization, experimentation and delivery. No deck or advance work. |
| 4. Cross-functional working session | Head of Sales + senior designer, 35 minutes | Candidate proposes a trial-conversion change. Sales resists because it changes incentives or workload; Design flags customer friction. Use identical opening objections, while responding naturally thereafter. Candidate must uncover interests, negotiate an achievable pilot and establish responsibilities and follow-through. Finish with reflection on what they would change. Test influence, customer judgment and collaboration—not whether they “win.” |
| 5. Offer conversation—not another assessment | CEO, 20 minutes, selected candidate only | Explain company direction, answer questions and help close. No informal culture-fit veto. |
Allow candidate questions in each assessment round. Share the format, competencies and preparation expectations beforehand; offer scheduling flexibility, breaks and reasonable accommodations. Assess reasoning and decisions, not presentation polish. The work sample is fictional, not free consulting.
Scorecard
Use 1 = clear contrary evidence, 2 = below bar, 3 = meets senior-role bar, 4 = exceptional evidence. “Not observed” is separate from a low score and must be resolved before deciding.
| Competency / primary evidence owner | Strong: meets or exceeds bar | Weak: below bar |
|---|---|---|
| Growth diagnosis and measurement / analyst | Defines conversion and expansion denominators, examines cohorts and segments, distinguishes correlation from causation, identifies missing evidence. Treats 9% as a starting point, not a diagnosis. | Jumps to tactics; misreads funnels; claims causality from movement; ignores retention or expansion economics. |
| Experiment design and prioritization / analyst, Amara | Chooses a defensible hypothesis, outcome and guardrails; weighs impact, effort and confidence; handles limited traffic and inconclusive results. | Produces a tactic list; promises unsupported lifts; assumes every weekly release can deliver statistically conclusive results. |
| Weekly delivery with a small team / Engineering | Slices a useful first release, identifies instrumentation and dependencies, trades scope deliberately, and plans the next decision. | Requires a major rebuild or more staff; ignores technical uncertainty; equates speed with skipping measurement or quality. |
| Influence without authority / Sales | Investigates incentives and constraints, makes a credible mutual-value case, secures specific commitments and establishes follow-through. Past examples show personal action and honest limits. | Relies on title, escalation or charisma; blames Sales; mistakes agreement in a meeting for changed behavior. |
| Customer and commercial judgment / Design | Connects accountancy-firm workflows to conversion and expansion; protects trust and usability; considers downstream retention and revenue quality. | Optimizes conversion through friction, pressure or misleading choices; ignores customer costs or churn. |
| Senior ownership and learning / Amara | Separates own contribution from team results, owns misses, changes course with evidence and creates a repeatable operating rhythm. | Offers vague achievements, cannot explain decisions, takes all credit or externalizes failures. |
Interviewers record concrete actions, quotes and reasoning alongside scores. “Smart,” “great presence” and “likeable” are not evidence.
Decision, capacity and timeline
Week 1: recruiter screens and schedules two finalists. Week 2: complete assessments. Week 3: decide, close and issue the offer.
Reserve each assessor 10 minutes for calibration, five minutes per finalist for scoring, and 20 minutes for debrief. Including interviews, Engineering and the analyst each use exactly 120 minutes; Amara, Sales and Design each use 110. CEO uses at most 40 minutes. If neither finalist qualifies, reopen rather than lower the bar or silently exceed the budget.
Submit scores independently before seeing others’ notes. In debrief, review evidence competency by competency; Amara speaks last. No show-of-hands vote.
Require at least 3 on every competency, especially influence and weekly delivery. Resolve disagreements through recorded evidence, not averaging or seniority. Missing critical evidence means no offer yet. Amara owns the final decision and documents the rationale, remaining risks and onboarding support.
Grades and run details
Decision model 92 · LLM judge 12 of 13 checks
Decision model checks
- passUses the supplied evidence correctly54%
- passAddresses the actual decision93%
- passRespects explicit constraints23%
- partialIdentifies material uncertainty51%
- partialAvoids unsupported claims29%
- passProduces the required deliverable91%
- passTests what the last hire failed at100%
- passFixes the take-home's cost to candidates80%
- passFits the people and the time75%
- passJudges on evidence, not presence100%
- passDefines good for this role first95%
- passKeeps each judgement independent100%
Run
- Run
- #1
- Time to output
- 62 s
- Submitted
- 5 Oct 2026