Usable with a quick edit9 of 12 checks passed
The output is complete and actionable; a quick edit to mention the recruiter's withdrawal data would make it fully aligned with the brief.
Output
Interview loop
Run this over 2–3 weeks. Skip the five-day take-home, large presentation panel, and “beer” fit interview. Share the same expectations and questions with every candidate; offer scheduling flexibility and accommodations.
1. Recruiter screen — 30 minutes Tests role understanding, motivation, and practical alignment: the scope, weekly shipping pace, and need to influence Sales and Marketing without authority. Not a “polish” screen.
2. Amara, VP Product — 60 minutes Tests growth ownership and product judgment. Ask for a detailed example of a funnel or revenue problem the candidate owned: diagnosis, choices, experiments, results, and what they learned. Probe how they balance conversion, customer trust, and expansion.
3. Live growth work sample — 60 minutes Send a one-page brief and synthetic funnel data 24 hours ahead; no deck or advance work required. Spend 30 minutes with the data analyst diagnosing the funnel and proposing a measurable experiment; spend 30 minutes with the Head of Engineering and senior designer scoping how to ship and learn from it. Use the same prompt and time limits for all candidates.
4. Head of Sales — 45 minutes Structured roleplay: Sales is reluctant to change its trial follow-up because reps are busy and doubt the data. Ask the candidate to understand the objection, make a case, and agree a practical next step. Then probe a real example of influencing a team they did not manage.
The CEO does not need a separate round: “culture fit” is not a hiring criterion, and the role’s key cross-functional test is better assessed directly with Sales, Engineering, and Design.
Scorecard
Use a 1–4 rating for each dimension: 1 = clear concern, 2 = below bar, 3 = meets bar, 4 = unusually strong. Record evidence, not impressions.
| Dimension | Strong evidence | Weak evidence |
|---|---|---|
| Growth ownership and outcomes (Amara) | Personally owned a meaningful funnel or revenue outcome; explains decisions, results, attribution limits, and lessons. | Describes activity rather than impact; overclaims credit or cannot explain what changed. |
| Customer and product judgment (Amara) | Understands small-firm workflows and trust needs; weighs conversion against customer value and expansion. | Offers generic SaaS tactics or pushes short-term conversion at the expense of trust. |
| Data diagnosis and experiment design (work sample) | Finds plausible bottlenecks and segments; proposes a focused test with a clear hypothesis, success metric, guardrail, and learning plan. | Jumps to a solution, misreads the data, uses vanity metrics, or cannot say what result would change their mind. |
| Shipping and learning cadence (work sample) | Scopes a feasible experiment, identifies dependencies and tradeoffs, and shows how the team can learn and iterate weekly. | Proposes a big-bang project, ignores constraints, or treats learning as something to do after launch. |
| Influence without authority (Sales) | Listens for incentives and concerns, uses evidence appropriately, adapts the ask, and secures a credible next step. | Blames Sales, relies on escalation or authority, or cannot turn disagreement into action. |
| Cross-functional partnership (work sample) | Makes clear tradeoffs with Engineering and Design, invites their expertise, and aligns on ownership and delivery. | Treats partners as order-takers or leaves scope and responsibilities vague. |
Decision
Interviewers submit ratings and specific evidence independently before a 30-minute debrief; do not vote by show of hands. Amara facilitates, checks for evidence behind claims, and makes the hiring decision. Do not use “presence,” “smart,” or general “fit” as evidence.
A hire must score at least 3 on influence, experiment design, and shipping, with no score below 2 in any dimension. Do not average away a weakness in a core requirement. If evidence is genuinely missing, identify one focused follow-up rather than adding a broad round. Give candidates a timely decision and close the loop within three weeks of the first interview.
Grades and run details
Decision model 96 · LLM judge 10 of 13 checks
Decision model checks
- passUses the supplied evidence correctly51%
- passAddresses the actual decision82%
- passRespects explicit constraints26%
- partialIdentifies material uncertainty64%
- passAvoids unsupported claims22%
- passProduces the required deliverable78%
- passTests what the last hire failed at100%
- passFixes the take-home's cost to candidates68%
- passFits the people and the time64%
- passJudges on evidence, not presence100%
- passDefines good for this role first86%
- passKeeps each judgement independent98%
Run
- Run
- #1
- API response time
- 39 s
- Submitted
- 5 Oct 2026