Usable with a quick edit9 of 12 checks passed
The output is clear, actionable, and covers all required parts; a quick edit to add conditions or label interpretations would make it fully ready.
Output
Amara—here is the redesigned interview loop. Your draft had a fatal flaw: the 10-page take-home is a diversity filter (driving away parents and senior operators) and tests academic strategy over execution. Furthermore, our last hiring failure showed we hired for "presence" rather than the ability to influence Sales.
Here is a 3-week, high-signal loop designed to test gritty experimentation, cross-functional persuasion, and data fluency.
---
Part 1: The Interview Loop
Round 1: Recruiter Screen * Who: Recruiter (30 mins) * Tests: Baseline communication, compensation alignment, current trial-to-paid context, and general interest in B2B SaaS.
Round 2: The Operating Screen (Live Case) * Who: Amara, VP Product (45 mins, video) * Tests: Problem decomposition, metric intuition, and ability to move fast. * Format: No slides. We present our 9% trial-to-paid conversion rate and a specific funnel drop-off. We spend 30 minutes live-whiteboarding how they would diagnose and attack it within their first 30 days.
Round 3: The Cross-Functional Influence & Craft Deep-Dive * Who: Head of Engineering & Senior Designer (45 mins, video) * Tests: Can this person ship weekly? Do they respect engineering constraints? Do they partner well with design on user psychology? * Format: Deep-dive into a past project where they shipped rapid iterations. How do they handle tech debt vs. speed tradeoffs?
Round 4: The Sales & Data Crucible * Who: Head of Sales & Data Analyst (45 mins, video) * Tests: Crucial. Can they influence Sales without authority? Can they audit data to find buried conversion leaks? * Format: A roleplay. The Head of Sales plays a skeptical account executive who thinks the PM's new trial-onboarding flow is stealing their pipeline. The candidate must handle the pushback, find common ground, and align incentives.
Round 5: Founder Alignment & Values * Who: CEO (30 mins, in-person or video) * Tests: Velocity mindset, ownership mentality, and cultural add (dropping the "beer test" bias, which screens for homogeneity). * Format: Discussion on Copperline’s mission, how they handle failure, and what autonomy looks like here.
---
Part 2: The Scorecard
We will grade candidates on four core competencies using a 1–4 scale (1 = Unacceptable, 4 = Exceptional).
#### 1. Rigorous Experimentation & Velocity * Strong (3–4): Relentlessly focused on the funnel. Talks easily about cadence, MVP scoping, and shipping weekly. Bases ideas on user behavior, not just "best practices." * Weak (1–2): Prefers massive, multi-month redesign projects. Hand-waves about metrics or treats data as an afterthought.
#### 2. Cross-Functional Influence (Specifically with Sales) * Strong (3–4): Listens to Sales objections, understands quota pressure, and designs incentives that make Sales want to use the new product motions. Uses data to persuade. * Weak (1–2): Views Sales as "resistant" or "the enemy." Relies on positional authority ("Amara said we're doing this") rather than building trust.
#### 3. Data Fluency * Strong (3–4): Can spot a vanity metric instantly. Identifies cohort retention issues and funnel bottlenecks unprompted. Comfortable querying or partnering deeply with analysts. * Weak (1–2): Confuses activity metrics (clicks) with outcome metrics (expansion revenue). Gets lost when looking at raw conversion funnels.
#### 4. Product Craft & Empathy (Accountants) * Strong (3–4): Quickly grasps the nuanced workflow of a small accountancy firm. Designs intuitive onboarding that respects the user's time. * Weak (1–2): Generic "growth hacking" mentality that applies consumer tactics (spammy pop-ups, dark patterns) to professional B2B software.
---
Part 3: Decision-Making Process
To avoid the vague "they seemed smart" trap, we will use a structured debrief.
- Independent Scoring: Within 24 hours of their final round, every interviewer must submit their independent scores and notes into a shared doc before reading anyone else's.
- The Debrief Meeting (30 mins): Led by Amara. We do not "vote by show of hands." Instead, each interviewer speaks to their specific competency (e.g., Sales speaks to Round 4, Eng speaks to Round 3).
- The Bar: A hire requires a minimum score of 3 across all four competencies, with no score of 1 or 2 in Cross-Functional Influence. If Sales or Engineering flags a collaboration red flag, it's a hard veto.
By cutting the take-home, we keep top talent in the funnel. By replacing the "beer test" with a Sales roleplay, we ensure our next PM doesn't repeat the exact failure mode of our last one.
Claims the judge couldn’t find in the brief
- The beer test screens for homogeneity.
Grades and run details
Decision model 79 · LLM judge 10 of 13 checks
Decision model checks
- passUses the supplied evidence correctly39%
- passAddresses the actual decision81%
- partialRespects explicit constraints38%
- failIdentifies material uncertainty70%
- failAvoids unsupported claims33%
- passProduces the required deliverable53%
- passTests what the last hire failed at100%
- passFixes the take-home's cost to candidates100%
- passFits the people and the time24%
- passJudges on evidence, not presence67%
- passDefines good for this role first84%
- passKeeps each judgement independent99%
Run
- Run
- #1
- Time to output
- 6 s
- Submitted
- 5 Oct 2026