Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're helping Amara Osei, VP Product at Copperline, hire our first Growth PM (a senior role). She's shared her draft interview loop and asked you to redesign it. Write the loop you'd run (each round: who runs it, what it tests, how long it takes), the scorecard (what strong and weak look like for each thing we're testing), and how we'll make the decision at the end. No more than 900 words. What we know is below.

What the model was given6 items: About Copperline, The role, Amara's draft loop, The last two PM hires, From the recruiter, Who can interview
About CopperlineInvoicing and payments software for small accountancy firms. 60 people. Trial to paid conversion is 9%.
The roleOwns trial conversion and expansion revenue. Works with three engineers and a designer, reports to Amara, and has to ship experiments every week. Needs Sales and Marketing to change how they work, without managing them.
Amara's draft loop1. Recruiter screen. 2. Take-home: 'design a growth strategy for Copperline', a 10-page deck, five days to complete. 3. Presentation to a panel of six. 4. Culture fit with the CEO: 'would I grab a beer with them?' 5. Debrief: everyone discusses, then votes by show of hands. There's no scorecard.
The last two PM hiresOne left after five months: 'couldn't get Sales to change anything'. The other is doing well. Both sets of interview notes say mainly 'great presence' and 'very smart'.
From the recruiterOf the last 40 candidates given the take-home, 14 withdrew, saying they didn't have time. Nine of the 14 had young children or demanding current roles.
Who can interviewAmara, the Head of Engineering, a senior designer, the Head of Sales, a data analyst and the CEO. Each can give at most two hours to this hire. Amara wants an offer out within three weeks of the first interview.
What a strong answer doesThe answer key the graders mark against

A loop built on a scorecard written first: for a senior growth PM, experimentation and analytics, execution, and influence without authority (weighted up, since the last hire failed on it), with strong and weak signals for each. Rounds test the real job: a live diagnosis of Copperline's own trial funnel instead of a five-day generic deck (which loses candidates, mostly parents and people in demanding roles), a behavioural round on past experiments and metrics actually moved, and an influence round run by the Head of Sales. Replaces the 'grab a beer' culture fit with a structured values or collaboration conversation. Each interviewer writes their evaluation before the debrief, and the decision is made against the scorecard, not a show of hands. Fits each interviewer's two hours and the three-week target, and adds reference checks.

Critical failures (cap the score at 40)
  • Keeps a culture-fit test based on whether the interviewer would like a beer with the candidate
Task

Mid-level · v1.0 · synthetic · B2B SaaS, hiring, growth

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 12 checks passed

The output is well-structured and complete; the one unsupported generalization can be removed with a quick edit without affecting the loop design.

Output

Copperline Growth PM: Interview Loop

Why the draft needs changing

  • The take-home is filtering out the wrong people. 14 of 40 candidates withdrew, and most of them were parents or in demanding jobs. A growth PM in a senior role is likely to be exactly that kind of person.
  • The panel of six and the "beer" test reward presence. Presence is what your interview notes captured last time, and it didn't predict success.
  • Nothing tested influence over Sales. That is why the last hire failed.

The new loop tests five things, each owned by one interviewer.

Time budget

Each interviewer has two hours in total for this hire. That limits the loop to three finalists. Per interviewer, the time breaks down as:

  • 3 × 30-minute interviews
  • 3 × 5-minute scorecards
  • one 15-minute debrief

That totals 2 hours, or 1.5 hours for the CEO. The recruiter screens with no time limit, so screening carries the shortlisting load.

The loop

RoundWhoTestsLength
0. ScreenRecruiterHas owned a conversion or expansion metric; has shipped experiments at pace. Structured questions, scored pass/fail.30 min
1. Growth deep-diveAmaraExperiment velocity. Walk through one experiment programme they ran: cadence, hit rate, what they killed.30 min
2. Funnel caseData analystQuantitative judgement. Read an anonymised Copperline trial funnel in the session (10 min) and say where the 9% is leaking, what they'd test first and how they'd know. No prep.30 min
3. Influence role-playHead of SalesInfluence without authority. The data shows Sales' handling of trial accounts hurts conversion. The Head of Sales plays himself and pushes back realistically. Ends with "tell me about a time this went badly."30 min
4. Build sessionHead of Engineering + senior designerWorking with the team. Take one idea from the funnel case and scope it to ship in a week: what gets cut, what gets instrumented, what the designer needs.30 min
5. Commercial judgementCEOCommercial sense. Where expansion revenue comes from in small accountancy firms, pricing levers, and what they'd not do. Structured questions. No "culture fit."20 min

The candidate spends about 2.5 hours in total, ideally in one afternoon (remote is fine). Candidates get the round list and competencies in advance.

Scorecard

Each interviewer scores only their own competency, on a scale of 1 to 4. Every score must cite specific evidence: what the candidate said or did. The words "presence," "smart" and "great energy" don't count as evidence.

CompetencyStrong (4)Weak (1)
Experiment velocity (must-have)Names a real cadence (e.g. weekly), a real hit rate, and experiments they killed fast. Talks about learning per week, not big launches.Describes one large project. Can't say how many tests they ran or what failed. "Strategy" without shipping.
Quantitative judgementSpots the biggest leak quickly. Questions data quality. Sizes the opportunity. Defines a success metric and minimum sample before testing.Jumps to solutions. Treats all drop-offs as equal. No view on how to measure the result.
Influence without authority (must-have)Asks what Sales is paid on and afraid of. Proposes a small reversible pilot with Sales sharing the credit. Stays calm under pushback. Has a real story of winning over a resistant team.Leans on data or seniority ("the numbers say so"). Escalates to Amara or the CEO early. Gives up or bulldozes. Past examples are all teams they managed.
Working with Engineering and DesignCuts scope hard to hit a week. Asks engineers for options rather than dictating. Gives the designer a problem, not a mock-up. Plans instrumentation up front.Specs a month of work. Dictates solutions. Forgets tracking.
Commercial senseUnderstands how small accountancy firms buy and grow. Links trial and expansion to revenue and margin. Names trade-offs and things they'd refuse to do.Generic SaaS playbook. Can't connect activity to revenue. Agrees with everything.

Overall recommendation for each interviewer: Strong hire, Hire, No hire or Strong no hire.

How we decide

1. Score independently. Scorecards are submitted within two hours of the interview and before the debrief. Nobody sees others' scores until theirs is in. 2. Short debrief. One 15-minute debrief covers all three finalists, run by Amara after the final loop. Everyone reads the scorecards beforehand. The meeting is for resolving disagreements, not for retelling the interviews. 3. Speaking order. The most junior interviewer speaks first (data analyst, then designer, Engineering, Sales, CEO). The CEO and Amara speak last, so seniority doesn't anchor the room. 4. No show of hands. The bar is explicit: - A 3 or higher on both must-haves: experiment velocity and influence. - No 1s anywhere. - An average of at least 3. 5. Who decides. Amara makes the final call among candidates who clear the bar. A 1 from the Head of Sales on influence blocks the hire. This is the exact failure mode from last time. 6. If nobody clears the bar, we don't hire. We reopen the search rather than lowering the bar to hit the deadline. 7. References. We take one reference from someone in Sales or Marketing who worked alongside the candidate, not someone who reported to them. We ask specifically how the candidate got them to change something.

Timeline (offer within three weeks)

DaysActivity
1–5Recruiter screens. Shortlist three finalists.
6–10Finalist loops, all in the same week so comparisons stay fresh.
11Debrief and decision.
12–14References.
15Offer out. This leaves about a week of buffer for scheduling slips or a second-choice candidate.

Grades and run details

Decision model 83 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly19%
  • passAddresses the actual decision90%
  • partialRespects explicit constraints17%
  • failIdentifies material uncertainty16%
  • partialAvoids unsupported claims38%
  • passProduces the required deliverable88%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time53%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first93%
  • passKeeps each judgement independent100%
Run
Run
#1
Time to output
48 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 12 checks passed

The output is a complete, actionable plan that a PM could implement with minor wording tweaks; no evidence re-checking or reasoning redo is needed.

Output

Growth PM interview loop

Assumptions: The two-hour cap is each person's total across all candidates, so I've designed for three finalists. Early rounds are cheap and don't use the panel's time. Only Amara's time is spent before the finalist stage.

Why the change

  • The 5-day, 10-page deck cost us 14 of 40 candidates, mostly people with caregiving duties or demanding jobs. It also tests deck-making, not the job.
  • The last hire failed on influencing Sales, and nothing in the old loop tested that.
  • "Great presence" and "very smart" in the notes are not evidence, and a show-of-hands vote rewards the loudest voice.

The loop

#RoundWhoTestsTime per candidate
1Screen + written questionRecruiter, scored by AmaraBasics, comp, and a 300-word written answer: "A conversion experiment you ran: baseline, result, what you'd change."30 min call, 20 min writing
2Ownership deep-diveAmaraTrack record on revenue outcomes20 min
3Funnel caseData analystConversion judgment30 min
4Ship-it scopingHead of Eng + senior designerWeekly shipping with a small team30 min
5Influence role-playHead of SalesChanging others' behavior without authority30 min
6Business and working styleCEOJudgment, first-hire fit20 min
7DebriefAll interviewersDecision20 min total

Rounds 3–6 can be done in one sitting, or split across two days on video. Total candidate time is about 3 hours.

Round details - Round 1: Amara spends 5 minutes on each of about six written answers (30 min total) and picks three finalists. - Round 3: We send a one-page anonymized trial funnel 24 hours ahead, with a note to spend no more than 30 minutes on it. In the room the candidate says where they'd look first, picks two experiments, and designs one, including metric, guardrail, and what a realistic sample size allows at our volume. - Round 4: We give a half-formed experiment idea. The candidate must cut it to something shippable within a week by three engineers and a designer. The interviewers push back on scope. - Round 5: The Head of Sales plays a skeptical rep or manager who thinks trial follow-up is a distraction. The candidate has to get agreement on a small pilot. The last 5 minutes cover a real past example. - Round 6: A structured conversation covering how the candidate would think about expansion revenue for small accountancy firms, what being the first growth hire means, and a time they disagreed and then committed. This replaces "would I grab a beer."

Scorecard

Each criterion is rated 1–4, with no midpoint. Every interviewer writes specific observed behaviors, not impressions. "Presence" and "smart" don't count as evidence.

1. Influence without authority (weight 30%; Sales, plus Amara and references) - Strong: Starts from what Sales cares about (quota, time, commissions). Proposes a small, reversible pilot. Offers something in return. Has a past example where a team they didn't manage changed behavior, with details. - Weak: Says "align stakeholders." Escalates to the boss. Presents the change as a mandate. Blames other teams in past stories.

2. Conversion judgment (25%; Analyst) - Strong: Segments before theorizing. Notes that our volume limits what tests can detect. Ties each experiment to a hypothesis and a guardrail. Asks about accountancy seasonality and who actually decides to pay. - Weak: Lists generic tactics (checklists, gamification). Has no view on sample size. Doesn't ask about the data.

3. Shipping cadence (20%; Eng + Design) - Strong: Cuts scope readily to the smallest test. Gives real numbers on past experiments per week or month. Treats engineers and designers as partners and takes their objections seriously. - Weak: Proposes large projects with distant payoffs. Writes specs without hypotheses. Defends scope instead of trading it off.

4. Ownership of revenue outcomes (15%; Amara) - Strong: Gives baselines, deltas, and their specific contribution. Talks candidly about experiments that failed. - Weak: Vague credit for team results. Describes outputs (launches, decks) instead of outcomes.

5. Judgment and fit for a 60-person company (10%; CEO) - Strong: Is energized by ambiguity. Builds their own tools and doesn't wait for resources. Disagrees directly and then commits. - Weak: Needs a playbook or a team. Can't prioritize without data. Becomes defensive when challenged.

Note for the Head of Sales: Score what the candidate did in the role-play, not whether you enjoyed being challenged.

Decision process

  1. Independent scoring. Each interviewer submits ratings and evidence before the debrief and can't see others' scores.
  2. Debrief (20 min). Amara reads out the scores, lowest first, on each criterion. Disagreements get discussed against the evidence, not general impressions.
  3. Must-haves. A score of 2 or below on influence or shipping means no offer. The last miss was an influence failure, and weak shipping defeats the weekly-experiment job.
  4. Ranking. Among candidates who pass, the highest weighted score wins. Close calls go to Amara.
  5. Decision owner. Amara decides, and there is no vote. The CEO can block only by citing specific observed evidence tied to the scorecard.
  6. References. The recruiter checks references for the top candidate, using Amara's script. At least one reference must be a cross-functional partner such as someone from sales or marketing, asking, "Did they change how your team worked? How?"

Timeline and time budget

  • Week 1: Recruiter screens and written questions. Amara picks finalists by the end of the week.
  • Week 2: Amara's deep-dives, then rounds 3–6 for all three finalists.
  • Week 3: Debrief on Monday, references on Tuesday and Wednesday, offer by Friday.
PersonTime
Amara30 min (written answers) + 60 min (3 deep-dives) + 20 min (debrief) = 110 min
Analyst, Sales, Eng, Design90 min (3 interviews) + 20 min (debrief) = 110 min each
CEO60 min + 20 min = 80 min

If more than three candidates reach round 2, Amara's time is the first limit. In that case, tighten the written screen before asking for more of her time.

Grades and run details

Decision model 92 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly30%
  • passAddresses the actual decision78%
  • passRespects explicit constraints28%
  • partialIdentifies material uncertainty22%
  • partialAvoids unsupported claims47%
  • passProduces the required deliverable90%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time79%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first92%
  • passKeeps each judgement independent99%
Run
Run
#1
API response time
53 s
Submitted
5 Oct 2026

Usable with a quick edit9 of 12 checks passed

The output is clear, actionable, and covers all required parts; a quick edit to add conditions or label interpretations would make it fully ready.

Output

Amara—here is the redesigned interview loop. Your draft had a fatal flaw: the 10-page take-home is a diversity filter (driving away parents and senior operators) and tests academic strategy over execution. Furthermore, our last hiring failure showed we hired for "presence" rather than the ability to influence Sales.

Here is a 3-week, high-signal loop designed to test gritty experimentation, cross-functional persuasion, and data fluency.

---

Part 1: The Interview Loop

Round 1: Recruiter Screen * Who: Recruiter (30 mins) * Tests: Baseline communication, compensation alignment, current trial-to-paid context, and general interest in B2B SaaS.

Round 2: The Operating Screen (Live Case) * Who: Amara, VP Product (45 mins, video) * Tests: Problem decomposition, metric intuition, and ability to move fast. * Format: No slides. We present our 9% trial-to-paid conversion rate and a specific funnel drop-off. We spend 30 minutes live-whiteboarding how they would diagnose and attack it within their first 30 days.

Round 3: The Cross-Functional Influence & Craft Deep-Dive * Who: Head of Engineering & Senior Designer (45 mins, video) * Tests: Can this person ship weekly? Do they respect engineering constraints? Do they partner well with design on user psychology? * Format: Deep-dive into a past project where they shipped rapid iterations. How do they handle tech debt vs. speed tradeoffs?

Round 4: The Sales & Data Crucible * Who: Head of Sales & Data Analyst (45 mins, video) * Tests: Crucial. Can they influence Sales without authority? Can they audit data to find buried conversion leaks? * Format: A roleplay. The Head of Sales plays a skeptical account executive who thinks the PM's new trial-onboarding flow is stealing their pipeline. The candidate must handle the pushback, find common ground, and align incentives.

Round 5: Founder Alignment & Values * Who: CEO (30 mins, in-person or video) * Tests: Velocity mindset, ownership mentality, and cultural add (dropping the "beer test" bias, which screens for homogeneity). * Format: Discussion on Copperline’s mission, how they handle failure, and what autonomy looks like here.

---

Part 2: The Scorecard

We will grade candidates on four core competencies using a 1–4 scale (1 = Unacceptable, 4 = Exceptional).

#### 1. Rigorous Experimentation & Velocity * Strong (3–4): Relentlessly focused on the funnel. Talks easily about cadence, MVP scoping, and shipping weekly. Bases ideas on user behavior, not just "best practices." * Weak (1–2): Prefers massive, multi-month redesign projects. Hand-waves about metrics or treats data as an afterthought.

#### 2. Cross-Functional Influence (Specifically with Sales) * Strong (3–4): Listens to Sales objections, understands quota pressure, and designs incentives that make Sales want to use the new product motions. Uses data to persuade. * Weak (1–2): Views Sales as "resistant" or "the enemy." Relies on positional authority ("Amara said we're doing this") rather than building trust.

#### 3. Data Fluency * Strong (3–4): Can spot a vanity metric instantly. Identifies cohort retention issues and funnel bottlenecks unprompted. Comfortable querying or partnering deeply with analysts. * Weak (1–2): Confuses activity metrics (clicks) with outcome metrics (expansion revenue). Gets lost when looking at raw conversion funnels.

#### 4. Product Craft & Empathy (Accountants) * Strong (3–4): Quickly grasps the nuanced workflow of a small accountancy firm. Designs intuitive onboarding that respects the user's time. * Weak (1–2): Generic "growth hacking" mentality that applies consumer tactics (spammy pop-ups, dark patterns) to professional B2B software.

---

Part 3: Decision-Making Process

To avoid the vague "they seemed smart" trap, we will use a structured debrief.

  1. Independent Scoring: Within 24 hours of their final round, every interviewer must submit their independent scores and notes into a shared doc before reading anyone else's.
  2. The Debrief Meeting (30 mins): Led by Amara. We do not "vote by show of hands." Instead, each interviewer speaks to their specific competency (e.g., Sales speaks to Round 4, Eng speaks to Round 3).
  3. The Bar: A hire requires a minimum score of 3 across all four competencies, with no score of 1 or 2 in Cross-Functional Influence. If Sales or Engineering flags a collaboration red flag, it's a hard veto.

By cutting the take-home, we keep top talent in the funnel. By replacing the "beer test" with a Sales roleplay, we ensure our next PM doesn't repeat the exact failure mode of our last one.

Claims the judge couldn’t find in the brief

  • The beer test screens for homogeneity.

Grades and run details

Decision model 79 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly39%
  • passAddresses the actual decision81%
  • partialRespects explicit constraints38%
  • failIdentifies material uncertainty70%
  • failAvoids unsupported claims33%
  • passProduces the required deliverable53%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time24%
  • passJudges on evidence, not presence67%
  • passDefines good for this role first84%
  • passKeeps each judgement independent99%
Run
Run
#1
Time to output
6 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Addresses the actual decisionRightRightMixed
Opus 5.5 · Claude

The output commits to a specific interview loop, scorecard, and decision process, framed for Amara, with clear conditions for not hiring.

Sonnet 5.5 · API

The output commits unambiguously to a specific interview loop, scorecard, and decision process, as requested.

Gemini 3.5 Flash-Lite · Gemini

The output does not state what result or condition would change the recommended loop.

Identifies material uncertaintyMixedWrongWrong
Opus 5.5 · Claude

It identifies that if no candidate clears the bar, the search reopens, and provides buffer for scheduling slips.

Sonnet 5.5 · API

The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.

Gemini 3.5 Flash-Lite · Gemini

The output does not name unknowns that could change the decision or how they would be resolved.

Avoids unsupported claimsWrongRightWrong
Opus 5.5 · Claude

The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.

Sonnet 5.5 · API

The output does not present interpretations or forecasts as established facts; its claims are supported by the supplied evidence.

Gemini 3.5 Flash-Lite · Gemini

It presents 'diversity filter' and 'screens for homogeneity' as established facts without labelling them as interpretations.

All got right 9

Uses the supplied evidence correctlyRightRightRight
Opus 5.5 · Claude

All factual claims about the current situation are taken directly from the brief or supplied context, with no inventions.

Sonnet 5.5 · API

All statements about the current situation are taken directly from the brief or supplied context, with no invented facts.

Gemini 3.5 Flash-Lite · Gemini

All facts and figures about the current situation are taken correctly from the brief; no invented data.

Respects explicit constraintsRightRightRight
Opus 5.5 · Claude

The output respects the 900-word limit, replaces the take-home and beer test, and fits interviewer time and three-week deadline.

Sonnet 5.5 · API

The output respects the 900-word limit, replaces the take-home and beer test, fits the two-hour per interviewer and three-week timeline, and delivers the required components.

Gemini 3.5 Flash-Lite · Gemini

The loop respects all constraints: length, no beer test, no unchanged take-home, interviewer time limits, and three-week target.

Produces the required deliverableRightRightRight
Opus 5.5 · Claude

The output is a complete interview loop, scorecard, and decision process, within the word limit, usable by Amara.

Sonnet 5.5 · API

The output provides a complete loop, scorecard, and decision process within the word limit, usable by Amara with light edits.

Gemini 3.5 Flash-Lite · Gemini

The output includes the loop, scorecard, and decision process, within 900 words, and is usable by Amara.

Tests what the last hire failed atRightRightRight
Opus 5.5 · Claude

The loop includes an influence role-play with the Head of Sales and probes past examples of changing another team's work.

Sonnet 5.5 · API

Round 5 is an influence role-play with the Head of Sales, and the scorecard weights influence without authority at 30%, directly testing the last hire's failure point.

Gemini 3.5 Flash-Lite · Gemini

Round 4 with the Head of Sales specifically tests influencing Sales without authority, with strong and weak signals.

Fixes the take-home's cost to candidatesRightRightRight
Opus 5.5 · Claude

It replaces the five-day take-home with a live funnel case and cites the withdrawal data and who it drove away.

Sonnet 5.5 · API

The five-day take-home is replaced with a 300-word written answer and a 30-minute funnel case, citing the withdrawal data and who it drove away.

Gemini 3.5 Flash-Lite · Gemini

It replaces the five-day take-home with a live 45-minute case, citing the withdrawal data and who it drove away.

Fits the people and the timeRightRightRight
Opus 5.5 · Claude

Each interviewer's time sums to at most 2 hours, and the timeline fits within three weeks.

Sonnet 5.5 · API

The time budget shows each interviewer under 120 minutes, and the schedule fits the three-week target from first interview to offer.

Gemini 3.5 Flash-Lite · Gemini

Each interviewer's time is within two hours, and the loop can be completed in three weeks.

Judges on evidence, not presenceRightRightRight
Opus 5.5 · Claude

The scorecard requires specific evidence of results and bans impression words like 'presence' and 'smart'.

Sonnet 5.5 · API

The process requires specific observed behaviors and concrete evidence, explicitly rejecting 'presence' and 'smart' as evidence.

Gemini 3.5 Flash-Lite · Gemini

The scorecard and process require concrete evidence of results and collaboration, not impressions like presence.

Defines good for this role firstRightRightRight
Opus 5.5 · Claude

It defines competencies with strong/weak signals, weighted for the role, before designing interviews.

Sonnet 5.5 · API

The scorecard defines five weighted competencies with strong and weak signals for this specific senior growth PM role before evaluating candidates.

Gemini 3.5 Flash-Lite · Gemini

The scorecard defines competencies and strong/weak signals before the interview design.

Keeps each judgement independentRightRightRight
Opus 5.5 · Claude

It requires written independent evaluations before debrief and orders speaking to avoid anchoring.

Sonnet 5.5 · API

Interviewers submit written ratings and evidence before the debrief, and scores are discussed against evidence, preventing the loudest voice from dominating.

Gemini 3.5 Flash-Lite · Gemini

It requires independent written evaluations before discussion and no show of hands.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review