Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're helping Amara Osei, VP Product at Copperline, hire our first Growth PM (a senior role). She's shared her draft interview loop and asked you to redesign it. Write the loop you'd run (each round: who runs it, what it tests, how long it takes), the scorecard (what strong and weak look like for each thing we're testing), and how we'll make the decision at the end. No more than 900 words. What we know is below.

What the model was given6 items: About Copperline, The role, Amara's draft loop, The last two PM hires, From the recruiter, Who can interview
About CopperlineInvoicing and payments software for small accountancy firms. 60 people. Trial to paid conversion is 9%.
The roleOwns trial conversion and expansion revenue. Works with three engineers and a designer, reports to Amara, and has to ship experiments every week. Needs Sales and Marketing to change how they work, without managing them.
Amara's draft loop1. Recruiter screen. 2. Take-home: 'design a growth strategy for Copperline', a 10-page deck, five days to complete. 3. Presentation to a panel of six. 4. Culture fit with the CEO: 'would I grab a beer with them?' 5. Debrief: everyone discusses, then votes by show of hands. There's no scorecard.
The last two PM hiresOne left after five months: 'couldn't get Sales to change anything'. The other is doing well. Both sets of interview notes say mainly 'great presence' and 'very smart'.
From the recruiterOf the last 40 candidates given the take-home, 14 withdrew, saying they didn't have time. Nine of the 14 had young children or demanding current roles.
Who can interviewAmara, the Head of Engineering, a senior designer, the Head of Sales, a data analyst and the CEO. Each can give at most two hours to this hire. Amara wants an offer out within three weeks of the first interview.
What a strong answer doesThe answer key the graders mark against

A loop built on a scorecard written first: for a senior growth PM, experimentation and analytics, execution, and influence without authority (weighted up, since the last hire failed on it), with strong and weak signals for each. Rounds test the real job: a live diagnosis of Copperline's own trial funnel instead of a five-day generic deck (which loses candidates, mostly parents and people in demanding roles), a behavioural round on past experiments and metrics actually moved, and an influence round run by the Head of Sales. Replaces the 'grab a beer' culture fit with a structured values or collaboration conversation. Each interviewer writes their evaluation before the debrief, and the decision is made against the scorecard, not a show of hands. Fits each interviewer's two hours and the three-week target, and adds reference checks.

Critical failures (cap the score at 40)
  • Keeps a culture-fit test based on whether the interviewer would like a beer with the candidate
Task

Mid-level · v1.0 · synthetic · B2B SaaS, hiring, growth

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 12 checks passed

The output is clear, actionable, and covers all required parts; a quick edit to add conditions or label interpretations would make it fully ready.

Output

Amara—here is the redesigned interview loop. Your draft had a fatal flaw: the 10-page take-home is a diversity filter (driving away parents and senior operators) and tests academic strategy over execution. Furthermore, our last hiring failure showed we hired for "presence" rather than the ability to influence Sales.

Here is a 3-week, high-signal loop designed to test gritty experimentation, cross-functional persuasion, and data fluency.

---

Part 1: The Interview Loop

Round 1: Recruiter Screen * Who: Recruiter (30 mins) * Tests: Baseline communication, compensation alignment, current trial-to-paid context, and general interest in B2B SaaS.

Round 2: The Operating Screen (Live Case) * Who: Amara, VP Product (45 mins, video) * Tests: Problem decomposition, metric intuition, and ability to move fast. * Format: No slides. We present our 9% trial-to-paid conversion rate and a specific funnel drop-off. We spend 30 minutes live-whiteboarding how they would diagnose and attack it within their first 30 days.

Round 3: The Cross-Functional Influence & Craft Deep-Dive * Who: Head of Engineering & Senior Designer (45 mins, video) * Tests: Can this person ship weekly? Do they respect engineering constraints? Do they partner well with design on user psychology? * Format: Deep-dive into a past project where they shipped rapid iterations. How do they handle tech debt vs. speed tradeoffs?

Round 4: The Sales & Data Crucible * Who: Head of Sales & Data Analyst (45 mins, video) * Tests: Crucial. Can they influence Sales without authority? Can they audit data to find buried conversion leaks? * Format: A roleplay. The Head of Sales plays a skeptical account executive who thinks the PM's new trial-onboarding flow is stealing their pipeline. The candidate must handle the pushback, find common ground, and align incentives.

Round 5: Founder Alignment & Values * Who: CEO (30 mins, in-person or video) * Tests: Velocity mindset, ownership mentality, and cultural add (dropping the "beer test" bias, which screens for homogeneity). * Format: Discussion on Copperline’s mission, how they handle failure, and what autonomy looks like here.

---

Part 2: The Scorecard

We will grade candidates on four core competencies using a 1–4 scale (1 = Unacceptable, 4 = Exceptional).

#### 1. Rigorous Experimentation & Velocity * Strong (3–4): Relentlessly focused on the funnel. Talks easily about cadence, MVP scoping, and shipping weekly. Bases ideas on user behavior, not just "best practices." * Weak (1–2): Prefers massive, multi-month redesign projects. Hand-waves about metrics or treats data as an afterthought.

#### 2. Cross-Functional Influence (Specifically with Sales) * Strong (3–4): Listens to Sales objections, understands quota pressure, and designs incentives that make Sales want to use the new product motions. Uses data to persuade. * Weak (1–2): Views Sales as "resistant" or "the enemy." Relies on positional authority ("Amara said we're doing this") rather than building trust.

#### 3. Data Fluency * Strong (3–4): Can spot a vanity metric instantly. Identifies cohort retention issues and funnel bottlenecks unprompted. Comfortable querying or partnering deeply with analysts. * Weak (1–2): Confuses activity metrics (clicks) with outcome metrics (expansion revenue). Gets lost when looking at raw conversion funnels.

#### 4. Product Craft & Empathy (Accountants) * Strong (3–4): Quickly grasps the nuanced workflow of a small accountancy firm. Designs intuitive onboarding that respects the user's time. * Weak (1–2): Generic "growth hacking" mentality that applies consumer tactics (spammy pop-ups, dark patterns) to professional B2B software.

---

Part 3: Decision-Making Process

To avoid the vague "they seemed smart" trap, we will use a structured debrief.

  1. Independent Scoring: Within 24 hours of their final round, every interviewer must submit their independent scores and notes into a shared doc before reading anyone else's.
  2. The Debrief Meeting (30 mins): Led by Amara. We do not "vote by show of hands." Instead, each interviewer speaks to their specific competency (e.g., Sales speaks to Round 4, Eng speaks to Round 3).
  3. The Bar: A hire requires a minimum score of 3 across all four competencies, with no score of 1 or 2 in Cross-Functional Influence. If Sales or Engineering flags a collaboration red flag, it's a hard veto.

By cutting the take-home, we keep top talent in the funnel. By replacing the "beer test" with a Sales roleplay, we ensure our next PM doesn't repeat the exact failure mode of our last one.

Claims the judge couldn’t find in the brief

  • The beer test screens for homogeneity.

Grades and run details

Decision model 79 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly39%
  • passAddresses the actual decision81%
  • partialRespects explicit constraints38%
  • failIdentifies material uncertainty70%
  • failAvoids unsupported claims33%
  • passProduces the required deliverable53%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time24%
  • passJudges on evidence, not presence67%
  • passDefines good for this role first84%
  • passKeeps each judgement independent99%
Run
Run
#1
Time to output
6 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 12 checks passed

The output is a complete, actionable plan that a product manager could implement with minimal editing, though adding a note on uncertainty would strengthen it.

Output

Here is the redesigned loop, scorecard, and decision framework tailored to Copperline’s constraints and failure modes.

---

1. The Redesigned Loop

We replace the 10-page take-home (which filters out working parents and busy seniors) with a Live Working Session using a 1-page data brief sent 24 hours prior (max 30 minutes prep). We eliminate the 6-person panel and beer test in favor of targeted, structured interviews.

Total interviewer commitment: Amara (105 min), Data Analyst (75 min), Head of Sales (75 min), Head of Eng (75 min), Designer (75 min), CEO (45 min). All under the 2-hour cap.

RoundInterviewer(s)Focus & FormatDuration
:---:---:---:---
0. Recruiter ScreenRecruiterComp, logistics, baseline communication, role alignment.30 min
1. HM ScreenAmara (VP Product)Growth philosophy, track record of weekly experiment cadences, why Copperline.30 min
2. Live Growth JamAmara & Data AnalystSimulated Work: 1-page brief on Copperline’s 9% trial conversion. Candidate leads a 30-min problem-solving session to generate, prioritize, and design 3 testable hypotheses.45 min
3. Influence & Stakeholder AlignmentHead of SalesPast Behavioral: Changing sales/marketing behavior without authority, handling resistance, aligning self-serve vs. assisted motions.45 min
4. Pod Execution & VelocityHead of Eng & Senior DesignerCollaboration: How they run a 4-person growth pod, maintain weekly shipping velocity, make scoping trade-offs, and treat engineering/design as thought partners.45 min
5. Values & AmbitionCEOStructured interview assessing ownership, comfort with 60-person ambiguity, and customer obsession for small accounting firms.30 min
6. Async DebriefAmara (Lead)Written-first calibration and final decision.30 min

Pacing: Rounds 2–5 can be scheduled across two consecutive days. Total timeline from screen to offer: 10–14 business days.

---

2. The Scorecard

Interviewers evaluate against specific evidence, not “smartness” or “presence.”

#### A. Influence Without Authority (Crucial copper-plate failure mode) Evaluated by: Head of Sales, Amara * Strong: Uses shared metrics to align incentives; brings Sales qualitative feedback into experiment design; creates lightweight workflows that save reps time; shows empathy for quota-carrying teams; demonstrates wins where an adversarial relationship was turned into a partnership. * Weak: Relies on executive mandate (“Amara said so”); views Sales as an execution arm; blames others when processes fail (“Sales refused to update the CRM”); uses heavy-handed documentation instead of relationship-building.

#### B. Experiment Velocity & Growth Rigor (9% trial-to-paid target) Evaluated by: Amara, Data Analyst * Strong: Understands the mechanics of PLG onboarding for B2B accounting workflows; defines clear leading indicators; comfortable shipping imperfect MVPs to learn fast; ruthlessly prioritizes based on impact vs. effort; embraces invalidated tests as learning. * Weak: Proposes massive, multi-month redesigns instead of iterative tests; lacks mathematical intuition around sample sizes and statistical significance; fixates on vanity top-of-funnel metrics instead of revenue and conversion.

#### C. Cross-Pod Execution & Scoping Evaluated by: Head of Engineering, Senior Designer * Strong: Involves design and engineering early in hypothesis generation; defines crisp acceptance criteria; scopes experiments down to hours/days, not weeks; protects pod morale by celebrating learning over pure "wins." * Weak: Hands fully formed PRDs over the fence; treats engineers as ticket-takers; expects pixel-perfection on throwaway split tests; cannot articulate trade-offs between tech debt and learning speed.

#### D. Customer Domain Empathy & Ownership Evaluated by: CEO * Strong: Shows genuine curiosity about the mundane operational pains of small accountants; low ego, high accountability; self-directed; articulates clear reasons for joining an early-stage B2B SaaS business. * Weak: Generic answers suited to consumer tech; dismissive of non-glamorous industries; expects corporate infrastructure and large support teams.

---

3. Decision-Making Framework

To prevent the groupthink, bias, and anchoring seen in show-of-hands voting:

1. Independent Scorecard Submission: Within 2 hours of their interview, every interviewer must submit their score (Strong Hire, Hire, Lean No Hire, Strong No Hire) and written evidence mapped to their specific competencies. Scores are locked; no one can view others' ratings until their own is submitted. 2. The 30-Minute Debrief: * Amara chairs the meeting. * If any interviewer marks Lean No Hire or Strong No Hire, the discussion starts entirely on those specific concerns. * The Head of Sales holds an effective veto on stakeholder influence. If the Head of Sales says, "This person will alienate my team," the candidate is disqualified. * "Culture fit" is explicitly redefined: We do not hire people we want to grab a beer with; we hire people who elevate our operating standard and live our core values. 3. Single Decision-Maker: This is not a democracy. Amara owns the final decision and accountability for the hire. If the scorecard is clean and Sales approves, Amara makes the call immediately.

This process directly solves candidate drop-off, screens for Copperline's exact historical failure mode, and enables an offer inside two weeks.

Grades and run details

Decision model 83 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly17%
  • passAddresses the actual decision66%
  • partialRespects explicit constraints31%
  • failIdentifies material uncertainty51%
  • partialAvoids unsupported claims54%
  • passProduces the required deliverable89%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time87%
  • passJudges on evidence, not presence96%
  • passDefines good for this role first86%
  • passKeeps each judgement independent100%
Run
Run
#1
API response time
29 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 12 checks passed

The output is complete and usable; the only missing piece is a brief mention of the recruiter's withdrawal data, which is a quick edit.

Output

Recommended loop

Replace the unpaid strategy deck, six-person presentation and “beer test” with structured evidence of doing this job. The previous hires’ notes cannot explain why one succeeded and one failed; cross-functional influence must be a hiring requirement, not a personality impression.

Rounds

RoundWho / durationWhat it tests and how
1. Structured screenRecruiter, 20 minutesConfirm scope, motivation, compensation and logistics. Ask for one conversion or expansion initiative: personal contribution, measured result and partner teams involved. Use consistent questions; accept relevant experience without requiring a Growth PM title. Advance at most two finalists under the interviewer-time budget.
2. Ownership deep diveAmara, 35 minutesExamine one growth initiative and one unsuccessful attempt to change another team’s behavior. Probe baseline, decision-making, individual contribution, results, uncertainty and lessons. Test senior ownership, growth judgment and influence through specific actions—not polished storytelling.
3. Live growth work sampleHead of Engineering + data analyst, 40 minutesGive everyone the same short, synthetic packet: a funnel beginning at Copperline’s 9% trial conversion, cohort/segment data, expansion opportunities, customer observations and delivery constraints. Allow eight minutes of quiet reading within the session. Ask candidates to diagnose, prioritize conversion versus expansion opportunities, and scope the first weekly experiment with three engineers and a designer. Introduce one standardized complication. Test analysis, prioritization, experimentation and delivery. No deck or advance work.
4. Cross-functional working sessionHead of Sales + senior designer, 35 minutesCandidate proposes a trial-conversion change. Sales resists because it changes incentives or workload; Design flags customer friction. Use identical opening objections, while responding naturally thereafter. Candidate must uncover interests, negotiate an achievable pilot and establish responsibilities and follow-through. Finish with reflection on what they would change. Test influence, customer judgment and collaboration—not whether they “win.”
5. Offer conversation—not another assessmentCEO, 20 minutes, selected candidate onlyExplain company direction, answer questions and help close. No informal culture-fit veto.

Allow candidate questions in each assessment round. Share the format, competencies and preparation expectations beforehand; offer scheduling flexibility, breaks and reasonable accommodations. Assess reasoning and decisions, not presentation polish. The work sample is fictional, not free consulting.

Scorecard

Use 1 = clear contrary evidence, 2 = below bar, 3 = meets senior-role bar, 4 = exceptional evidence. “Not observed” is separate from a low score and must be resolved before deciding.

Competency / primary evidence ownerStrong: meets or exceeds barWeak: below bar
Growth diagnosis and measurement / analystDefines conversion and expansion denominators, examines cohorts and segments, distinguishes correlation from causation, identifies missing evidence. Treats 9% as a starting point, not a diagnosis.Jumps to tactics; misreads funnels; claims causality from movement; ignores retention or expansion economics.
Experiment design and prioritization / analyst, AmaraChooses a defensible hypothesis, outcome and guardrails; weighs impact, effort and confidence; handles limited traffic and inconclusive results.Produces a tactic list; promises unsupported lifts; assumes every weekly release can deliver statistically conclusive results.
Weekly delivery with a small team / EngineeringSlices a useful first release, identifies instrumentation and dependencies, trades scope deliberately, and plans the next decision.Requires a major rebuild or more staff; ignores technical uncertainty; equates speed with skipping measurement or quality.
Influence without authority / SalesInvestigates incentives and constraints, makes a credible mutual-value case, secures specific commitments and establishes follow-through. Past examples show personal action and honest limits.Relies on title, escalation or charisma; blames Sales; mistakes agreement in a meeting for changed behavior.
Customer and commercial judgment / DesignConnects accountancy-firm workflows to conversion and expansion; protects trust and usability; considers downstream retention and revenue quality.Optimizes conversion through friction, pressure or misleading choices; ignores customer costs or churn.
Senior ownership and learning / AmaraSeparates own contribution from team results, owns misses, changes course with evidence and creates a repeatable operating rhythm.Offers vague achievements, cannot explain decisions, takes all credit or externalizes failures.

Interviewers record concrete actions, quotes and reasoning alongside scores. “Smart,” “great presence” and “likeable” are not evidence.

Decision, capacity and timeline

Week 1: recruiter screens and schedules two finalists. Week 2: complete assessments. Week 3: decide, close and issue the offer.

Reserve each assessor 10 minutes for calibration, five minutes per finalist for scoring, and 20 minutes for debrief. Including interviews, Engineering and the analyst each use exactly 120 minutes; Amara, Sales and Design each use 110. CEO uses at most 40 minutes. If neither finalist qualifies, reopen rather than lower the bar or silently exceed the budget.

Submit scores independently before seeing others’ notes. In debrief, review evidence competency by competency; Amara speaks last. No show-of-hands vote.

Require at least 3 on every competency, especially influence and weekly delivery. Resolve disagreements through recorded evidence, not averaging or seniority. Missing critical evidence means no offer yet. Amara owns the final decision and documents the rationale, remaining risks and onboarding support.

Grades and run details

Decision model 92 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly54%
  • passAddresses the actual decision93%
  • passRespects explicit constraints23%
  • partialIdentifies material uncertainty51%
  • partialAvoids unsupported claims29%
  • passProduces the required deliverable91%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates80%
  • passFits the people and the time75%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first95%
  • passKeeps each judgement independent100%
Run
Run
#1
Time to output
62 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Addresses the actual decisionMixedRightRight
Gemini 3.5 Flash-Lite · Gemini

The output does not state what result or condition would change the recommended loop.

Gemini 3.8 Flash · API

The output commits to a specific redesigned loop, scorecard, and decision framework, which is the answer requested.

GPT-6 Astra · ChatGPT

The output commits to a specific redesigned loop and decision process, with Amara owning the final call and a clear condition to reopen if no candidate meets the bar.

Identifies material uncertaintyWrongWrongRight
Gemini 3.5 Flash-Lite · Gemini

The output does not name unknowns that could change the decision or how they would be resolved.

Gemini 3.8 Flash · API

The output does not name any unknowns that could change the loop design or decision process, nor does it say how they would be resolved.

GPT-6 Astra · ChatGPT

It names the condition that would change the decision (if neither finalist qualifies, reopen) and that missing critical evidence means no offer.

Avoids unsupported claimsWrongWrongRight
Gemini 3.5 Flash-Lite · Gemini

It presents 'diversity filter' and 'screens for homogeneity' as established facts without labelling them as interpretations.

Gemini 3.8 Flash · API

The output presents forecasts like 'This process directly solves candidate drop-off' as established fact without labelling them as hypotheses.

GPT-6 Astra · ChatGPT

The only interpretation (notes can't explain outcomes) is directly supported by the brief; no other claims are presented as established fact.

Fixes the take-home's cost to candidatesRightRightMixed
Gemini 3.5 Flash-Lite · Gemini

It replaces the five-day take-home with a live 45-minute case, citing the withdrawal data and who it drove away.

Gemini 3.8 Flash · API

The five-day take-home is replaced with a live, bounded exercise, and the withdrawal data is cited to justify the change.

GPT-6 Astra · ChatGPT

The output replaces the take-home with a live exercise but never cites the recruiter's withdrawal data or who the take-home was driving away.

All got right 8

Uses the supplied evidence correctlyRightRightRight
Gemini 3.5 Flash-Lite · Gemini

All facts and figures about the current situation are taken correctly from the brief; no invented data.

Gemini 3.8 Flash · API

All statements about the current situation are taken directly from the supplied context.

GPT-6 Astra · ChatGPT

The output uses the 9% conversion rate and the insufficient interview notes correctly, and invents no facts about the current situation.

Respects explicit constraintsRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The loop respects all constraints: length, no beer test, no unchanged take-home, interviewer time limits, and three-week target.

Gemini 3.8 Flash · API

The output is under 900 words, addresses the VP Product, and respects all stated constraints.

GPT-6 Astra · ChatGPT

The output is a loop, scorecard, and decision process within 900 words, addressed to Amara, and respects all stated constraints.

Produces the required deliverableRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The output includes the loop, scorecard, and decision process, within 900 words, and is usable by Amara.

Gemini 3.8 Flash · API

The output provides a complete loop, scorecard, and decision framework that a product manager could act on with light edits.

GPT-6 Astra · ChatGPT

The output provides a complete loop, scorecard, and decision process that a PM could act on with light edits.

Tests what the last hire failed atRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Round 4 with the Head of Sales specifically tests influencing Sales without authority, with strong and weak signals.

Gemini 3.8 Flash · API

A round with the Head of Sales specifically tests influencing teams without authority, with strong and weak signals defined.

GPT-6 Astra · ChatGPT

Round 4 is a cross-functional working session with the Head of Sales, and the scorecard includes 'Influence without authority' with strong/weak signals.

Fits the people and the timeRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Each interviewer's time is within two hours, and the loop can be completed in three weeks.

Gemini 3.8 Flash · API

Each interviewer's time is under two hours, and the total timeline of 10-14 business days fits the three-week target.

GPT-6 Astra · ChatGPT

Interviewer times are explicitly kept within two hours (Engineering and analyst 120 min, others 110 min, CEO 40 min) and the three-week schedule is detailed.

Judges on evidence, not presenceRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The scorecard and process require concrete evidence of results and collaboration, not impressions like presence.

Gemini 3.8 Flash · API

The scorecard and process require concrete evidence of results and collaboration, not impressions like presence or smartness.

GPT-6 Astra · ChatGPT

The scorecard demands concrete actions, quotes, and reasoning, and explicitly rejects 'smart', 'great presence', and 'likeable' as evidence.

Defines good for this role firstRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The scorecard defines competencies and strong/weak signals before the interview design.

Gemini 3.8 Flash · API

The scorecard defines the weighted competencies for this role with strong and weak signals before the decision framework.

GPT-6 Astra · ChatGPT

The scorecard defines six competencies with strong and weak signals before any interview design, weighting influence and delivery for this role.

Keeps each judgement independentRightRightRight
Gemini 3.5 Flash-Lite · Gemini

It requires independent written evaluations before discussion and no show of hands.

Gemini 3.8 Flash · API

Interviewers must submit written evaluations independently before any group discussion, preventing groupthink.

GPT-6 Astra · ChatGPT

Interviewers must submit scores independently before seeing others' notes, and Amara speaks last; no show-of-hands vote.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review