Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're helping Amara Osei, VP Product at Copperline, hire our first Growth PM (a senior role). She's shared her draft interview loop and asked you to redesign it. Write the loop you'd run (each round: who runs it, what it tests, how long it takes), the scorecard (what strong and weak look like for each thing we're testing), and how we'll make the decision at the end. No more than 900 words. What we know is below.

What the model was given6 items: About Copperline, The role, Amara's draft loop, The last two PM hires, From the recruiter, Who can interview
About CopperlineInvoicing and payments software for small accountancy firms. 60 people. Trial to paid conversion is 9%.
The roleOwns trial conversion and expansion revenue. Works with three engineers and a designer, reports to Amara, and has to ship experiments every week. Needs Sales and Marketing to change how they work, without managing them.
Amara's draft loop1. Recruiter screen. 2. Take-home: 'design a growth strategy for Copperline', a 10-page deck, five days to complete. 3. Presentation to a panel of six. 4. Culture fit with the CEO: 'would I grab a beer with them?' 5. Debrief: everyone discusses, then votes by show of hands. There's no scorecard.
The last two PM hiresOne left after five months: 'couldn't get Sales to change anything'. The other is doing well. Both sets of interview notes say mainly 'great presence' and 'very smart'.
From the recruiterOf the last 40 candidates given the take-home, 14 withdrew, saying they didn't have time. Nine of the 14 had young children or demanding current roles.
Who can interviewAmara, the Head of Engineering, a senior designer, the Head of Sales, a data analyst and the CEO. Each can give at most two hours to this hire. Amara wants an offer out within three weeks of the first interview.
What a strong answer doesThe answer key the graders mark against

A loop built on a scorecard written first: for a senior growth PM, experimentation and analytics, execution, and influence without authority (weighted up, since the last hire failed on it), with strong and weak signals for each. Rounds test the real job: a live diagnosis of Copperline's own trial funnel instead of a five-day generic deck (which loses candidates, mostly parents and people in demanding roles), a behavioural round on past experiments and metrics actually moved, and an influence round run by the Head of Sales. Replaces the 'grab a beer' culture fit with a structured values or collaboration conversation. Each interviewer writes their evaluation before the debrief, and the decision is made against the scorecard, not a show of hands. Fits each interviewer's two hours and the three-week target, and adds reference checks.

Critical failures (cap the score at 40)
  • Keeps a culture-fit test based on whether the interviewer would like a beer with the candidate
Task

Mid-level · v1.0 · synthetic · B2B SaaS, hiring, growth

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 12 checks passed

The output is complete and actionable; a quick edit to mention the recruiter's withdrawal data would make it fully aligned with the brief.

Output

Interview loop

Run this over 2–3 weeks. Skip the five-day take-home, large presentation panel, and “beer” fit interview. Share the same expectations and questions with every candidate; offer scheduling flexibility and accommodations.

1. Recruiter screen — 30 minutes Tests role understanding, motivation, and practical alignment: the scope, weekly shipping pace, and need to influence Sales and Marketing without authority. Not a “polish” screen.

2. Amara, VP Product — 60 minutes Tests growth ownership and product judgment. Ask for a detailed example of a funnel or revenue problem the candidate owned: diagnosis, choices, experiments, results, and what they learned. Probe how they balance conversion, customer trust, and expansion.

3. Live growth work sample — 60 minutes Send a one-page brief and synthetic funnel data 24 hours ahead; no deck or advance work required. Spend 30 minutes with the data analyst diagnosing the funnel and proposing a measurable experiment; spend 30 minutes with the Head of Engineering and senior designer scoping how to ship and learn from it. Use the same prompt and time limits for all candidates.

4. Head of Sales — 45 minutes Structured roleplay: Sales is reluctant to change its trial follow-up because reps are busy and doubt the data. Ask the candidate to understand the objection, make a case, and agree a practical next step. Then probe a real example of influencing a team they did not manage.

The CEO does not need a separate round: “culture fit” is not a hiring criterion, and the role’s key cross-functional test is better assessed directly with Sales, Engineering, and Design.

Scorecard

Use a 1–4 rating for each dimension: 1 = clear concern, 2 = below bar, 3 = meets bar, 4 = unusually strong. Record evidence, not impressions.

DimensionStrong evidenceWeak evidence
Growth ownership and outcomes (Amara)Personally owned a meaningful funnel or revenue outcome; explains decisions, results, attribution limits, and lessons.Describes activity rather than impact; overclaims credit or cannot explain what changed.
Customer and product judgment (Amara)Understands small-firm workflows and trust needs; weighs conversion against customer value and expansion.Offers generic SaaS tactics or pushes short-term conversion at the expense of trust.
Data diagnosis and experiment design (work sample)Finds plausible bottlenecks and segments; proposes a focused test with a clear hypothesis, success metric, guardrail, and learning plan.Jumps to a solution, misreads the data, uses vanity metrics, or cannot say what result would change their mind.
Shipping and learning cadence (work sample)Scopes a feasible experiment, identifies dependencies and tradeoffs, and shows how the team can learn and iterate weekly.Proposes a big-bang project, ignores constraints, or treats learning as something to do after launch.
Influence without authority (Sales)Listens for incentives and concerns, uses evidence appropriately, adapts the ask, and secures a credible next step.Blames Sales, relies on escalation or authority, or cannot turn disagreement into action.
Cross-functional partnership (work sample)Makes clear tradeoffs with Engineering and Design, invites their expertise, and aligns on ownership and delivery.Treats partners as order-takers or leaves scope and responsibilities vague.

Decision

Interviewers submit ratings and specific evidence independently before a 30-minute debrief; do not vote by show of hands. Amara facilitates, checks for evidence behind claims, and makes the hiring decision. Do not use “presence,” “smart,” or general “fit” as evidence.

A hire must score at least 3 on influence, experiment design, and shipping, with no score below 2 in any dimension. Do not average away a weakness in a core requirement. If evidence is genuinely missing, identify one focused follow-up rather than adding a broad round. Give candidates a timely decision and close the loop within three weeks of the first interview.

Grades and run details

Decision model 96 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly51%
  • passAddresses the actual decision82%
  • passRespects explicit constraints26%
  • partialIdentifies material uncertainty64%
  • passAvoids unsupported claims22%
  • passProduces the required deliverable78%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates68%
  • passFits the people and the time64%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first86%
  • passKeeps each judgement independent98%
Run
Run
#1
API response time
39 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 12 checks passed

The output is complete and usable; the only missing piece is a brief mention of the recruiter's withdrawal data, which is a quick edit.

Output

Recommended loop

Replace the unpaid strategy deck, six-person presentation and “beer test” with structured evidence of doing this job. The previous hires’ notes cannot explain why one succeeded and one failed; cross-functional influence must be a hiring requirement, not a personality impression.

Rounds

RoundWho / durationWhat it tests and how
1. Structured screenRecruiter, 20 minutesConfirm scope, motivation, compensation and logistics. Ask for one conversion or expansion initiative: personal contribution, measured result and partner teams involved. Use consistent questions; accept relevant experience without requiring a Growth PM title. Advance at most two finalists under the interviewer-time budget.
2. Ownership deep diveAmara, 35 minutesExamine one growth initiative and one unsuccessful attempt to change another team’s behavior. Probe baseline, decision-making, individual contribution, results, uncertainty and lessons. Test senior ownership, growth judgment and influence through specific actions—not polished storytelling.
3. Live growth work sampleHead of Engineering + data analyst, 40 minutesGive everyone the same short, synthetic packet: a funnel beginning at Copperline’s 9% trial conversion, cohort/segment data, expansion opportunities, customer observations and delivery constraints. Allow eight minutes of quiet reading within the session. Ask candidates to diagnose, prioritize conversion versus expansion opportunities, and scope the first weekly experiment with three engineers and a designer. Introduce one standardized complication. Test analysis, prioritization, experimentation and delivery. No deck or advance work.
4. Cross-functional working sessionHead of Sales + senior designer, 35 minutesCandidate proposes a trial-conversion change. Sales resists because it changes incentives or workload; Design flags customer friction. Use identical opening objections, while responding naturally thereafter. Candidate must uncover interests, negotiate an achievable pilot and establish responsibilities and follow-through. Finish with reflection on what they would change. Test influence, customer judgment and collaboration—not whether they “win.”
5. Offer conversation—not another assessmentCEO, 20 minutes, selected candidate onlyExplain company direction, answer questions and help close. No informal culture-fit veto.

Allow candidate questions in each assessment round. Share the format, competencies and preparation expectations beforehand; offer scheduling flexibility, breaks and reasonable accommodations. Assess reasoning and decisions, not presentation polish. The work sample is fictional, not free consulting.

Scorecard

Use 1 = clear contrary evidence, 2 = below bar, 3 = meets senior-role bar, 4 = exceptional evidence. “Not observed” is separate from a low score and must be resolved before deciding.

Competency / primary evidence ownerStrong: meets or exceeds barWeak: below bar
Growth diagnosis and measurement / analystDefines conversion and expansion denominators, examines cohorts and segments, distinguishes correlation from causation, identifies missing evidence. Treats 9% as a starting point, not a diagnosis.Jumps to tactics; misreads funnels; claims causality from movement; ignores retention or expansion economics.
Experiment design and prioritization / analyst, AmaraChooses a defensible hypothesis, outcome and guardrails; weighs impact, effort and confidence; handles limited traffic and inconclusive results.Produces a tactic list; promises unsupported lifts; assumes every weekly release can deliver statistically conclusive results.
Weekly delivery with a small team / EngineeringSlices a useful first release, identifies instrumentation and dependencies, trades scope deliberately, and plans the next decision.Requires a major rebuild or more staff; ignores technical uncertainty; equates speed with skipping measurement or quality.
Influence without authority / SalesInvestigates incentives and constraints, makes a credible mutual-value case, secures specific commitments and establishes follow-through. Past examples show personal action and honest limits.Relies on title, escalation or charisma; blames Sales; mistakes agreement in a meeting for changed behavior.
Customer and commercial judgment / DesignConnects accountancy-firm workflows to conversion and expansion; protects trust and usability; considers downstream retention and revenue quality.Optimizes conversion through friction, pressure or misleading choices; ignores customer costs or churn.
Senior ownership and learning / AmaraSeparates own contribution from team results, owns misses, changes course with evidence and creates a repeatable operating rhythm.Offers vague achievements, cannot explain decisions, takes all credit or externalizes failures.

Interviewers record concrete actions, quotes and reasoning alongside scores. “Smart,” “great presence” and “likeable” are not evidence.

Decision, capacity and timeline

Week 1: recruiter screens and schedules two finalists. Week 2: complete assessments. Week 3: decide, close and issue the offer.

Reserve each assessor 10 minutes for calibration, five minutes per finalist for scoring, and 20 minutes for debrief. Including interviews, Engineering and the analyst each use exactly 120 minutes; Amara, Sales and Design each use 110. CEO uses at most 40 minutes. If neither finalist qualifies, reopen rather than lower the bar or silently exceed the budget.

Submit scores independently before seeing others’ notes. In debrief, review evidence competency by competency; Amara speaks last. No show-of-hands vote.

Require at least 3 on every competency, especially influence and weekly delivery. Resolve disagreements through recorded evidence, not averaging or seniority. Missing critical evidence means no offer yet. Amara owns the final decision and documents the rationale, remaining risks and onboarding support.

Grades and run details

Decision model 92 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly54%
  • passAddresses the actual decision93%
  • passRespects explicit constraints23%
  • partialIdentifies material uncertainty51%
  • partialAvoids unsupported claims29%
  • passProduces the required deliverable91%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates80%
  • passFits the people and the time75%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first95%
  • passKeeps each judgement independent100%
Run
Run
#1
Time to output
62 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 12 checks passed

The output is a complete, actionable plan that a PM could implement with minor wording tweaks; no evidence re-checking or reasoning redo is needed.

Output

Growth PM interview loop

Assumptions: The two-hour cap is each person's total across all candidates, so I've designed for three finalists. Early rounds are cheap and don't use the panel's time. Only Amara's time is spent before the finalist stage.

Why the change

  • The 5-day, 10-page deck cost us 14 of 40 candidates, mostly people with caregiving duties or demanding jobs. It also tests deck-making, not the job.
  • The last hire failed on influencing Sales, and nothing in the old loop tested that.
  • "Great presence" and "very smart" in the notes are not evidence, and a show-of-hands vote rewards the loudest voice.

The loop

#RoundWhoTestsTime per candidate
1Screen + written questionRecruiter, scored by AmaraBasics, comp, and a 300-word written answer: "A conversion experiment you ran: baseline, result, what you'd change."30 min call, 20 min writing
2Ownership deep-diveAmaraTrack record on revenue outcomes20 min
3Funnel caseData analystConversion judgment30 min
4Ship-it scopingHead of Eng + senior designerWeekly shipping with a small team30 min
5Influence role-playHead of SalesChanging others' behavior without authority30 min
6Business and working styleCEOJudgment, first-hire fit20 min
7DebriefAll interviewersDecision20 min total

Rounds 3–6 can be done in one sitting, or split across two days on video. Total candidate time is about 3 hours.

Round details - Round 1: Amara spends 5 minutes on each of about six written answers (30 min total) and picks three finalists. - Round 3: We send a one-page anonymized trial funnel 24 hours ahead, with a note to spend no more than 30 minutes on it. In the room the candidate says where they'd look first, picks two experiments, and designs one, including metric, guardrail, and what a realistic sample size allows at our volume. - Round 4: We give a half-formed experiment idea. The candidate must cut it to something shippable within a week by three engineers and a designer. The interviewers push back on scope. - Round 5: The Head of Sales plays a skeptical rep or manager who thinks trial follow-up is a distraction. The candidate has to get agreement on a small pilot. The last 5 minutes cover a real past example. - Round 6: A structured conversation covering how the candidate would think about expansion revenue for small accountancy firms, what being the first growth hire means, and a time they disagreed and then committed. This replaces "would I grab a beer."

Scorecard

Each criterion is rated 1–4, with no midpoint. Every interviewer writes specific observed behaviors, not impressions. "Presence" and "smart" don't count as evidence.

1. Influence without authority (weight 30%; Sales, plus Amara and references) - Strong: Starts from what Sales cares about (quota, time, commissions). Proposes a small, reversible pilot. Offers something in return. Has a past example where a team they didn't manage changed behavior, with details. - Weak: Says "align stakeholders." Escalates to the boss. Presents the change as a mandate. Blames other teams in past stories.

2. Conversion judgment (25%; Analyst) - Strong: Segments before theorizing. Notes that our volume limits what tests can detect. Ties each experiment to a hypothesis and a guardrail. Asks about accountancy seasonality and who actually decides to pay. - Weak: Lists generic tactics (checklists, gamification). Has no view on sample size. Doesn't ask about the data.

3. Shipping cadence (20%; Eng + Design) - Strong: Cuts scope readily to the smallest test. Gives real numbers on past experiments per week or month. Treats engineers and designers as partners and takes their objections seriously. - Weak: Proposes large projects with distant payoffs. Writes specs without hypotheses. Defends scope instead of trading it off.

4. Ownership of revenue outcomes (15%; Amara) - Strong: Gives baselines, deltas, and their specific contribution. Talks candidly about experiments that failed. - Weak: Vague credit for team results. Describes outputs (launches, decks) instead of outcomes.

5. Judgment and fit for a 60-person company (10%; CEO) - Strong: Is energized by ambiguity. Builds their own tools and doesn't wait for resources. Disagrees directly and then commits. - Weak: Needs a playbook or a team. Can't prioritize without data. Becomes defensive when challenged.

Note for the Head of Sales: Score what the candidate did in the role-play, not whether you enjoyed being challenged.

Decision process

  1. Independent scoring. Each interviewer submits ratings and evidence before the debrief and can't see others' scores.
  2. Debrief (20 min). Amara reads out the scores, lowest first, on each criterion. Disagreements get discussed against the evidence, not general impressions.
  3. Must-haves. A score of 2 or below on influence or shipping means no offer. The last miss was an influence failure, and weak shipping defeats the weekly-experiment job.
  4. Ranking. Among candidates who pass, the highest weighted score wins. Close calls go to Amara.
  5. Decision owner. Amara decides, and there is no vote. The CEO can block only by citing specific observed evidence tied to the scorecard.
  6. References. The recruiter checks references for the top candidate, using Amara's script. At least one reference must be a cross-functional partner such as someone from sales or marketing, asking, "Did they change how your team worked? How?"

Timeline and time budget

  • Week 1: Recruiter screens and written questions. Amara picks finalists by the end of the week.
  • Week 2: Amara's deep-dives, then rounds 3–6 for all three finalists.
  • Week 3: Debrief on Monday, references on Tuesday and Wednesday, offer by Friday.
PersonTime
Amara30 min (written answers) + 60 min (3 deep-dives) + 20 min (debrief) = 110 min
Analyst, Sales, Eng, Design90 min (3 interviews) + 20 min (debrief) = 110 min each
CEO60 min + 20 min = 80 min

If more than three candidates reach round 2, Amara's time is the first limit. In that case, tighten the written screen before asking for more of her time.

Grades and run details

Decision model 92 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly30%
  • passAddresses the actual decision78%
  • passRespects explicit constraints28%
  • partialIdentifies material uncertainty22%
  • partialAvoids unsupported claims47%
  • passProduces the required deliverable90%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time79%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first92%
  • passKeeps each judgement independent99%
Run
Run
#1
API response time
53 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Addresses the actual decisionMixedRightRight
GPT-6 Luna · API

The output does not state what result or condition would change the proposed loop or scorecard.

GPT-6 Astra · ChatGPT

The output commits to a specific redesigned loop and decision process, with Amara owning the final call and a clear condition to reopen if no candidate meets the bar.

Sonnet 5.5 · API

The output commits unambiguously to a specific interview loop, scorecard, and decision process, as requested.

Identifies material uncertaintyWrongRightWrong
GPT-6 Luna · API

The output does not name any unknowns that could change the decision or how they would be resolved.

GPT-6 Astra · ChatGPT

It names the condition that would change the decision (if neither finalist qualifies, reopen) and that missing critical evidence means no offer.

Sonnet 5.5 · API

The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.

Fixes the take-home's cost to candidatesMixedMixedRight
GPT-6 Luna · API

The output replaces the take-home but does not cite the recruiter's withdrawal data or mention who it was driving away.

GPT-6 Astra · ChatGPT

The output replaces the take-home with a live exercise but never cites the recruiter's withdrawal data or who the take-home was driving away.

Sonnet 5.5 · API

The five-day take-home is replaced with a 300-word written answer and a 30-minute funnel case, citing the withdrawal data and who it drove away.

All got right 9

Uses the supplied evidence correctlyRightRightRight
GPT-6 Luna · API

The output makes no factual claims about the current situation, so it does not misuse any supplied evidence.

GPT-6 Astra · ChatGPT

The output uses the 9% conversion rate and the insufficient interview notes correctly, and invents no facts about the current situation.

Sonnet 5.5 · API

All statements about the current situation are taken directly from the brief or supplied context, with no invented facts.

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

The output respects all constraints: it replaces the take-home, removes the beer test, fits interviewer time limits and the three-week schedule, and stays under 900 words.

GPT-6 Astra · ChatGPT

The output is a loop, scorecard, and decision process within 900 words, addressed to Amara, and respects all stated constraints.

Sonnet 5.5 · API

The output respects the 900-word limit, replaces the take-home and beer test, fits the two-hour per interviewer and three-week timeline, and delivers the required components.

Avoids unsupported claimsRightRightRight
GPT-6 Luna · API

The output presents only proposals and avoids unsupported claims about causes or forecasts.

GPT-6 Astra · ChatGPT

The only interpretation (notes can't explain outcomes) is directly supported by the brief; no other claims are presented as established fact.

Sonnet 5.5 · API

The output does not present interpretations or forecasts as established facts; its claims are supported by the supplied evidence.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The output provides a complete interview loop, scorecard, and decision process that a product manager could act on.

GPT-6 Astra · ChatGPT

The output provides a complete loop, scorecard, and decision process that a PM could act on with light edits.

Sonnet 5.5 · API

The output provides a complete loop, scorecard, and decision process within the word limit, usable by Amara with light edits.

Tests what the last hire failed atRightRightRight
GPT-6 Luna · API

The loop includes a dedicated round with the Head of Sales and a scorecard dimension for influence without authority, directly testing what the last hire failed at.

GPT-6 Astra · ChatGPT

Round 4 is a cross-functional working session with the Head of Sales, and the scorecard includes 'Influence without authority' with strong/weak signals.

Sonnet 5.5 · API

Round 5 is an influence role-play with the Head of Sales, and the scorecard weights influence without authority at 30%, directly testing the last hire's failure point.

Fits the people and the timeRightRightRight
GPT-6 Luna · API

Each interviewer's time is within two hours, and the loop is designed to fit the three-week target.

GPT-6 Astra · ChatGPT

Interviewer times are explicitly kept within two hours (Engineering and analyst 120 min, others 110 min, CEO 40 min) and the three-week schedule is detailed.

Sonnet 5.5 · API

The time budget shows each interviewer under 120 minutes, and the schedule fits the three-week target from first interview to offer.

Judges on evidence, not presenceRightRightRight
GPT-6 Luna · API

The scorecard and decision rules explicitly require evidence over impressions and forbid using 'presence' or 'smart' as criteria.

GPT-6 Astra · ChatGPT

The scorecard demands concrete actions, quotes, and reasoning, and explicitly rejects 'smart', 'great presence', and 'likeable' as evidence.

Sonnet 5.5 · API

The process requires specific observed behaviors and concrete evidence, explicitly rejecting 'presence' and 'smart' as evidence.

Defines good for this role firstRightRightRight
GPT-6 Luna · API

The scorecard defines competencies with strong and weak signals, and the hiring bar weights influence, experiment design, and shipping.

GPT-6 Astra · ChatGPT

The scorecard defines six competencies with strong and weak signals before any interview design, weighting influence and delivery for this role.

Sonnet 5.5 · API

The scorecard defines five weighted competencies with strong and weak signals for this specific senior growth PM role before evaluating candidates.

Keeps each judgement independentRightRightRight
GPT-6 Luna · API

Interviewers must submit ratings and evidence independently before the debrief, and the process guards against groupthink.

GPT-6 Astra · ChatGPT

Interviewers must submit scores independently before seeing others' notes, and Amara speaks last; no show-of-hands vote.

Sonnet 5.5 · API

Interviewers submit written ratings and evidence before the debrief, and scores are discussed against evidence, preventing the loudest voice from dominating.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review