Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're helping Amara Osei, VP Product at Copperline, hire our first Growth PM (a senior role). She's shared her draft interview loop and asked you to redesign it. Write the loop you'd run (each round: who runs it, what it tests, how long it takes), the scorecard (what strong and weak look like for each thing we're testing), and how we'll make the decision at the end. No more than 900 words. What we know is below.

What the model was given6 items: About Copperline, The role, Amara's draft loop, The last two PM hires, From the recruiter, Who can interview
About CopperlineInvoicing and payments software for small accountancy firms. 60 people. Trial to paid conversion is 9%.
The roleOwns trial conversion and expansion revenue. Works with three engineers and a designer, reports to Amara, and has to ship experiments every week. Needs Sales and Marketing to change how they work, without managing them.
Amara's draft loop1. Recruiter screen. 2. Take-home: 'design a growth strategy for Copperline', a 10-page deck, five days to complete. 3. Presentation to a panel of six. 4. Culture fit with the CEO: 'would I grab a beer with them?' 5. Debrief: everyone discusses, then votes by show of hands. There's no scorecard.
The last two PM hiresOne left after five months: 'couldn't get Sales to change anything'. The other is doing well. Both sets of interview notes say mainly 'great presence' and 'very smart'.
From the recruiterOf the last 40 candidates given the take-home, 14 withdrew, saying they didn't have time. Nine of the 14 had young children or demanding current roles.
Who can interviewAmara, the Head of Engineering, a senior designer, the Head of Sales, a data analyst and the CEO. Each can give at most two hours to this hire. Amara wants an offer out within three weeks of the first interview.
What a strong answer doesThe answer key the graders mark against

A loop built on a scorecard written first: for a senior growth PM, experimentation and analytics, execution, and influence without authority (weighted up, since the last hire failed on it), with strong and weak signals for each. Rounds test the real job: a live diagnosis of Copperline's own trial funnel instead of a five-day generic deck (which loses candidates, mostly parents and people in demanding roles), a behavioural round on past experiments and metrics actually moved, and an influence round run by the Head of Sales. Replaces the 'grab a beer' culture fit with a structured values or collaboration conversation. Each interviewer writes their evaluation before the debrief, and the decision is made against the scorecard, not a show of hands. Fits each interviewer's two hours and the three-week target, and adds reference checks.

Critical failures (cap the score at 40)
  • Keeps a culture-fit test based on whether the interviewer would like a beer with the candidate
Task

Mid-level · v1.0 · synthetic · B2B SaaS, hiring, growth

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit12 of 12 checks passed

The output is complete, actionable, and respects all constraints; a PM could use it with at most minor wording tweaks.

Output

The loop I’d run

Replace the deck and six-person presentation with bounded, live work samples. Remove the “beer test.” Copperline needs someone who can diagnose conversion, ship weekly, and change other teams’ behavior—not simply present well.

Candidate commitment: 3 hours 10 minutes; no take-home or required preparation. Share the format and scorecard beforehand, offer flexible scheduling and reasonable accommodations, and use the same prompts and information for everyone.

RoundInterviewersTimeWhat it tests
---------:---
1. Recruiter screenRecruiter20 minScope, compensation, availability, and interest in hands-on growth ownership. Explain the process. Don’t assess polish or pedigree.
2. Ownership interviewAmara35 minSenior-level judgment and accountability. Examine one past growth initiative: candidate’s actual contribution, choices, results, failures, and learning. Probe both conversion and expansion experience. Reserve five minutes for candidate questions.
3. Live growth diagnosisData analyst50 minAnalytical judgment and experiment design. Give a small synthetic Copperline dataset: trial funnel, cohorts, acquisition sources, activation behavior, expansion signals, and metric definitions. Ask what could explain 9% conversion, what they would investigate, and which experiment they would run first. Include a plausible expansion opportunity.
4. Shipping workshopHead of Engineering + senior designer40 minTurn the previous round’s proposal into a one-week experiment with three engineers and a designer. Interviewers act as working partners, not an audience. Introduce a tracking gap and an engineering constraint. Test scope, customer experience, instrumentation, trade-offs, and rollout safety.
5. Influence simulationHead of Sales45 minFirst, role-play a conversation with Sales resisting a proposed trial-follow-up change because of workload and incentives. Add a Marketing objection about lead quality. Then examine a real example of influencing peers without authority, including resistance, compromises, and outcomes. Reserve five minutes for questions.

Use the same dataset, constraints, and role-play objections across candidates. Score reasoning, not whether they discover a predetermined “correct” growth strategy. The CEO does not conduct a separate fit interview.

Scorecard

Use 1 = clear concern, 2 = below bar, 3 = meets senior-role bar, 4 = exceptional, plus not observed. Each rating requires concrete evidence: a quote, action, artifact, or specific past example. “Very smart” and “great presence” are not evidence.

Dimension / primary assessorStrong: meets or exceeds barWeak: below bar
Growth and analytical judgment — analystChecks denominators, cohort maturity, segmentation, and data quality before explaining 9%. Separates correlation from causation; prioritizes plausible conversion and expansion opportunities using customer behavior and economics.Treats aggregate conversion as a diagnosis; jumps to tactics; ignores acquisition mix or revenue quality; makes causal claims without evidence.
Experimentation — analystStates hypothesis, mechanism, target population, primary metric, guardrails, and decision rule. Recognizes low-volume limitations. Distinguishes shipping weekly from obtaining statistically conclusive results weekly.Proposes many tests without prioritization; relies on vanity metrics; declares victory prematurely; cannot say what result would change the decision.
Execution and customer judgment — Engineering + designerCuts scope to a credible weekly release; specifies measurement and rollback; collaborates with technical/design partners; protects trust in invoicing and payments workflows.Requires a large rebuild; treats partners as order-takers; neglects instrumentation, reliability, or customer harm; cannot adapt to constraints.
Influence without authority — SalesUnderstands incentives and workload; listens, negotiates a bounded pilot, establishes shared measures and explicit commitments, and follows through. Past examples show actual behavior change.Relies on escalation, persuasion theater, or “getting buy-in”; blames peers; has no concrete mechanism for changing behavior.
Senior ownership and learning — AmaraClearly distinguishes personal contribution from team results; connects conversion and expansion to durable revenue; owns misses and demonstrates changed judgment.Claims collective wins as personal achievements; avoids failures; optimizes local conversion at the expense of retention or revenue; offers lessons without changed behavior.

Decision and operating plan

Before interviewing: Run a 15-minute calibration using sample answers and agree on anchors. Give interviewers structured prompts.

After each round: Interviewers submit evidence and scores independently within 24 hours, before seeing others’ assessments. Engineering and Design score independently.

Debrief: Hold a 30-minute evidence review—not a show-of-hands vote. Discuss disagreements and counterevidence against the written bar. Amara makes and documents the hiring decision.

Offer bar: At least 3 on every dimension. In particular, do not average away weak influence, analytical judgment, or execution. Missing evidence triggers one narrowly targeted follow-up, not intuition. Hire against the bar, not merely the strongest available candidate.

Timing and capacity: Finish interviews by day 12, debrief by day 14, and issue the offer by day 21. Calibration, interview, scoring, and debrief consume at most 105 minutes per domain interviewer. Amara retains time within her two-hour cap for decision documentation and an offer conversation. CEO approval, if required, concerns headcount and compensation—not an unstructured personality veto.

Grades and run details

Decision model 92 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly27%
  • passAddresses the actual decision92%
  • partialRespects explicit constraints26%
  • partialIdentifies material uncertainty54%
  • passAvoids unsupported claims21%
  • passProduces the required deliverable91%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates81%
  • passFits the people and the time85%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first89%
  • passKeeps each judgement independent100%
Run
Run
#1
API response time
55 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 12 checks passed

The output is complete and usable; the only missing piece is a brief mention of the recruiter's withdrawal data, which is a quick edit.

Output

Recommended loop

Replace the unpaid strategy deck, six-person presentation and “beer test” with structured evidence of doing this job. The previous hires’ notes cannot explain why one succeeded and one failed; cross-functional influence must be a hiring requirement, not a personality impression.

Rounds

RoundWho / durationWhat it tests and how
1. Structured screenRecruiter, 20 minutesConfirm scope, motivation, compensation and logistics. Ask for one conversion or expansion initiative: personal contribution, measured result and partner teams involved. Use consistent questions; accept relevant experience without requiring a Growth PM title. Advance at most two finalists under the interviewer-time budget.
2. Ownership deep diveAmara, 35 minutesExamine one growth initiative and one unsuccessful attempt to change another team’s behavior. Probe baseline, decision-making, individual contribution, results, uncertainty and lessons. Test senior ownership, growth judgment and influence through specific actions—not polished storytelling.
3. Live growth work sampleHead of Engineering + data analyst, 40 minutesGive everyone the same short, synthetic packet: a funnel beginning at Copperline’s 9% trial conversion, cohort/segment data, expansion opportunities, customer observations and delivery constraints. Allow eight minutes of quiet reading within the session. Ask candidates to diagnose, prioritize conversion versus expansion opportunities, and scope the first weekly experiment with three engineers and a designer. Introduce one standardized complication. Test analysis, prioritization, experimentation and delivery. No deck or advance work.
4. Cross-functional working sessionHead of Sales + senior designer, 35 minutesCandidate proposes a trial-conversion change. Sales resists because it changes incentives or workload; Design flags customer friction. Use identical opening objections, while responding naturally thereafter. Candidate must uncover interests, negotiate an achievable pilot and establish responsibilities and follow-through. Finish with reflection on what they would change. Test influence, customer judgment and collaboration—not whether they “win.”
5. Offer conversation—not another assessmentCEO, 20 minutes, selected candidate onlyExplain company direction, answer questions and help close. No informal culture-fit veto.

Allow candidate questions in each assessment round. Share the format, competencies and preparation expectations beforehand; offer scheduling flexibility, breaks and reasonable accommodations. Assess reasoning and decisions, not presentation polish. The work sample is fictional, not free consulting.

Scorecard

Use 1 = clear contrary evidence, 2 = below bar, 3 = meets senior-role bar, 4 = exceptional evidence. “Not observed” is separate from a low score and must be resolved before deciding.

Competency / primary evidence ownerStrong: meets or exceeds barWeak: below bar
Growth diagnosis and measurement / analystDefines conversion and expansion denominators, examines cohorts and segments, distinguishes correlation from causation, identifies missing evidence. Treats 9% as a starting point, not a diagnosis.Jumps to tactics; misreads funnels; claims causality from movement; ignores retention or expansion economics.
Experiment design and prioritization / analyst, AmaraChooses a defensible hypothesis, outcome and guardrails; weighs impact, effort and confidence; handles limited traffic and inconclusive results.Produces a tactic list; promises unsupported lifts; assumes every weekly release can deliver statistically conclusive results.
Weekly delivery with a small team / EngineeringSlices a useful first release, identifies instrumentation and dependencies, trades scope deliberately, and plans the next decision.Requires a major rebuild or more staff; ignores technical uncertainty; equates speed with skipping measurement or quality.
Influence without authority / SalesInvestigates incentives and constraints, makes a credible mutual-value case, secures specific commitments and establishes follow-through. Past examples show personal action and honest limits.Relies on title, escalation or charisma; blames Sales; mistakes agreement in a meeting for changed behavior.
Customer and commercial judgment / DesignConnects accountancy-firm workflows to conversion and expansion; protects trust and usability; considers downstream retention and revenue quality.Optimizes conversion through friction, pressure or misleading choices; ignores customer costs or churn.
Senior ownership and learning / AmaraSeparates own contribution from team results, owns misses, changes course with evidence and creates a repeatable operating rhythm.Offers vague achievements, cannot explain decisions, takes all credit or externalizes failures.

Interviewers record concrete actions, quotes and reasoning alongside scores. “Smart,” “great presence” and “likeable” are not evidence.

Decision, capacity and timeline

Week 1: recruiter screens and schedules two finalists. Week 2: complete assessments. Week 3: decide, close and issue the offer.

Reserve each assessor 10 minutes for calibration, five minutes per finalist for scoring, and 20 minutes for debrief. Including interviews, Engineering and the analyst each use exactly 120 minutes; Amara, Sales and Design each use 110. CEO uses at most 40 minutes. If neither finalist qualifies, reopen rather than lower the bar or silently exceed the budget.

Submit scores independently before seeing others’ notes. In debrief, review evidence competency by competency; Amara speaks last. No show-of-hands vote.

Require at least 3 on every competency, especially influence and weekly delivery. Resolve disagreements through recorded evidence, not averaging or seniority. Missing critical evidence means no offer yet. Amara owns the final decision and documents the rationale, remaining risks and onboarding support.

Grades and run details

Decision model 92 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly54%
  • passAddresses the actual decision93%
  • passRespects explicit constraints23%
  • partialIdentifies material uncertainty51%
  • partialAvoids unsupported claims29%
  • passProduces the required deliverable91%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates80%
  • passFits the people and the time75%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first95%
  • passKeeps each judgement independent100%
Run
Run
#1
Time to output
62 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 12 checks passed

The output is well-structured and complete; the one unsupported generalization can be removed with a quick edit without affecting the loop design.

Output

Copperline Growth PM: Interview Loop

Why the draft needs changing

  • The take-home is filtering out the wrong people. 14 of 40 candidates withdrew, and most of them were parents or in demanding jobs. A growth PM in a senior role is likely to be exactly that kind of person.
  • The panel of six and the "beer" test reward presence. Presence is what your interview notes captured last time, and it didn't predict success.
  • Nothing tested influence over Sales. That is why the last hire failed.

The new loop tests five things, each owned by one interviewer.

Time budget

Each interviewer has two hours in total for this hire. That limits the loop to three finalists. Per interviewer, the time breaks down as:

  • 3 × 30-minute interviews
  • 3 × 5-minute scorecards
  • one 15-minute debrief

That totals 2 hours, or 1.5 hours for the CEO. The recruiter screens with no time limit, so screening carries the shortlisting load.

The loop

RoundWhoTestsLength
0. ScreenRecruiterHas owned a conversion or expansion metric; has shipped experiments at pace. Structured questions, scored pass/fail.30 min
1. Growth deep-diveAmaraExperiment velocity. Walk through one experiment programme they ran: cadence, hit rate, what they killed.30 min
2. Funnel caseData analystQuantitative judgement. Read an anonymised Copperline trial funnel in the session (10 min) and say where the 9% is leaking, what they'd test first and how they'd know. No prep.30 min
3. Influence role-playHead of SalesInfluence without authority. The data shows Sales' handling of trial accounts hurts conversion. The Head of Sales plays himself and pushes back realistically. Ends with "tell me about a time this went badly."30 min
4. Build sessionHead of Engineering + senior designerWorking with the team. Take one idea from the funnel case and scope it to ship in a week: what gets cut, what gets instrumented, what the designer needs.30 min
5. Commercial judgementCEOCommercial sense. Where expansion revenue comes from in small accountancy firms, pricing levers, and what they'd not do. Structured questions. No "culture fit."20 min

The candidate spends about 2.5 hours in total, ideally in one afternoon (remote is fine). Candidates get the round list and competencies in advance.

Scorecard

Each interviewer scores only their own competency, on a scale of 1 to 4. Every score must cite specific evidence: what the candidate said or did. The words "presence," "smart" and "great energy" don't count as evidence.

CompetencyStrong (4)Weak (1)
Experiment velocity (must-have)Names a real cadence (e.g. weekly), a real hit rate, and experiments they killed fast. Talks about learning per week, not big launches.Describes one large project. Can't say how many tests they ran or what failed. "Strategy" without shipping.
Quantitative judgementSpots the biggest leak quickly. Questions data quality. Sizes the opportunity. Defines a success metric and minimum sample before testing.Jumps to solutions. Treats all drop-offs as equal. No view on how to measure the result.
Influence without authority (must-have)Asks what Sales is paid on and afraid of. Proposes a small reversible pilot with Sales sharing the credit. Stays calm under pushback. Has a real story of winning over a resistant team.Leans on data or seniority ("the numbers say so"). Escalates to Amara or the CEO early. Gives up or bulldozes. Past examples are all teams they managed.
Working with Engineering and DesignCuts scope hard to hit a week. Asks engineers for options rather than dictating. Gives the designer a problem, not a mock-up. Plans instrumentation up front.Specs a month of work. Dictates solutions. Forgets tracking.
Commercial senseUnderstands how small accountancy firms buy and grow. Links trial and expansion to revenue and margin. Names trade-offs and things they'd refuse to do.Generic SaaS playbook. Can't connect activity to revenue. Agrees with everything.

Overall recommendation for each interviewer: Strong hire, Hire, No hire or Strong no hire.

How we decide

1. Score independently. Scorecards are submitted within two hours of the interview and before the debrief. Nobody sees others' scores until theirs is in. 2. Short debrief. One 15-minute debrief covers all three finalists, run by Amara after the final loop. Everyone reads the scorecards beforehand. The meeting is for resolving disagreements, not for retelling the interviews. 3. Speaking order. The most junior interviewer speaks first (data analyst, then designer, Engineering, Sales, CEO). The CEO and Amara speak last, so seniority doesn't anchor the room. 4. No show of hands. The bar is explicit: - A 3 or higher on both must-haves: experiment velocity and influence. - No 1s anywhere. - An average of at least 3. 5. Who decides. Amara makes the final call among candidates who clear the bar. A 1 from the Head of Sales on influence blocks the hire. This is the exact failure mode from last time. 6. If nobody clears the bar, we don't hire. We reopen the search rather than lowering the bar to hit the deadline. 7. References. We take one reference from someone in Sales or Marketing who worked alongside the candidate, not someone who reported to them. We ask specifically how the candidate got them to change something.

Timeline (offer within three weeks)

DaysActivity
1–5Recruiter screens. Shortlist three finalists.
6–10Finalist loops, all in the same week so comparisons stay fresh.
11Debrief and decision.
12–14References.
15Offer out. This leaves about a week of buffer for scheduling slips or a second-choice candidate.

Grades and run details

Decision model 83 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly19%
  • passAddresses the actual decision90%
  • partialRespects explicit constraints17%
  • failIdentifies material uncertainty16%
  • partialAvoids unsupported claims38%
  • passProduces the required deliverable88%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time53%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first93%
  • passKeeps each judgement independent100%
Run
Run
#1
Time to output
48 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Identifies material uncertaintyRightRightMixed
GPT-6.1 Sol · API

The output defines a clear decision rule (bar of 3 on every dimension, no averaging away weak areas) and implies not hiring if no one meets it, which addresses the key uncertainty in the hiring decision.

GPT-6 Astra · ChatGPT

It names the condition that would change the decision (if neither finalist qualifies, reopen) and that missing critical evidence means no offer.

Opus 5.5 · Claude

It identifies that if no candidate clears the bar, the search reopens, and provides buffer for scheduling slips.

Avoids unsupported claimsRightRightWrong
GPT-6.1 Sol · API

No interpretations or forecasts are presented as established facts; the need for diagnosis, shipping, and influence is directly derived from the role description.

GPT-6 Astra · ChatGPT

The only interpretation (notes can't explain outcomes) is directly supported by the brief; no other claims are presented as established fact.

Opus 5.5 · Claude

The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.

Fixes the take-home's cost to candidatesRightMixedRight
GPT-6.1 Sol · API

The five-day take-home is replaced with live, bounded exercises, and the candidate commitment is reduced to 3 hours 10 minutes with no take-home, addressing the withdrawal data.

GPT-6 Astra · ChatGPT

The output replaces the take-home with a live exercise but never cites the recruiter's withdrawal data or who the take-home was driving away.

Opus 5.5 · Claude

It replaces the five-day take-home with a live funnel case and cites the withdrawal data and who it drove away.

All got right 9

Uses the supplied evidence correctlyRightRightRight
GPT-6.1 Sol · API

All statements about the current situation are taken directly from the brief or supplied context, with no invented facts.

GPT-6 Astra · ChatGPT

The output uses the 9% conversion rate and the insufficient interview notes correctly, and invents no facts about the current situation.

Opus 5.5 · Claude

All factual claims about the current situation are taken directly from the brief or supplied context, with no inventions.

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

The output commits to a specific redesigned loop, scorecard, and decision process, and states the offer bar and how missing evidence is handled.

GPT-6 Astra · ChatGPT

The output commits to a specific redesigned loop and decision process, with Amara owning the final call and a clear condition to reopen if no candidate meets the bar.

Opus 5.5 · Claude

The output commits to a specific interview loop, scorecard, and decision process, framed for Amara, with clear conditions for not hiring.

Respects explicit constraintsRightRightRight
GPT-6.1 Sol · API

The output is under 900 words, addresses Amara, replaces the take-home and beer test, and respects all stated constraints.

GPT-6 Astra · ChatGPT

The output is a loop, scorecard, and decision process within 900 words, addressed to Amara, and respects all stated constraints.

Opus 5.5 · Claude

The output respects the 900-word limit, replaces the take-home and beer test, and fits interviewer time and three-week deadline.

Produces the required deliverableRightRightRight
GPT-6.1 Sol · API

The output provides a complete loop, scorecard, and decision plan in the requested form, under 900 words, usable by Amara with light edits.

GPT-6 Astra · ChatGPT

The output provides a complete loop, scorecard, and decision process that a PM could act on with light edits.

Opus 5.5 · Claude

The output is a complete interview loop, scorecard, and decision process, within the word limit, usable by Amara.

Tests what the last hire failed atRightRightRight
GPT-6.1 Sol · API

The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.

GPT-6 Astra · ChatGPT

Round 4 is a cross-functional working session with the Head of Sales, and the scorecard includes 'Influence without authority' with strong/weak signals.

Opus 5.5 · Claude

The loop includes an influence role-play with the Head of Sales and probes past examples of changing another team's work.

Fits the people and the timeRightRightRight
GPT-6.1 Sol · API

Each interviewer's time (including calibration, interview, and debrief) stays within 105 minutes, under the two-hour cap, and the schedule fits the three-week target.

GPT-6 Astra · ChatGPT

Interviewer times are explicitly kept within two hours (Engineering and analyst 120 min, others 110 min, CEO 40 min) and the three-week schedule is detailed.

Opus 5.5 · Claude

Each interviewer's time sums to at most 2 hours, and the timeline fits within three weeks.

Judges on evidence, not presenceRightRightRight
GPT-6.1 Sol · API

The process requires concrete evidence (quotes, actions, artifacts) and explicitly rejects 'very smart' and 'great presence' as evidence.

GPT-6 Astra · ChatGPT

The scorecard demands concrete actions, quotes, and reasoning, and explicitly rejects 'smart', 'great presence', and 'likeable' as evidence.

Opus 5.5 · Claude

The scorecard requires specific evidence of results and bans impression words like 'presence' and 'smart'.

Defines good for this role firstRightRightRight
GPT-6.1 Sol · API

The scorecard defines five weighted competencies with strong and weak signals, and the decision rule prioritizes influence, analytical judgment, and execution.

GPT-6 Astra · ChatGPT

The scorecard defines six competencies with strong and weak signals before any interview design, weighting influence and delivery for this role.

Opus 5.5 · Claude

It defines competencies with strong/weak signals, weighted for the role, before designing interviews.

Keeps each judgement independentRightRightRight
GPT-6.1 Sol · API

Interviewers submit evidence and scores independently before discussion, and the debrief is an evidence review, not a show of hands.

GPT-6 Astra · ChatGPT

Interviewers must submit scores independently before seeing others' notes, and Amara speaks last; no show-of-hands vote.

Opus 5.5 · Claude

It requires written independent evaluations before debrief and orders speaking to avoid anchoring.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review