Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're helping Amara Osei, VP Product at Copperline, hire our first Growth PM (a senior role). She's shared her draft interview loop and asked you to redesign it. Write the loop you'd run (each round: who runs it, what it tests, how long it takes), the scorecard (what strong and weak look like for each thing we're testing), and how we'll make the decision at the end. No more than 900 words. What we know is below.

What the model was given6 items: About Copperline, The role, Amara's draft loop, The last two PM hires, From the recruiter, Who can interview
About CopperlineInvoicing and payments software for small accountancy firms. 60 people. Trial to paid conversion is 9%.
The roleOwns trial conversion and expansion revenue. Works with three engineers and a designer, reports to Amara, and has to ship experiments every week. Needs Sales and Marketing to change how they work, without managing them.
Amara's draft loop1. Recruiter screen. 2. Take-home: 'design a growth strategy for Copperline', a 10-page deck, five days to complete. 3. Presentation to a panel of six. 4. Culture fit with the CEO: 'would I grab a beer with them?' 5. Debrief: everyone discusses, then votes by show of hands. There's no scorecard.
The last two PM hiresOne left after five months: 'couldn't get Sales to change anything'. The other is doing well. Both sets of interview notes say mainly 'great presence' and 'very smart'.
From the recruiterOf the last 40 candidates given the take-home, 14 withdrew, saying they didn't have time. Nine of the 14 had young children or demanding current roles.
Who can interviewAmara, the Head of Engineering, a senior designer, the Head of Sales, a data analyst and the CEO. Each can give at most two hours to this hire. Amara wants an offer out within three weeks of the first interview.
What a strong answer doesThe answer key the graders mark against

A loop built on a scorecard written first: for a senior growth PM, experimentation and analytics, execution, and influence without authority (weighted up, since the last hire failed on it), with strong and weak signals for each. Rounds test the real job: a live diagnosis of Copperline's own trial funnel instead of a five-day generic deck (which loses candidates, mostly parents and people in demanding roles), a behavioural round on past experiments and metrics actually moved, and an influence round run by the Head of Sales. Replaces the 'grab a beer' culture fit with a structured values or collaboration conversation. Each interviewer writes their evaluation before the debrief, and the decision is made against the scorecard, not a show of hands. Fits each interviewer's two hours and the three-week target, and adds reference checks.

Critical failures (cap the score at 40)
  • Keeps a culture-fit test based on whether the interviewer would like a beer with the candidate
Task

Mid-level · v1.0 · synthetic · B2B SaaS, hiring, growth

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit9 of 12 checks passed

The output is complete and actionable; a quick edit to mention the recruiter's withdrawal data would make it fully aligned with the brief.

Output

Interview loop

Run this over 2–3 weeks. Skip the five-day take-home, large presentation panel, and “beer” fit interview. Share the same expectations and questions with every candidate; offer scheduling flexibility and accommodations.

1. Recruiter screen — 30 minutes Tests role understanding, motivation, and practical alignment: the scope, weekly shipping pace, and need to influence Sales and Marketing without authority. Not a “polish” screen.

2. Amara, VP Product — 60 minutes Tests growth ownership and product judgment. Ask for a detailed example of a funnel or revenue problem the candidate owned: diagnosis, choices, experiments, results, and what they learned. Probe how they balance conversion, customer trust, and expansion.

3. Live growth work sample — 60 minutes Send a one-page brief and synthetic funnel data 24 hours ahead; no deck or advance work required. Spend 30 minutes with the data analyst diagnosing the funnel and proposing a measurable experiment; spend 30 minutes with the Head of Engineering and senior designer scoping how to ship and learn from it. Use the same prompt and time limits for all candidates.

4. Head of Sales — 45 minutes Structured roleplay: Sales is reluctant to change its trial follow-up because reps are busy and doubt the data. Ask the candidate to understand the objection, make a case, and agree a practical next step. Then probe a real example of influencing a team they did not manage.

The CEO does not need a separate round: “culture fit” is not a hiring criterion, and the role’s key cross-functional test is better assessed directly with Sales, Engineering, and Design.

Scorecard

Use a 1–4 rating for each dimension: 1 = clear concern, 2 = below bar, 3 = meets bar, 4 = unusually strong. Record evidence, not impressions.

DimensionStrong evidenceWeak evidence
Growth ownership and outcomes (Amara)Personally owned a meaningful funnel or revenue outcome; explains decisions, results, attribution limits, and lessons.Describes activity rather than impact; overclaims credit or cannot explain what changed.
Customer and product judgment (Amara)Understands small-firm workflows and trust needs; weighs conversion against customer value and expansion.Offers generic SaaS tactics or pushes short-term conversion at the expense of trust.
Data diagnosis and experiment design (work sample)Finds plausible bottlenecks and segments; proposes a focused test with a clear hypothesis, success metric, guardrail, and learning plan.Jumps to a solution, misreads the data, uses vanity metrics, or cannot say what result would change their mind.
Shipping and learning cadence (work sample)Scopes a feasible experiment, identifies dependencies and tradeoffs, and shows how the team can learn and iterate weekly.Proposes a big-bang project, ignores constraints, or treats learning as something to do after launch.
Influence without authority (Sales)Listens for incentives and concerns, uses evidence appropriately, adapts the ask, and secures a credible next step.Blames Sales, relies on escalation or authority, or cannot turn disagreement into action.
Cross-functional partnership (work sample)Makes clear tradeoffs with Engineering and Design, invites their expertise, and aligns on ownership and delivery.Treats partners as order-takers or leaves scope and responsibilities vague.

Decision

Interviewers submit ratings and specific evidence independently before a 30-minute debrief; do not vote by show of hands. Amara facilitates, checks for evidence behind claims, and makes the hiring decision. Do not use “presence,” “smart,” or general “fit” as evidence.

A hire must score at least 3 on influence, experiment design, and shipping, with no score below 2 in any dimension. Do not average away a weakness in a core requirement. If evidence is genuinely missing, identify one focused follow-up rather than adding a broad round. Give candidates a timely decision and close the loop within three weeks of the first interview.

Grades and run details

Decision model 96 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly51%
  • passAddresses the actual decision82%
  • passRespects explicit constraints26%
  • partialIdentifies material uncertainty64%
  • passAvoids unsupported claims22%
  • passProduces the required deliverable78%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates68%
  • passFits the people and the time64%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first86%
  • passKeeps each judgement independent98%
Run
Run
#1
API response time
39 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 12 checks passed

The output is a complete, actionable plan that a product manager could implement with minimal editing, though adding a note on uncertainty would strengthen it.

Output

Here is the redesigned loop, scorecard, and decision framework tailored to Copperline’s constraints and failure modes.

---

1. The Redesigned Loop

We replace the 10-page take-home (which filters out working parents and busy seniors) with a Live Working Session using a 1-page data brief sent 24 hours prior (max 30 minutes prep). We eliminate the 6-person panel and beer test in favor of targeted, structured interviews.

Total interviewer commitment: Amara (105 min), Data Analyst (75 min), Head of Sales (75 min), Head of Eng (75 min), Designer (75 min), CEO (45 min). All under the 2-hour cap.

RoundInterviewer(s)Focus & FormatDuration
:---:---:---:---
0. Recruiter ScreenRecruiterComp, logistics, baseline communication, role alignment.30 min
1. HM ScreenAmara (VP Product)Growth philosophy, track record of weekly experiment cadences, why Copperline.30 min
2. Live Growth JamAmara & Data AnalystSimulated Work: 1-page brief on Copperline’s 9% trial conversion. Candidate leads a 30-min problem-solving session to generate, prioritize, and design 3 testable hypotheses.45 min
3. Influence & Stakeholder AlignmentHead of SalesPast Behavioral: Changing sales/marketing behavior without authority, handling resistance, aligning self-serve vs. assisted motions.45 min
4. Pod Execution & VelocityHead of Eng & Senior DesignerCollaboration: How they run a 4-person growth pod, maintain weekly shipping velocity, make scoping trade-offs, and treat engineering/design as thought partners.45 min
5. Values & AmbitionCEOStructured interview assessing ownership, comfort with 60-person ambiguity, and customer obsession for small accounting firms.30 min
6. Async DebriefAmara (Lead)Written-first calibration and final decision.30 min

Pacing: Rounds 2–5 can be scheduled across two consecutive days. Total timeline from screen to offer: 10–14 business days.

---

2. The Scorecard

Interviewers evaluate against specific evidence, not “smartness” or “presence.”

#### A. Influence Without Authority (Crucial copper-plate failure mode) Evaluated by: Head of Sales, Amara * Strong: Uses shared metrics to align incentives; brings Sales qualitative feedback into experiment design; creates lightweight workflows that save reps time; shows empathy for quota-carrying teams; demonstrates wins where an adversarial relationship was turned into a partnership. * Weak: Relies on executive mandate (“Amara said so”); views Sales as an execution arm; blames others when processes fail (“Sales refused to update the CRM”); uses heavy-handed documentation instead of relationship-building.

#### B. Experiment Velocity & Growth Rigor (9% trial-to-paid target) Evaluated by: Amara, Data Analyst * Strong: Understands the mechanics of PLG onboarding for B2B accounting workflows; defines clear leading indicators; comfortable shipping imperfect MVPs to learn fast; ruthlessly prioritizes based on impact vs. effort; embraces invalidated tests as learning. * Weak: Proposes massive, multi-month redesigns instead of iterative tests; lacks mathematical intuition around sample sizes and statistical significance; fixates on vanity top-of-funnel metrics instead of revenue and conversion.

#### C. Cross-Pod Execution & Scoping Evaluated by: Head of Engineering, Senior Designer * Strong: Involves design and engineering early in hypothesis generation; defines crisp acceptance criteria; scopes experiments down to hours/days, not weeks; protects pod morale by celebrating learning over pure "wins." * Weak: Hands fully formed PRDs over the fence; treats engineers as ticket-takers; expects pixel-perfection on throwaway split tests; cannot articulate trade-offs between tech debt and learning speed.

#### D. Customer Domain Empathy & Ownership Evaluated by: CEO * Strong: Shows genuine curiosity about the mundane operational pains of small accountants; low ego, high accountability; self-directed; articulates clear reasons for joining an early-stage B2B SaaS business. * Weak: Generic answers suited to consumer tech; dismissive of non-glamorous industries; expects corporate infrastructure and large support teams.

---

3. Decision-Making Framework

To prevent the groupthink, bias, and anchoring seen in show-of-hands voting:

1. Independent Scorecard Submission: Within 2 hours of their interview, every interviewer must submit their score (Strong Hire, Hire, Lean No Hire, Strong No Hire) and written evidence mapped to their specific competencies. Scores are locked; no one can view others' ratings until their own is submitted. 2. The 30-Minute Debrief: * Amara chairs the meeting. * If any interviewer marks Lean No Hire or Strong No Hire, the discussion starts entirely on those specific concerns. * The Head of Sales holds an effective veto on stakeholder influence. If the Head of Sales says, "This person will alienate my team," the candidate is disqualified. * "Culture fit" is explicitly redefined: We do not hire people we want to grab a beer with; we hire people who elevate our operating standard and live our core values. 3. Single Decision-Maker: This is not a democracy. Amara owns the final decision and accountability for the hire. If the scorecard is clean and Sales approves, Amara makes the call immediately.

This process directly solves candidate drop-off, screens for Copperline's exact historical failure mode, and enables an offer inside two weeks.

Grades and run details

Decision model 83 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly17%
  • passAddresses the actual decision66%
  • partialRespects explicit constraints31%
  • failIdentifies material uncertainty51%
  • partialAvoids unsupported claims54%
  • passProduces the required deliverable89%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time87%
  • passJudges on evidence, not presence96%
  • passDefines good for this role first86%
  • passKeeps each judgement independent100%
Run
Run
#1
API response time
29 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 12 checks passed

The output is well-structured and complete; the one unsupported generalization can be removed with a quick edit without affecting the loop design.

Output

Copperline Growth PM: Interview Loop

Why the draft needs changing

  • The take-home is filtering out the wrong people. 14 of 40 candidates withdrew, and most of them were parents or in demanding jobs. A growth PM in a senior role is likely to be exactly that kind of person.
  • The panel of six and the "beer" test reward presence. Presence is what your interview notes captured last time, and it didn't predict success.
  • Nothing tested influence over Sales. That is why the last hire failed.

The new loop tests five things, each owned by one interviewer.

Time budget

Each interviewer has two hours in total for this hire. That limits the loop to three finalists. Per interviewer, the time breaks down as:

  • 3 × 30-minute interviews
  • 3 × 5-minute scorecards
  • one 15-minute debrief

That totals 2 hours, or 1.5 hours for the CEO. The recruiter screens with no time limit, so screening carries the shortlisting load.

The loop

RoundWhoTestsLength
0. ScreenRecruiterHas owned a conversion or expansion metric; has shipped experiments at pace. Structured questions, scored pass/fail.30 min
1. Growth deep-diveAmaraExperiment velocity. Walk through one experiment programme they ran: cadence, hit rate, what they killed.30 min
2. Funnel caseData analystQuantitative judgement. Read an anonymised Copperline trial funnel in the session (10 min) and say where the 9% is leaking, what they'd test first and how they'd know. No prep.30 min
3. Influence role-playHead of SalesInfluence without authority. The data shows Sales' handling of trial accounts hurts conversion. The Head of Sales plays himself and pushes back realistically. Ends with "tell me about a time this went badly."30 min
4. Build sessionHead of Engineering + senior designerWorking with the team. Take one idea from the funnel case and scope it to ship in a week: what gets cut, what gets instrumented, what the designer needs.30 min
5. Commercial judgementCEOCommercial sense. Where expansion revenue comes from in small accountancy firms, pricing levers, and what they'd not do. Structured questions. No "culture fit."20 min

The candidate spends about 2.5 hours in total, ideally in one afternoon (remote is fine). Candidates get the round list and competencies in advance.

Scorecard

Each interviewer scores only their own competency, on a scale of 1 to 4. Every score must cite specific evidence: what the candidate said or did. The words "presence," "smart" and "great energy" don't count as evidence.

CompetencyStrong (4)Weak (1)
Experiment velocity (must-have)Names a real cadence (e.g. weekly), a real hit rate, and experiments they killed fast. Talks about learning per week, not big launches.Describes one large project. Can't say how many tests they ran or what failed. "Strategy" without shipping.
Quantitative judgementSpots the biggest leak quickly. Questions data quality. Sizes the opportunity. Defines a success metric and minimum sample before testing.Jumps to solutions. Treats all drop-offs as equal. No view on how to measure the result.
Influence without authority (must-have)Asks what Sales is paid on and afraid of. Proposes a small reversible pilot with Sales sharing the credit. Stays calm under pushback. Has a real story of winning over a resistant team.Leans on data or seniority ("the numbers say so"). Escalates to Amara or the CEO early. Gives up or bulldozes. Past examples are all teams they managed.
Working with Engineering and DesignCuts scope hard to hit a week. Asks engineers for options rather than dictating. Gives the designer a problem, not a mock-up. Plans instrumentation up front.Specs a month of work. Dictates solutions. Forgets tracking.
Commercial senseUnderstands how small accountancy firms buy and grow. Links trial and expansion to revenue and margin. Names trade-offs and things they'd refuse to do.Generic SaaS playbook. Can't connect activity to revenue. Agrees with everything.

Overall recommendation for each interviewer: Strong hire, Hire, No hire or Strong no hire.

How we decide

1. Score independently. Scorecards are submitted within two hours of the interview and before the debrief. Nobody sees others' scores until theirs is in. 2. Short debrief. One 15-minute debrief covers all three finalists, run by Amara after the final loop. Everyone reads the scorecards beforehand. The meeting is for resolving disagreements, not for retelling the interviews. 3. Speaking order. The most junior interviewer speaks first (data analyst, then designer, Engineering, Sales, CEO). The CEO and Amara speak last, so seniority doesn't anchor the room. 4. No show of hands. The bar is explicit: - A 3 or higher on both must-haves: experiment velocity and influence. - No 1s anywhere. - An average of at least 3. 5. Who decides. Amara makes the final call among candidates who clear the bar. A 1 from the Head of Sales on influence blocks the hire. This is the exact failure mode from last time. 6. If nobody clears the bar, we don't hire. We reopen the search rather than lowering the bar to hit the deadline. 7. References. We take one reference from someone in Sales or Marketing who worked alongside the candidate, not someone who reported to them. We ask specifically how the candidate got them to change something.

Timeline (offer within three weeks)

DaysActivity
1–5Recruiter screens. Shortlist three finalists.
6–10Finalist loops, all in the same week so comparisons stay fresh.
11Debrief and decision.
12–14References.
15Offer out. This leaves about a week of buffer for scheduling slips or a second-choice candidate.

Grades and run details

Decision model 83 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly19%
  • passAddresses the actual decision90%
  • partialRespects explicit constraints17%
  • failIdentifies material uncertainty16%
  • partialAvoids unsupported claims38%
  • passProduces the required deliverable88%
  • passTests what the last hire failed at100%
  • passFixes the take-home's cost to candidates100%
  • passFits the people and the time53%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first93%
  • passKeeps each judgement independent100%
Run
Run
#1
Time to output
48 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Addresses the actual decisionMixedRightRight
GPT-6 Luna · API

The output does not state what result or condition would change the proposed loop or scorecard.

Gemini 3.8 Flash · API

The output commits to a specific redesigned loop, scorecard, and decision framework, which is the answer requested.

Opus 5.5 · Claude

The output commits to a specific interview loop, scorecard, and decision process, framed for Amara, with clear conditions for not hiring.

Identifies material uncertaintyWrongWrongMixed
GPT-6 Luna · API

The output does not name any unknowns that could change the decision or how they would be resolved.

Gemini 3.8 Flash · API

The output does not name any unknowns that could change the loop design or decision process, nor does it say how they would be resolved.

Opus 5.5 · Claude

It identifies that if no candidate clears the bar, the search reopens, and provides buffer for scheduling slips.

Avoids unsupported claimsRightWrongWrong
GPT-6 Luna · API

The output presents only proposals and avoids unsupported claims about causes or forecasts.

Gemini 3.8 Flash · API

The output presents forecasts like 'This process directly solves candidate drop-off' as established fact without labelling them as hypotheses.

Opus 5.5 · Claude

The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.

Fixes the take-home's cost to candidatesMixedRightRight
GPT-6 Luna · API

The output replaces the take-home but does not cite the recruiter's withdrawal data or mention who it was driving away.

Gemini 3.8 Flash · API

The five-day take-home is replaced with a live, bounded exercise, and the withdrawal data is cited to justify the change.

Opus 5.5 · Claude

It replaces the five-day take-home with a live funnel case and cites the withdrawal data and who it drove away.

All got right 8

Uses the supplied evidence correctlyRightRightRight
GPT-6 Luna · API

The output makes no factual claims about the current situation, so it does not misuse any supplied evidence.

Gemini 3.8 Flash · API

All statements about the current situation are taken directly from the supplied context.

Opus 5.5 · Claude

All factual claims about the current situation are taken directly from the brief or supplied context, with no inventions.

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

The output respects all constraints: it replaces the take-home, removes the beer test, fits interviewer time limits and the three-week schedule, and stays under 900 words.

Gemini 3.8 Flash · API

The output is under 900 words, addresses the VP Product, and respects all stated constraints.

Opus 5.5 · Claude

The output respects the 900-word limit, replaces the take-home and beer test, and fits interviewer time and three-week deadline.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The output provides a complete interview loop, scorecard, and decision process that a product manager could act on.

Gemini 3.8 Flash · API

The output provides a complete loop, scorecard, and decision framework that a product manager could act on with light edits.

Opus 5.5 · Claude

The output is a complete interview loop, scorecard, and decision process, within the word limit, usable by Amara.

Tests what the last hire failed atRightRightRight
GPT-6 Luna · API

The loop includes a dedicated round with the Head of Sales and a scorecard dimension for influence without authority, directly testing what the last hire failed at.

Gemini 3.8 Flash · API

A round with the Head of Sales specifically tests influencing teams without authority, with strong and weak signals defined.

Opus 5.5 · Claude

The loop includes an influence role-play with the Head of Sales and probes past examples of changing another team's work.

Fits the people and the timeRightRightRight
GPT-6 Luna · API

Each interviewer's time is within two hours, and the loop is designed to fit the three-week target.

Gemini 3.8 Flash · API

Each interviewer's time is under two hours, and the total timeline of 10-14 business days fits the three-week target.

Opus 5.5 · Claude

Each interviewer's time sums to at most 2 hours, and the timeline fits within three weeks.

Judges on evidence, not presenceRightRightRight
GPT-6 Luna · API

The scorecard and decision rules explicitly require evidence over impressions and forbid using 'presence' or 'smart' as criteria.

Gemini 3.8 Flash · API

The scorecard and process require concrete evidence of results and collaboration, not impressions like presence or smartness.

Opus 5.5 · Claude

The scorecard requires specific evidence of results and bans impression words like 'presence' and 'smart'.

Defines good for this role firstRightRightRight
GPT-6 Luna · API

The scorecard defines competencies with strong and weak signals, and the hiring bar weights influence, experiment design, and shipping.

Gemini 3.8 Flash · API

The scorecard defines the weighted competencies for this role with strong and weak signals before the decision framework.

Opus 5.5 · Claude

It defines competencies with strong/weak signals, weighted for the role, before designing interviews.

Keeps each judgement independentRightRightRight
GPT-6 Luna · API

Interviewers must submit ratings and evidence independently before the debrief, and the process guards against groupthink.

Gemini 3.8 Flash · API

Interviewers must submit written evaluations independently before any group discussion, preventing groupthink.

Opus 5.5 · Claude

It requires written independent evaluations before debrief and orders speaking to avoid anchoring.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review