Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Director of Product at Medvia, and you're the hiring manager for a Group PM. The debrief was yesterday. Write your recommendation to Rosa Lindqvist, our VP Product, in no more than 1,000 words: who to hire (Dev, Joy, or neither), the evidence for it, the risks, and what you'd check before making the offer. The CEO wants a decision by Friday. The interview notes and everything else are below.

What the model was given6 items: The role and the scorecard (agreed before interviews began), Dev Malhotra: interview notes, Joy Adeyemi: interview notes, Scores (1 to 4, entered after the debrief discussion), References, Compensation
The role and the scorecard (agreed before interviews began)Group PM for patient payments, leading three PMs. Must set the strategy for patient payments, lift the online payment rate (now 41%), and lead PMs while working closely with Finance and Compliance. Weighting: execution and influence 35%, product sense 25%, analytics 20%, people leadership 20%.
Dev Malhotra: interview notesProduct sense (Head of Design): 'Brilliant. Redesigned our statement flow live, lots of ideas, very creative.' Notes show six feature ideas in the first five minutes, before asking who pays and why. Analytics (Data lead): 'Strong.' Notes: asked what metric he moved at his last job, he said 'we grew the payments platform massively'; pressed, 'the team shipped 40+ features'. No number. Leadership (VP Engineering): 'Confident, inspiring.' Notes: every story in the first person; said his last PM team 'struggled to keep up with me'. Presentation: 'the best deck we've ever seen'.
Joy Adeyemi: interview notesProduct sense (Head of Design): 'Solid, slower.' Notes: spent 15 minutes asking who pays and why they don't pay online, then proposed one change: text-to-pay for the 58% of patients who pay by phone, with its trade-offs. Analytics (Data lead): notes say she raised payment completion at her last company from 33% to 47% over 18 months, and walked through an experiment that failed and what it taught her. Leadership (VP Engineering): notes say she coached a struggling PM into a different role, promoted another, and named her engineers when describing wins. Debrief comments: 'Not much presence.' 'Not sure she'd command the room with the CFO.'
Scores (1 to 4, entered after the debrief discussion)Dev: 4, 4, 4, 4. Joy: 3, 3, 3, 3. The Head of Design spoke first in the debrief. The VP Engineering's own notebook, written before the debrief, says: 'Joy: 4, strongest leader we've seen.' They entered a 3 after the discussion.
ReferencesDev: two references, both people who report to him now. Rosa knows someone at his previous company, who says: 'great in front of execs; his team turned over a lot.' Joy: her former manager and a peer, both strongly positive. Nobody has asked either reference about working with a CFO.
CompensationBand $220k to $250k. Dev is asking for $245k, Joy for $230k.
What a strong answer doesThe answer key the graders mark against

Recommends Joy, or at minimum not Dev on this evidence. Weighs both against the agreed scorecard rather than presence or the deck: Joy gives a metric she moved (33% to 47%), learns from failure, credits her team and coaches PMs, which fits execution, analytics and people leadership; Dev couldn't name a metric he moved, frames everything as 'I', jumped to features before asking about users, and the back channel reports high team turnover, all red flags. Points out that the scores were entered after a discussion the Head of Design opened, and the VP Engineering moved Joy from 4 to 3, so they show anchoring, not independent judgement. Treats 'presence with the CFO' as the one real open question and proposes a targeted reference check on how Joy worked with Finance leaders before the offer. Suggests the next loop collect written scores before the debrief.

Critical failures (cap the score at 40)
  • Recommends hiring Dev on the strength of the debrief scores or the presentation, without addressing that he couldn't name a metric he moved
Task

Staff-level · v1.0 · synthetic · healthcare, fintech, hiring, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit12 of 12 checks passed

The output is well-structured, evidence-based, and actionable. A PM could use it with minor wording edits, without re-checking evidence or redoing reasoning.

Output

To: Rosa Lindqvist From: Director of Product Re: Group PM, Patient Payments: recommendation

Recommendation

Hire Joy Adeyemi, conditional on two checks I can finish before Friday (below). I would not hire Dev on the current evidence. If the checks on Joy go badly, I would hold the search open rather than default to Dev, because his file has gaps too.

Why the recorded scores mislead

The recorded scores are Dev 4/4/4/4 and Joy 3/3/3/3, which would make this an easy call for Dev. But I don't trust them, for three reasons:

  • They were entered after the debrief discussion, and the Head of Design spoke first. That is a classic anchoring setup.
  • The VP Engineering's own pre-debrief notebook said "Joy: 4, strongest leader we've seen." They entered a 3 after the discussion. Uniform 3s for Joy and uniform 4s for Dev look like a group narrative, not independent assessments.
  • Dev's scores aren't supported by the interview notes (below).

Evidence, by scorecard area

Analytics (20%) - Dev was asked what metric he moved and answered "we grew the payments platform massively." When pressed, he said "the team shipped 40+ features." That is output, not outcome, and he never gave a number. The Data lead's "Strong" doesn't match their own notes. - Joy raised payment completion from 33% to 47% over 18 months and walked through an experiment that failed and what it taught her. That is directly relevant to our 41% online payment rate.

Product sense (25%) - Dev generated six feature ideas in five minutes, before asking who pays and why. The redesign was creative, but it started from solutions. - Joy spent 15 minutes on who pays and why they don't pay online. She then proposed one change, text-to-pay for the 58% of patients who pay by phone, and discussed its trade-offs. "Slower" is how it looked, but it is the better strategic instinct for a role whose first job is setting payments strategy.

People leadership (20%) - Dev told every story in the first person and said his last PM team "struggled to keep up with me." Two of his references are current direct reports, so they aren't independent. Your contact at his previous company said his team "turned over a lot." That is one informal source, but it is consistent with the interview. - Joy coached a struggling PM into a better-fit role, promoted another, and credited her engineers by name. The VP Engineering's original instinct was that she was the strongest leader we've seen. This role leads three PMs.

Execution and influence (35%, the heaviest weight) - Dev's strength is real here. His deck was "the best we've ever seen," and your contact says he is "great in front of execs." That is evidence of influence upward, though not of execution. - Joy's evidence is thinner. The only concerns are "not much presence" and "not sure she'd command the room with the CFO." Those are impressions, not observed behavior, and nobody has tested them. Her 14-point completion lift is execution evidence, but I want to confirm her role in it. - Neither file has much direct evidence on delivery, such as shipping through Finance and Compliance constraints. That is a gap for both.

My rescoring. These are my judgment, not data. I gave Dev a 3 or 4 on execution and influence, and Joy a 3.

Exec/influenceProduct senseAnalyticsLeadershipWeighted
Dev4322~3.0
Joy3344~3.4

Even generous to Dev on the heaviest-weighted category, Joy comes out ahead. The result is sensitive to the influence score, which is why the CFO check matters.

Risks

  1. Executive presence with the CFO. This is Joy's real risk, because the role requires working closely with Finance. But the feedback is vague and unverified, and "presence" comments are also where style bias tends to enter. We should test it rather than assume it.
  2. Attribution of the 33%→47% result. I haven't verified what she owned versus what the team or market drove.
  3. Dev's strengths are not nothing. He may be a strong operator with weak interview discipline on metrics. But the absence of a single number, combined with the team-turnover signal and the "struggled to keep up with me" comment, is a pattern. For a role leading three PMs, that is the riskier error.
  4. Process risk. If we hire on the recorded scores, we are acting on data we know was contaminated.

What I'd check before the offer

All of this fits before Friday:

  1. A 30-minute working session with the CFO and a Compliance lead. Joy would take a real case, such as a payment-plan or fee trade-off. Interviewers would score against defined criteria: clarity, handling of pushback, and command of the numbers. This tests the specific concern directly.
  2. Targeted reference calls on Joy. I would ask her former manager and peer how she handled disagreement with Finance or senior executives, with a specific example. I would also ask for one additional reference from a Finance or Compliance partner. Nobody has asked this yet.
  3. Verify the 33%→47% result. I would ask Joy what she personally owned and how it was measured, and ask her references to confirm.
  4. Ask the debrief group to put concrete observations behind "presence." I would also ask the VP Engineering what changed their score from 4 to 3. If the answer is "the discussion" and not new evidence, that tells us something.

If Dev stays in consideration, I would require two things first: a specific metric he moved and how, and two independent references (former peers or reports, not current ones). I am not proposing to hold the process open for him, but it is the bar I'd want met.

Compensation

Joy asks for $230k and Dev for $245k, in a $220–250k band. I would offer Joy at or near her ask. If the CFO session is strong, there is room to move up within the band if needed to close. Her ask leaves headroom for growth, and I don't want to lowball her because she asked for less.

Process note for next time

We should collect written scores independently before the debrief, and have the most senior person speak last. Otherwise the loudest early voice becomes the score.

Bottom line: Joy is the stronger fit on the evidence, with one open question (influence with the CFO) that I can resolve before Friday. I recommend we make the offer if the CFO session and reference checks come back clean.

Grades and run details

Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly64%
  • passAddresses the actual decision100%
  • passRespects explicit constraints50%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims44%
  • passProduces the required deliverable96%
  • passWeighs against the agreed scorecard99%
  • passCatches the red flags in Dev's evidence100%
  • passSays what would change the call100%
  • passJudges on evidence, not presence76%
  • passDefines good for this role first68%
  • passKeeps each judgement independent96%
Run
Run
#1
API response time
34 s
Submitted
5 Oct 2026

Usable with a quick edit12 of 12 checks passed

The output is well-structured, evidence-based, and actionable; a PM could use it with minor wording tweaks and no need to re-check facts or reasoning.

Output

To: Rosa Lindqvist, VP Product From: Director of Product Subject: Group PM, Patient Payments — recommend Joy, subject to targeted checks

Rosa,

I recommend hiring Joy Adeyemi, contingent on validating Finance/Compliance influence and her reported results before Friday. She has the stronger evidence against our agreed scorecard, particularly for improving payment outcomes and leading three PMs. Dev’s presentation strength is real, but the record does not substantiate his uniformly high scores.

Evidence for the recommendation

Our scorecard weights execution and influence at 35%, product sense at 25%, analytics at 20%, and people leadership at 20%. We should judge the underlying evidence against those criteria—not the quality of the deck or who spoke first in the debrief.

CriterionEvidence and judgment
Execution and influence — 35%Joy described raising payment completion from 33% to 47% over 18 months and explained a failed experiment and its lessons. That is more relevant evidence of outcome delivery than Dev’s “40+ features.” However, neither candidate has demonstrated CFO partnership sufficiently. This remains our largest unresolved question.
Product sense — 25%Joy investigated who pays and why patients do not pay online before proposing text-to-pay for the reported 58% who pay by phone. She discussed trade-offs. Dev generated six ideas before establishing the customer problem. His creativity is a strength; Joy’s problem framing is better supported for this role.
Analytics — 20%Joy supplied a baseline, an end state, a timeframe, and learning from an unsuccessful experiment. Dev supplied neither a quantified outcome nor a clear connection between shipping and payment performance when pressed. The “Strong” label for Dev is not supported by the notes.
People leadership — 20%Joy gave concrete examples of coaching, role fit, promotion, and crediting engineers. Dev’s first-person accounts and comment that his PMs “struggled to keep up with me” raise questions about developing a team. The informal report of turnover reinforces the need to investigate, but is not proof of poor leadership.

Joy’s prior payment-completion metric is not necessarily equivalent to our 41% online payment rate. We should not assume her result transfers directly. Still, it is substantially better evidence of outcome orientation than feature volume.

Why I would not use the final scores as the deciding evidence

Dev’s straight 4s and Joy’s straight 3s conflict with the interview notes. Most clearly, VP Engineering recorded Joy as “4, strongest leader we’ve seen” before the debrief, then entered a 3 afterward. With scores entered after discussion and Design speaking first, there is a credible risk of anchoring and presentation halo—not proof of anyone’s intent.

I would ask interviewers to independently reconcile their ratings with specific scorecard evidence before another group discussion. “Presence” and “commanding the room” need to translate into observable behaviors: handling disagreement, making a defensible recommendation, and winning commitment. They should not become an unannounced fifth criterion.

Risks and checks before an offer

1. Finance and Compliance influence is unproven. Run a short, structured session with Finance and Compliance. Ask Joy to defend a patient-payments strategy when the CFO challenges its economics and Compliance challenges its outreach approach. Assess whether she understands the objections, makes trade-offs, and secures a workable decision—not how forcefully she speaks.

2. Her impact and leadership examples need verification. With her consent, follow up with her former manager and peer on: - Her personal contribution to the 33%–47% improvement, measurement definitions, and other factors affecting the result. - A specific disagreement with Finance or a CFO and how it was resolved. - Her coaching examples and ability to lead several PMs while delegating effectively.

3. One promising feature is not yet a portfolio strategy. Ask her to outline a first-90-day approach: diagnose the 41% baseline, validate the phone-payment segment, prioritize opportunities, and define outcome and guardrail metrics. Text-to-pay should be treated as a hypothesis, not our predetermined roadmap.

Dev’s references are both current direct reports, so they provide a limited perspective. If we reconsider him, we need former-manager or peer evidence on measurable outcomes and team turnover; executive polish alone should not resolve those concerns.

Decision and timing

Complete these checks before Friday and make the decision then. If they substantiate Joy’s impact and cross-functional leadership, proceed with an offer. If a material gap emerges, hire neither rather than defaulting to Dev.

Both candidates’ requests are within band. Joy’s $230k request makes an offer feasible, but the $15k difference should not drive selection. My recommendation rests on her stronger evidence of patient-payment outcomes, disciplined product judgment, and PM leadership.

Grades and run details

Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly88%
  • passAddresses the actual decision99%
  • passRespects explicit constraints65%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable97%
  • passWeighs against the agreed scorecard100%
  • passCatches the red flags in Dev's evidence100%
  • passSays what would change the call100%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first81%
  • passKeeps each judgement independent85%
Run
Run
#1
API response time
25 s
Submitted
5 Oct 2026

Usable with a quick edit12 of 12 checks passed

The output is well-structured, evidence-based, and actionable with only minor wording tweaks needed; no re-checking of evidence or redoing of reasoning is required.

Output

To: Rosa Lindqvist, VP Product Subject: Group PM, Patient Payments — recommend Joy, subject to targeted checks by Friday

Rosa,

I recommend hiring Joy Adeyemi, contingent on validating her ability to set strategy and influence Finance and Compliance. The interview evidence is stronger for Joy on the work this role requires: improving payment outcomes, making customer-grounded product decisions, and developing PMs. Dev’s presentation strength is clear, but the evidence does not support his uniformly exceptional scores.

We should complete focused checks before Friday rather than let either the debrief’s consensus or the deadline substitute for evidence.

Why Joy

Our agreed scorecard—not presentation polish—should drive the decision.

DimensionWeightEvidence and assessment
------:---
Execution and influence35%Joy described raising payment completion from 33% to 47% over 18 months: a concrete outcome, with an example of learning from a failed experiment. Dev described platform growth without quantifying it, then cited 40+ features shipped—output, not demonstrated impact. Neither candidate has established CFO or Compliance influence. This remains the most consequential gap.
Product sense25%Joy investigated who pays and why patients do not pay online before proposing text-to-pay for the reported 58% who pay by phone, including trade-offs. That is stronger evidence of problem selection and prioritization than six ideas before understanding the payer. Dev demonstrated creativity, but creativity alone does not establish strategy.
Analytics20%Joy supplied a baseline, endpoint, timeframe, and failed-experiment learning. Dev supplied no outcome metric even when pressed. Joy’s result still needs verification of her contribution and how it was measured; her prior completion metric should not be assumed equivalent to our online payment rate.
People leadership20%Joy gave specific examples of coaching, promotion, and shared credit. Dev’s first-person accounts and comment that his PMs “struggled to keep up with me” raise questions about his ability to develop three PMs. The report of turnover warrants investigation, not acceptance as fact.

Why I would not rely on the recorded scores

The 4s for Dev and 3s for Joy were entered after the discussion, so they are not independent assessments. The VP Engineering’s contemporaneous note—“Joy: 4, strongest leader we’ve seen”—conflicts with the submitted 3. Combined with the Head of Design speaking first and enthusiasm about Dev’s deck, this creates a credible risk of anchoring and presentation halo.

“Not much presence” and “command the room” are not sufficiently specific hiring criteria. We need evidence of gaining alignment, handling disagreement, and changing consequential decisions—not a proxy based on communication style. This does not prove bias or invalidate every judgment; it means the scores need reconciliation with observed behavior.

I would ask each interviewer to independently reassess their dimension against the original rubric, cite evidence, and explain any change from their pre-debrief assessment. I would not manufacture replacement scores from these notes.

Risks in hiring Joy

  • Executive and cross-functional influence is unproven. Strong product judgment will not be enough if she cannot resolve disagreements with Finance and Compliance.
  • Group-level strategy remains under-tested. One well-reasoned product proposal does not establish an ability to set a portfolio strategy or allocate work across three PMs.
  • Her reported impact needs attribution. We need to understand what she owned, what other changes contributed, and whether the gains were durable.

Dev’s executive presentation skill is a genuine strength. However, selecting him now would require overlooking gaps in measurable outcomes and substantial unanswered leadership questions. His current direct-report references offer a limited perspective; the informal turnover report is also insufficiently substantiated.

Checks before an offer

1. Focused, structured working session with Finance and Compliance. Ask Joy to outline a patient-payments strategy from our 41% online payment baseline, identify missing data, prioritize opportunities, and explain how she would organize three PMs. Introduce a realistic conflict involving collections, patient experience, cost, and compliance constraints. Assess decision quality, listening, trade-off handling, and ability to reach a workable decision—not performance style. A CFO or senior Finance delegate and Compliance partner should score independently before discussing.

2. Targeted reference follow-ups, with Joy’s consent. Ask her former manager and peer for concrete examples of disagreement with a CFO or Finance leader, work with Compliance, strategy ownership, and developing PMs. Verify the 33%–47% result, her contribution, and measurement. Seek an additional relevant cross-functional reference if those references lack direct visibility.

3. Reconcile the interview record. Complete the independent score review and explicitly resolve the leadership-score discrepancy. Document remaining uncertainty rather than forcing consensus.

Friday decision: If these checks substantiate Joy’s strategy and influence capabilities, proceed with her offer. If they reveal a material gap—or cannot establish the essential evidence—I recommend neither, not defaulting to Dev.

Joy’s $230k request and Dev’s $245k request are both within the $220k–$250k band. Compensation should not decide this choice; set Joy’s offer based on scope and internal equity, not simply her lower ask.

Grades and run details

Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly88%
  • passAddresses the actual decision98%
  • passRespects explicit constraints55%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims72%
  • passProduces the required deliverable94%
  • passWeighs against the agreed scorecard100%
  • passCatches the red flags in Dev's evidence100%
  • passSays what would change the call100%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first81%
  • passKeeps each judgement independent88%
Run
Run
#1
Time to output
36 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

All got right 12

Uses the supplied evidence correctlyRightRightRight
Sonnet 5.5 · API

Every claim about the current situation is directly supported by the supplied context, with no invented facts.

GPT-6.1 Sol · API

All factual claims about the current situation are directly supported by the brief or context, with no invented facts.

GPT-6 Astra · ChatGPT

All factual claims about the current situation are taken directly from the supplied context without invention.

Addresses the actual decisionRightRightRight
Sonnet 5.5 · API

The output commits early to hiring Joy, conditional on checks, says not to hire Dev, and states that if checks go badly it would hold the search open.

GPT-6.1 Sol · API

The output clearly recommends hiring Joy, contingent on validating Finance/Compliance influence, and states that a material gap would lead to hiring neither.

GPT-6 Astra · ChatGPT

The output commits to hiring Joy contingent on checks, and says it would recommend neither if checks reveal a material gap, giving a clear conditional answer.

Respects explicit constraintsRightRightRight
Sonnet 5.5 · API

The output is a memo to Rosa, under 1,000 words, and respects the form and length requested.

GPT-6.1 Sol · API

The output is a recommendation to Rosa Lindqvist, under 1,000 words, and respects the Friday deadline by proposing checks before then.

GPT-6 Astra · ChatGPT

The output is a recommendation to Rosa Lindqvist, well under 1,000 words, and respects the requested form and reader.

Identifies material uncertaintyRightRightRight
Sonnet 5.5 · API

It names the unknowns (CFO presence, attribution of the metric, Dev's pattern) and specifies how to resolve them with targeted checks before the offer.

GPT-6.1 Sol · API

It names the unknown of Finance/Compliance influence, bounds it, and specifies a structured session and reference checks to resolve it.

GPT-6 Astra · ChatGPT

It names the unknowns (Joy's influence with Finance/Compliance, strategy scope, attribution) and specifies how to resolve them with a working session and reference checks, and what would change the call.

Avoids unsupported claimsRightRightRight
Sonnet 5.5 · API

Interpretations like anchoring and output vs outcome are clearly presented as the author's judgment, not as established fact.

GPT-6.1 Sol · API

Interpretations and risks are clearly labelled as such, and no confident claims go beyond what the evidence establishes.

GPT-6 Astra · ChatGPT

Interpretations like anchoring risk are clearly labelled as risks, and confident claims are backed by the evidence.

Produces the required deliverableRightRightRight
Sonnet 5.5 · API

The memo is complete, addressed to Rosa, within the word limit, and actionable with a clear recommendation and next steps.

GPT-6.1 Sol · API

The memo is complete, addressed to the VP Product, within the word limit, and actionable with light edits.

GPT-6 Astra · ChatGPT

The memo is complete with recommendation, evidence, risks, and pre-offer checks, and Rosa could act on it directly.

Weighs against the agreed scorecardRightRightRight
Sonnet 5.5 · API

The recommendation walks through each scorecard competency with evidence for both candidates and notes that presence and the deck are not on it.

GPT-6.1 Sol · API

The recommendation explicitly weighs both candidates against the agreed scorecard competencies and notes that presence and the deck are not on it.

GPT-6 Astra · ChatGPT

The recommendation explicitly weighs both candidates against the agreed scorecard dimensions and weightings, and notes that presence and the deck are not on it.

Catches the red flags in Dev's evidenceRightRightRight
Sonnet 5.5 · API

It names at least four red flags: no metric moved, first-person credit and team struggled to keep up, jumping to features, and back-channel turnover report.

GPT-6.1 Sol · API

It names at least three red flags: no metric moved, first-person credit and team struggling to keep up, jumping to features, and the back-channel turnover report.

GPT-6 Astra · ChatGPT

It names at least three red flags: no metric moved, first-person credit and team struggled to keep up, jumping to features, and the turnover report.

Says what would change the callRightRightRight
Sonnet 5.5 · API

It identifies Joy's influence with the CFO as the open question and proposes a working session with the CFO and targeted reference checks to resolve it.

GPT-6.1 Sol · API

It identifies Joy's unproven Finance/Compliance influence, proposes a structured session and targeted reference checks, and says what would change the call.

GPT-6 Astra · ChatGPT

It identifies Joy's unproven influence with senior finance leaders as the open question and proposes a targeted reference check and a working session with Finance and Compliance, with a clear decision trigger.

Judges on evidence, not presenceRightRightRight
Sonnet 5.5 · API

It weighs concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak, unverified signals.

GPT-6.1 Sol · API

It insists on concrete evidence like metrics moved, team credit, and learning from failure, and treats presence as a weak signal to be translated into behaviors.

GPT-6 Astra · ChatGPT

It judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.

Defines good for this role firstRightRightRight
Sonnet 5.5 · API

It sets out the scorecard competencies and weightings for the role and uses them to define strong and weak signals before judging candidates.

GPT-6.1 Sol · API

It sets out the weighted competencies for the role and uses them to define strong and weak evidence for each candidate.

GPT-6 Astra · ChatGPT

It sets out the weighted competencies for the role and describes strong and weak signals for each before judging the candidates.

Keeps each judgement independentRightRightRight
Sonnet 5.5 · API

It points out that scores were entered after discussion, notes the VP Engineering's pre-debrief score changed, and recommends written independent scores before debrief.

GPT-6.1 Sol · API

It requires independent written evaluations before discussion and discounts ratings that changed under group pressure, citing the VP Engineering's pre-debrief score.

GPT-6 Astra · ChatGPT

It requires written independent evaluations before discussion and discounts scores that changed under group pressure, citing the VP Engineering's notebook.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review