Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Director of Product at Medvia, and you're the hiring manager for a Group PM. The debrief was yesterday. Write your recommendation to Rosa Lindqvist, our VP Product, in no more than 1,000 words: who to hire (Dev, Joy, or neither), the evidence for it, the risks, and what you'd check before making the offer. The CEO wants a decision by Friday. The interview notes and everything else are below.

What the model was given6 items: The role and the scorecard (agreed before interviews began), Dev Malhotra: interview notes, Joy Adeyemi: interview notes, Scores (1 to 4, entered after the debrief discussion), References, Compensation
The role and the scorecard (agreed before interviews began)Group PM for patient payments, leading three PMs. Must set the strategy for patient payments, lift the online payment rate (now 41%), and lead PMs while working closely with Finance and Compliance. Weighting: execution and influence 35%, product sense 25%, analytics 20%, people leadership 20%.
Dev Malhotra: interview notesProduct sense (Head of Design): 'Brilliant. Redesigned our statement flow live, lots of ideas, very creative.' Notes show six feature ideas in the first five minutes, before asking who pays and why. Analytics (Data lead): 'Strong.' Notes: asked what metric he moved at his last job, he said 'we grew the payments platform massively'; pressed, 'the team shipped 40+ features'. No number. Leadership (VP Engineering): 'Confident, inspiring.' Notes: every story in the first person; said his last PM team 'struggled to keep up with me'. Presentation: 'the best deck we've ever seen'.
Joy Adeyemi: interview notesProduct sense (Head of Design): 'Solid, slower.' Notes: spent 15 minutes asking who pays and why they don't pay online, then proposed one change: text-to-pay for the 58% of patients who pay by phone, with its trade-offs. Analytics (Data lead): notes say she raised payment completion at her last company from 33% to 47% over 18 months, and walked through an experiment that failed and what it taught her. Leadership (VP Engineering): notes say she coached a struggling PM into a different role, promoted another, and named her engineers when describing wins. Debrief comments: 'Not much presence.' 'Not sure she'd command the room with the CFO.'
Scores (1 to 4, entered after the debrief discussion)Dev: 4, 4, 4, 4. Joy: 3, 3, 3, 3. The Head of Design spoke first in the debrief. The VP Engineering's own notebook, written before the debrief, says: 'Joy: 4, strongest leader we've seen.' They entered a 3 after the discussion.
ReferencesDev: two references, both people who report to him now. Rosa knows someone at his previous company, who says: 'great in front of execs; his team turned over a lot.' Joy: her former manager and a peer, both strongly positive. Nobody has asked either reference about working with a CFO.
CompensationBand $220k to $250k. Dev is asking for $245k, Joy for $230k.
What a strong answer doesThe answer key the graders mark against

Recommends Joy, or at minimum not Dev on this evidence. Weighs both against the agreed scorecard rather than presence or the deck: Joy gives a metric she moved (33% to 47%), learns from failure, credits her team and coaches PMs, which fits execution, analytics and people leadership; Dev couldn't name a metric he moved, frames everything as 'I', jumped to features before asking about users, and the back channel reports high team turnover, all red flags. Points out that the scores were entered after a discussion the Head of Design opened, and the VP Engineering moved Joy from 4 to 3, so they show anchoring, not independent judgement. Treats 'presence with the CFO' as the one real open question and proposes a targeted reference check on how Joy worked with Finance leaders before the offer. Suggests the next loop collect written scores before the debrief.

Critical failures (cap the score at 40)
  • Recommends hiring Dev on the strength of the debrief scores or the presentation, without addressing that he couldn't name a metric he moved
Task

Staff-level · v1.0 · synthetic · healthcare, fintech, hiring, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit12 of 12 checks passed

The output is well-structured, evidence-based, and actionable; a PM could use it with minor wording tweaks and no need to re-check facts or reasoning.

Output

To: Rosa Lindqvist, VP Product From: Director of Product Subject: Group PM, Patient Payments — recommend Joy, subject to targeted checks

Rosa,

I recommend hiring Joy Adeyemi, contingent on validating Finance/Compliance influence and her reported results before Friday. She has the stronger evidence against our agreed scorecard, particularly for improving payment outcomes and leading three PMs. Dev’s presentation strength is real, but the record does not substantiate his uniformly high scores.

Evidence for the recommendation

Our scorecard weights execution and influence at 35%, product sense at 25%, analytics at 20%, and people leadership at 20%. We should judge the underlying evidence against those criteria—not the quality of the deck or who spoke first in the debrief.

CriterionEvidence and judgment
Execution and influence — 35%Joy described raising payment completion from 33% to 47% over 18 months and explained a failed experiment and its lessons. That is more relevant evidence of outcome delivery than Dev’s “40+ features.” However, neither candidate has demonstrated CFO partnership sufficiently. This remains our largest unresolved question.
Product sense — 25%Joy investigated who pays and why patients do not pay online before proposing text-to-pay for the reported 58% who pay by phone. She discussed trade-offs. Dev generated six ideas before establishing the customer problem. His creativity is a strength; Joy’s problem framing is better supported for this role.
Analytics — 20%Joy supplied a baseline, an end state, a timeframe, and learning from an unsuccessful experiment. Dev supplied neither a quantified outcome nor a clear connection between shipping and payment performance when pressed. The “Strong” label for Dev is not supported by the notes.
People leadership — 20%Joy gave concrete examples of coaching, role fit, promotion, and crediting engineers. Dev’s first-person accounts and comment that his PMs “struggled to keep up with me” raise questions about developing a team. The informal report of turnover reinforces the need to investigate, but is not proof of poor leadership.

Joy’s prior payment-completion metric is not necessarily equivalent to our 41% online payment rate. We should not assume her result transfers directly. Still, it is substantially better evidence of outcome orientation than feature volume.

Why I would not use the final scores as the deciding evidence

Dev’s straight 4s and Joy’s straight 3s conflict with the interview notes. Most clearly, VP Engineering recorded Joy as “4, strongest leader we’ve seen” before the debrief, then entered a 3 afterward. With scores entered after discussion and Design speaking first, there is a credible risk of anchoring and presentation halo—not proof of anyone’s intent.

I would ask interviewers to independently reconcile their ratings with specific scorecard evidence before another group discussion. “Presence” and “commanding the room” need to translate into observable behaviors: handling disagreement, making a defensible recommendation, and winning commitment. They should not become an unannounced fifth criterion.

Risks and checks before an offer

1. Finance and Compliance influence is unproven. Run a short, structured session with Finance and Compliance. Ask Joy to defend a patient-payments strategy when the CFO challenges its economics and Compliance challenges its outreach approach. Assess whether she understands the objections, makes trade-offs, and secures a workable decision—not how forcefully she speaks.

2. Her impact and leadership examples need verification. With her consent, follow up with her former manager and peer on: - Her personal contribution to the 33%–47% improvement, measurement definitions, and other factors affecting the result. - A specific disagreement with Finance or a CFO and how it was resolved. - Her coaching examples and ability to lead several PMs while delegating effectively.

3. One promising feature is not yet a portfolio strategy. Ask her to outline a first-90-day approach: diagnose the 41% baseline, validate the phone-payment segment, prioritize opportunities, and define outcome and guardrail metrics. Text-to-pay should be treated as a hypothesis, not our predetermined roadmap.

Dev’s references are both current direct reports, so they provide a limited perspective. If we reconsider him, we need former-manager or peer evidence on measurable outcomes and team turnover; executive polish alone should not resolve those concerns.

Decision and timing

Complete these checks before Friday and make the decision then. If they substantiate Joy’s impact and cross-functional leadership, proceed with an offer. If a material gap emerges, hire neither rather than defaulting to Dev.

Both candidates’ requests are within band. Joy’s $230k request makes an offer feasible, but the $15k difference should not drive selection. My recommendation rests on her stronger evidence of patient-payment outcomes, disciplined product judgment, and PM leadership.

Grades and run details

Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly88%
  • passAddresses the actual decision99%
  • passRespects explicit constraints65%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims70%
  • passProduces the required deliverable97%
  • passWeighs against the agreed scorecard100%
  • passCatches the red flags in Dev's evidence100%
  • passSays what would change the call100%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first81%
  • passKeeps each judgement independent85%
Run
Run
#1
API response time
25 s
Submitted
5 Oct 2026

Usable with a quick edit10 of 12 checks passed

The output is well-structured and evidence-based; a quick edit to add the condition that would change the call (e.g., 'if the reference reveals she struggled with finance, we will not proceed') would make it fully complete.

Output

MEMORANDUM

TO: Rosa Lindqvist, VP of Product FROM: Director of Product, Medvia DATE: October 24, 2023 SUBJECT: Hiring Recommendation: Group PM, Patient Payments

*

Rosa,

Following yesterday’s debrief, I have reviewed the scorecards, interview transcripts, and reference notes for our open Group PM role. We are tasked with finding a leader who can set strategy, lift our online payment rate from its current 41%, lead a team of three PMs, and collaborate effectively with Finance and Compliance in a regulated healthcare environment.

The CEO needs our decision by Friday. My recommendation is that we extend an offer to Joy Adeyemi at $230k. We should not hire Dev Malhotra.

Below is the evidence, the associated risks, and the pre-offer checks required to secure Joy.

---

1. The Evidence: Why Joy Adeyemi over Dev Malhotra

While the final entered scores (Dev straight 4s; Joy straight 3s) suggest a clear winner, the raw interview notes and hidden dynamics of our debrief tell a different story.

  • Product Sense & Problem Solving: Dev is a classic "feature factory" ideator. The Head of Design noted Dev pitched six features in five minutes before asking who pays and why. Conversely, Joy spent her first 15 minutes investigating the root causes of friction—asking who pays and why others don’t—before proposing a targeted, high-leverage solution: text-to-pay tailored to the 58% of patients currently paying via phone. For a critical metric like our 41% online payment rate, we need clinical execution, not a scattershot feature list.
  • Analytics & Rigor: The Data Lead’s notes on Dev are a red flag. When pressed on metrics, Dev offered vague marketing speak ("grew massively," "shipped 40+ features") with zero data attribution. Joy provided hard, demonstrable impact: she raised payment completion at her previous company from 33% to 47% over 18 months. Crucially, she also walked through a failed experiment and extracted lessons—proving she understands the complex trade-offs of healthcare payments.
  • People Leadership & Culture: Dev speaks entirely in the first person ("I," "my wins") and casually remarked that his last team "struggled to keep up with me." Rosa, you likely noted your contact’s warning that his team suffered high turnover. Joy is the inverse: she coaches struggling reports, promotes talent, and attributes wins directly to her engineers. The VP of Engineering’s pre-debrief notebook literally rated Joy as a "4, strongest leader we've seen," before social pressure in the debrief caused the scores to normalize.
  • The "Presence" Bias: The primary knock on Joy during the debrief was a lack of "presence" and a concern about whether she could "command the room with the CFO." This is a subjective trap. We do not need a charismatic showman to present slide decks; we need a strategist who can build bulletproof business cases with Finance and navigate Compliance. Joy’s methodical, customer-first approach is precisely what Finance will trust.

---

2. Risks to Manage

  • The CFO/Executive Presence Risk: Because the interview panel worried Joy might be quiet in front of executive stakeholders, we must validate her ability to defend her strategy to our CFO and Compliance leads.
  • Compensation and Closing: Joy is asking for $230k (right in the middle of our $220k–$250k band), making her a financially sound hire. However, because she is measured and thoughtful, she may take a few days to deliberate. We need to move decisively.

---

3. Pre-Offer Checks (To Be Completed by Thursday EOD)

Before we release the offer to Joy, I will execute two final validation steps:

  1. Backchannel/Targeted Reference Check: Since our previous reference calls did not explicitly probe her dynamic with Finance or Compliance, I will spend 10 minutes with her former manager specifically asking: “How did Joy handle budget cross-examinations with Finance? How did she partner with Legal/Compliance when building payment flows?”
  2. The "CFO Alignment" Test: I’d like to spend 30 minutes with Joy and our current Director of Finance (or you and I) presenting a mock 10-minute briefing on our 41% online payment bottleneck, asking her to walk us through how she would construct the business case for the Q1 roadmap.

Conclusion

Dev is dazzling in a pitch meeting, but his metrics are hollow, his leadership style is toxic to retention, and his instinct is to build before understanding. Joy is rigorous, customer-obsessed, metrics-driven, and a multiplier for her team and cross-functional partners.

Let’s trust the substance over the theater. I am ready to initiate the final check and prepare the $230k offer for Friday morning.

Grades and run details

Decision model 92 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly26%
  • passAddresses the actual decision76%
  • passRespects explicit constraints60%
  • passIdentifies material uncertainty78%
  • partialAvoids unsupported claims56%
  • passProduces the required deliverable98%
  • passWeighs against the agreed scorecard94%
  • passCatches the red flags in Dev's evidence100%
  • passSays what would change the call48%
  • passJudges on evidence, not presence98%
  • partialDefines good for this role first18%
  • passKeeps each judgement independent68%
Run
Run
#1
Time to output
6 s
Submitted
5 Oct 2026

Usable with a quick edit12 of 12 checks passed

The output is well-structured, evidence-based, and actionable; a PM could use it with minor wording edits without re-checking evidence or redoing reasoning.

Output

To: Rosa Lindqvist, VP Product Subject: Group PM recommendation — Joy Adeyemi

Recommendation: Hire Joy, subject to focused reference and influence checks before the offer. She has the stronger evidence against the role’s core needs: measurable payment outcomes, thoughtful product judgment, and the ability to lead through other people. The CEO needs a decision by Friday; I would not let the debrief scores alone drive it.

Evidence

  • Execution and influence (35%): Joy gave a specific result from her last company—payment completion rose from 33% to 47% over 18 months—and explained a failed experiment and what she learned. Her leadership examples showed investment in others: coaching a struggling PM into a different role and promoting another. Dev’s evidence was less concrete: he described platform growth and 40+ shipped features but supplied no outcome metric when pressed. His stories were consistently framed around his own contribution, and he said his team struggled to keep up with him.
  • Product sense (25%): Joy investigated who pays and why before proposing text-to-pay for the 58% who currently pay by phone, including trade-offs. That is relevant to our 41% online payment rate and reflects problem-first judgment. Dev’s creativity and presentation were strengths, but six ideas in the first five minutes—before clarifying who pays and why—are a concern for this role.
  • Analytics (20%): Joy’s quantified result and discussion of a failed experiment are stronger evidence of analytical ownership. Dev did not provide a measurable result.
  • People leadership (20%): Joy’s examples demonstrate coaching, development, and shared credit. Dev’s confidence and presentation are positives, but they do not offset the limited evidence of developing a team or his comment about the team struggling to keep up.

Risks and scorecard

Joy’s main risk is whether she can influence senior stakeholders, particularly the CFO and Finance. The “not much presence” and “command the room” comments are concerns to test, not proof that she cannot do it; we should assess her ability to make a clear, evidence-based case rather than equate influence with a particular presentation style.

Dev may be compelling with executives, but we need to understand the team-turnover signal and whether his leadership style scales. His two references are current direct reports, which may limit candor. His former-company reference’s report of high turnover is important but unverified.

I would treat the 4/4/4/4 versus 3/3/3/3 ratings as weak evidence. The VP Engineering’s pre-debrief note called Joy the strongest leader seen and rated her a 4, then the score became a 3 after discussion; the Head of Design also spoke first. This does not establish why the ratings changed, but it is reason to rely on the underlying evidence and check for independent calibration.

Before making the offer

  1. Ask Joy’s former manager and peer specifically about her influence with Finance/CFO-level stakeholders, including a difficult disagreement and the outcome.
  2. Ask for one additional Joy reference who can speak to cross-functional execution; ask Dev’s references about team turnover, feedback, and how he develops PMs. Seek candid, comparable evidence where possible.
  3. In a short structured conversation, ask Joy to make the case for a payment-rate strategy to a skeptical CFO. Evaluate clarity, use of evidence, and handling of challenge—not polish alone.
  4. Confirm Joy’s 33%→47% result, her role in it, and the scope of her people leadership.

If those checks reveal a material gap in Joy’s influence or ownership, I would choose neither rather than default to Dev on presentation strength. Otherwise, proceed with Joy at her requested $230k, within band.

Grades and run details

Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly87%
  • passAddresses the actual decision100%
  • passRespects explicit constraints76%
  • passIdentifies material uncertainty100%
  • passAvoids unsupported claims60%
  • passProduces the required deliverable99%
  • passWeighs against the agreed scorecard100%
  • passCatches the red flags in Dev's evidence100%
  • passSays what would change the call100%
  • passJudges on evidence, not presence100%
  • passDefines good for this role first69%
  • passKeeps each judgement independent87%
Run
Run
#1
API response time
12 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 2

Says what would change the callRightMixedRight
GPT-6.1 Sol · API

It identifies Joy's unproven Finance/Compliance influence, proposes a structured session and targeted reference checks, and says what would change the call.

Gemini 3.5 Flash-Lite · Gemini

It names the CFO influence open question and proposes a targeted reference check and mock briefing, but does not state what answer or result from those checks would change the hiring decision.

GPT-6 Luna · API

Identifies Joy's influence with CFO as the open question and proposes specific reference checks and a structured conversation to test it.

Defines good for this role firstRightWrongRight
GPT-6.1 Sol · API

It sets out the weighted competencies for the role and uses them to define strong and weak evidence for each candidate.

Gemini 3.5 Flash-Lite · Gemini

It does not explicitly set out the weighted competencies and what strong and weak look like for this role before judging the candidates.

GPT-6 Luna · API

Sets out the weighted competencies and what strong evidence looks like before comparing candidates.

All got right 10

Uses the supplied evidence correctlyRightRightRight
GPT-6.1 Sol · API

All factual claims about the current situation are directly supported by the brief or context, with no invented facts.

Gemini 3.5 Flash-Lite · Gemini

Every claim about the current situation is directly supported by the supplied context, with no invented facts.

GPT-6 Luna · API

All factual claims about the current situation are directly from the brief, with no inventions.

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

The output clearly recommends hiring Joy, contingent on validating Finance/Compliance influence, and states that a material gap would lead to hiring neither.

Gemini 3.5 Flash-Lite · Gemini

The output unambiguously recommends hiring Joy and not Dev, framed for Rosa, and implies the offer is contingent on the pre-offer checks resolving the CFO concern.

GPT-6 Luna · API

Recommends hiring Joy early, framed for Rosa, and states that material gaps in influence or ownership would change the call to neither.

Respects explicit constraintsRightRightRight
GPT-6.1 Sol · API

The output is a recommendation to Rosa Lindqvist, under 1,000 words, and respects the Friday deadline by proposing checks before then.

Gemini 3.5 Flash-Lite · Gemini

The memo is addressed to Rosa, stays under 1,000 words, and includes the required elements: who to hire, evidence, risks, and pre-offer checks.

GPT-6 Luna · API

Memo to Rosa, under 1000 words, respects the Friday deadline by making a decision now with pre-offer checks.

Identifies material uncertaintyRightRightRight
GPT-6.1 Sol · API

It names the unknown of Finance/Compliance influence, bounds it, and specifies a structured session and reference checks to resolve it.

Gemini 3.5 Flash-Lite · Gemini

It names the CFO presence concern as the key unknown, bounds it to the panel's impression, and proposes specific checks to resolve it.

GPT-6 Luna · API

Names unknowns (Joy's CFO influence, Dev's turnover, debrief score reliability) and specifies how to resolve them with targeted checks.

Avoids unsupported claimsRightRightRight
GPT-6.1 Sol · API

Interpretations and risks are clearly labelled as such, and no confident claims go beyond what the evidence establishes.

Gemini 3.5 Flash-Lite · Gemini

Interpretations like 'feature factory' are clearly presented as conclusions drawn from the evidence, not as established facts.

GPT-6 Luna · API

Interpretations are clearly labelled as risks or weak evidence, not presented as established facts.

Produces the required deliverableRightRightRight
GPT-6.1 Sol · API

The memo is complete, addressed to the VP Product, within the word limit, and actionable with light edits.

Gemini 3.5 Flash-Lite · Gemini

The recommendation memo is complete, in the correct form, and could be acted on with light edits.

GPT-6 Luna · API

Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.

Weighs against the agreed scorecardRightRightRight
GPT-6.1 Sol · API

The recommendation explicitly weighs both candidates against the agreed scorecard competencies and notes that presence and the deck are not on it.

Gemini 3.5 Flash-Lite · Gemini

The output evaluates candidates on product sense, analytics, and people leadership, and explicitly dismisses presence and the deck as not on the scorecard.

GPT-6 Luna · API

Weighs both candidates against the agreed competencies (execution, product sense, analytics, people leadership) and notes that presence and deck are not on it.

Catches the red flags in Dev's evidenceRightRightRight
GPT-6.1 Sol · API

It names at least three red flags: no metric moved, first-person credit and team struggling to keep up, jumping to features, and the back-channel turnover report.

Gemini 3.5 Flash-Lite · Gemini

It catches Dev's missing metric, first-person credit and team struggle, jumping to features, and the back-channel turnover report—all four red flags.

GPT-6 Luna · API

Names at least three red flags: no metric moved, first-person credit and team struggled to keep up, jumping to features, and back-channel turnover report.

Judges on evidence, not presenceRightRightRight
GPT-6.1 Sol · API

It insists on concrete evidence like metrics moved, team credit, and learning from failure, and treats presence as a weak signal to be translated into behaviors.

Gemini 3.5 Flash-Lite · Gemini

It judges on concrete evidence (metric moved, team credit, learning from failure) and treats presence as a weak, subjective signal.

GPT-6 Luna · API

Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.

Keeps each judgement independentRightRightRight
GPT-6.1 Sol · API

It requires independent written evaluations before discussion and discounts ratings that changed under group pressure, citing the VP Engineering's pre-debrief score.

Gemini 3.5 Flash-Lite · Gemini

It relies on the VP Engineering's independent pre-debrief rating and discounts the score that changed under group pressure.

GPT-6 Luna · API

Relies on the VP Engineering's independent pre-debrief note and discounts group-influenced scores, keeping judgement independent.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review