Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Director of Product at Kestrel. Elif Demir, our VP Product, wants to fix the PM interview loop before we hire four more PMs next quarter (her note is below). Using the data we have, write her a memo of no more than 1,000 words: which parts of the loop seem to predict how PMs do once hired and which don't, how confident we can be, and the loop you'd run next quarter.\n\nThe files are attached.

What the model was given3 items: elif_note.md, current_loop.md, pm_candidates_last_18_months.csv (every PM candidate who reached the final rounds)
elif_note.md1 lines · Download
From Elif (VP Product): "We've made three PM hires in 18 months who struggled. I want to know what in our loop is and isn't working, using the data we have, and what to change before we hire four more PMs next quarter."
current_loop.md10 lines · Download
# Current PM interview loop

1. Recruiter screen.
2. Take-home: a product strategy deck, one week to complete.
3. Presentation of the take-home to a panel.
4. Behavioural round: "Tell me about a metric you moved", with follow-ups on what you did and what you learned.
5. VP Engineering conversation ("technical credibility").
…
pm_candidates_last_18_months.csv (every PM candidate who reached the final rounds)candidate,background,takehome_score,presentation_score,behavioural_metric_round_score,vp_engineering_score,outcome,rating_after_12_months C100,Big tech,4,4,3,2,Withdrew, C101,Startup,3,2,4,2,Hired,5 C102,Big tech,1,2,3,4,Rejected, C103,Agency,1,3,2,3,Rejected, C104,Big tech,4,3,3,4,Hired,3 C105,Startup,4,2,1,1,Rejected, C106,Startup,4,2,1,3,Rejected, C107,Big tech,4,4,2,4,Hired,2 C108,Consulting,1,4,2,1,Rejected, C109,Agency,4,2,1,2,Rejected, C110,Consulting,3,1,2,2,Rejected, C111,Big tech,2,2,1,4,Rejected, C112,Agency,3,1,2,1,Rejected, C113,Startup,4,2,3,1,Withdrew, C114,Big tech,2,2,3,2,Withdrew, C115,Consulting,3,1,1,1,Rejected, C116,Startup,2,2,1,1,Rejected, C117,Startup,2,1,4,3,Withdrew, C118,Startup,2,2,1,2,Withdrew, C119,Consulting,1,1,3,3,Rejected, C120,Consulting,2,4,1,3,Rejected, C121,Consulting,2,3,4,1,Rejected, C122,Agency,4,1,3,3,Withdrew, C123,Agency,2,4,4,1,Rejected, C124,Big tech,2,2,2,3,Rejected, C125,Startup,2,2,1,3,Rejected, C126,Consulting,4,1,2,2,Rejected, C127,Consulting,1,3,3,1,Rejected, C128,Startup,3,3,2,3,Rejected, C129,Agency,3,3,3,1,Rejected, C130,Startup,1,3,2,2,Rejected, C131,Agency,1,2,4,2,Rejected, C132,Big tech,3,4,2,4,Hired,2 C133,Agency,2,4,2,1,Rejected, C134,Consulting,4,4,3,3,Hired,3 C135,Consulting,3,4,4,3,Rejected, C136,Consulting,2,2,3,2,Rejected, C137,Consulting,3,4,3,1,Withdrew, C138,Startup,2,4,4,2,Withdrew, C139,Startup,3,2,2,3,Rejected, C140,Startup,1,1,3,3,Rejected, C141,Startup,3,3,4,3,Hired,5 C142,Big tech,4,2,3,4,Withdrew, C143,Consulting,1,4,3,2,Rejected, C144,Consulting,1,1,1,2,Rejected, C145,Agency,4,1,2,1,Rejected, C146,Startup,3,1,4,1,Rejected, C147,Big tech,4,1,1,4,Withdrew, C148,Big tech,4,4,2,4,Hired,2 C149,Agency,2,3,4,2,Hired,4 C150,Consulting,4,1,1,2,Rejected, C151,Startup,2,3,3,3,Hired,4
What a strong answer doesThe answer key the graders mark against

Uses the nine hires to compare each round with 12-month ratings: the behavioural 'metric you moved' round lines up strongly with how hires performed (higher scores, better ratings), while the take-home, the presentation and the VP Engineering conversation line up the wrong way: the hires they scored highest are the ones rated 2. The VP Engineering scores big-tech candidates higher, and all four big-tech hires were rated 2 or 3. Is explicit that nine hires is far too few to conclude much: the pattern is a reason to change and then measure, not proof. Proposes a loop with a scorecard defined up front, the behavioural round strengthened (more than one, structured), the week-long take-home replaced or cut down (it costs candidates, and some withdraw), the VP Engineering round given a defined bar and calibrated, and every interviewer writing a score before the debrief. Sets up tracking so the next hires test the change.

Critical failures (cap the score at 40)
  • Recommends keeping the take-home and presentation as the main deciding rounds
Task

Staff-level · v1.0 · synthetic · hiring, data files, staff-level

The outputs

5 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit9 of 12 checks passed

The output misses the VP Engineering big-tech pattern and does not define role competencies, so a PM would need to add those before using it.

Output

To: Elif Demir, VP Product From: Director of Product Subject: PM interview loop: what the data says and what to change next quarter

What seems to predict performance

We have records for 52 final-round candidates, including 8 hires with 12-month ratings. Those hires averaged 3.5, versus a company PM average of 3.4; four were below 3.4, including two rated 2.

Among the eight hires, the clearest signal is the behavioural metric round. Its score has a strong positive association with 12-month ratings (exploratory Pearson r ≈ +0.93). The four hires rated 4–5 averaged 3.75 on this round, versus 2.5 for the four rated 2–3. This is promising evidence for probing how candidates actually move metrics and what they learn—not proof that the round is a reliable predictor.

The other scored rounds show no positive relationship with ratings in this small hired sample:

RoundAssociation with 12-month rating among hires
------:
Take-homer ≈ −0.50
Presentationr ≈ −0.85
Behavioural metricr ≈ +0.93
VP Engineeringr ≈ −0.79

The negative estimates for presentation and VP Engineering should not be read as evidence that those skills hurt PM performance. With only eight hires, these figures are highly unstable. They do tell us that our current scores offer no reassuring evidence that those rounds predict success.

How confident we should be

Low confidence overall; modest confidence that the behavioural round is the best lead to investigate. Eight outcomes are far too few to set reliable weights or cutoffs, and interview scores are ordinal ratings, so the correlations are descriptive only. We also observe performance only for people we hired: panel decisions restrict the range of scores and can distort associations. Non-hires have no performance outcome for comparison.

We lack recruiter-screen scores, interviewer-level data, consistent evidence of scoring reliability, and context about role, manager, or rating differences. The company-average comparison is also a coarse measure. The data can guide a better experiment; it cannot establish causation or validate the loop.

Loop I’d run next quarter

  1. Keep the recruiter screen for eligibility, motivation, and logistics, but don’t treat it as a predictive assessment. We have no screen-score data to evaluate.
  2. Replace the week-long take-home deck and separate presentation with one time-boxed, job-relevant work sample. Use a live or short-prep product problem, not a polished deck. Score problem framing, user evidence, prioritization and tradeoffs, success measures, and communication against anchored criteria. This reduces candidate burden and stops us treating two related presentation exercises as independent evidence.
  3. Keep and strengthen the behavioural round. Ask for a specific metric the candidate owned or influenced; probe the baseline, their actions, evidence of impact, attribution, tradeoffs, and what they learned. Use the same questions and anchored scorecard for every candidate.
  4. Refocus the VP Engineering conversation on PM–engineering collaboration and concrete technical tradeoffs—not general “technical credibility.” Use a structured scenario and score it consistently; don’t make an unstructured impression a veto.
  5. Have interviewers score independently before the panel discussion. The hiring manager can decide after reviewing evidence against the role criteria, with any disagreement recorded. Do not introduce numerical weights based on these eight hires.

For each of the four hires, preserve round-level scores and interviewer notes, then collect comparable 6- and 12-month outcomes with role and manager context. Review the evidence after the cohort has matured. Next quarter’s loop should be a more structured, lower-burden test—not a claim that we have already found a validated predictor.

Grades and run details

Decision model 83 · LLM judge 10 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly21%
  • passAddresses the actual decision61%
  • passRespects explicit constraints50%
  • passIdentifies material uncertainty77%
  • passAvoids unsupported claims59%
  • passProduces the required deliverable64%
  • passFinds which rounds line up with performance92%
  • passSays how little nine hires can show99%
  • partialSpots the interviewer pattern86%
  • passJudges on evidence, not presence97%
  • partialDefines good for this role first28%
  • passKeeps each judgement independent97%
Run
Run
#1
API response time
33 s
Submitted
5 Oct 2026

Needs more than a quick edit9 of 12 checks passed

The memo contains arithmetic errors in the take-home and presentation averages and the VP Engineering count, so a PM would need to re-check the data before using it.

Output

To: Elif Demir, VP Product From: Director of Product Re: What our PM loop predicts, and the loop for next quarter

Summary

Of the five stages in our loop, only the behavioural "metric you moved" round tracks how hires perform after 12 months. The take-home, the presentation and the VP Engineering conversation all point the wrong way: hires who scored higher on them performed worse. The evidence is thin (nine hires), so I'd treat this as a strong hint rather than proof. Even so, it is consistent enough to change the loop now and to measure the new loop properly.

What the data shows

Over 18 months, 52 candidates reached the final rounds. We hired 9, rejected 33, and 10 withdrew. Our hires averaged 3.3 at 12 months, slightly below the PM average of 3.4. The three who struggled are C107, C132 and C148, all rated 2.

Average 12-month rating of hires, by interview score

StageHires who scored 4Hires who scored 3Hires who scored 2Correlation with rating
Behavioural (metric moved)4.7 (n=3)3.3 (n=3)2.0 (n=3)+0.94
Take-home2.5 (n=4)4.0 (n=3)4.0 (n=2)−0.57
Presentation2.25 (n=4)4.0 (n=4)5.0 (n=1)−0.87
VP Engineering2.25 (n=4)4.0 (n=3)4.5 (n=2)−0.82

The three strugglers share an identical profile:

  • 4 on the presentation
  • 4 on VP Engineering
  • 3 or 4 on the take-home
  • 2 on the behavioural round
  • Big-tech background

Our two strongest hires, C101 and C141 (both rated 5), had middling presentation and VP Engineering scores but scored 4 on the behavioural round.

In other words, we have been hiring for polish and for technical fluency, and overriding the one signal that worked.

Background is tangled up in this. All four big-tech hires were rated 2 or 3, averaging 2.25. The five others averaged 4.2. The VP Engineering round looks especially exposed to this:

  • Every VP Engineering score of 4 in the dataset went to a big-tech candidate (9 of 11 big-tech candidates).
  • No other candidate scored 4.

The round may be measuring familiarity with big-tech engineering culture rather than how well someone will work with our engineers.

The behavioural result holds within each group. This is the most reassuring part of the analysis.

  • Among non-big-tech hires, the behavioural score still tracks performance. Those who scored 4 were rated 5, 5 and 4. Those who scored 3 were rated 3 and 4.
  • Among big-tech hires, the one who scored 3 was rated 3 and the three who scored 2 were rated 2.

So the behavioural signal isn't simply a stand-in for background. Within non-big-tech hires, the presentation and take-home scores still lean negative.

How confident we can be

I'm only moderately confident, for five reasons:

  • Nine data points. One or two different outcomes could change the picture, and none of these correlations would survive a strict significance test.
  • We only see outcomes for people we hired. Hires were selected partly on high presentation and VP Engineering scores, so we can't see how low scorers on those stages would have done. We also don't know about strong behavioural candidates we rejected (C121, C123, C131, C135, C146), or about withdrawals. C135, for example, was rejected despite scoring 3, 4, 4 and 3, among the strongest sets of scores in the pool.
  • Background is confounded. With four big-tech hires, I can't fully separate "big-tech PMs struggle here" from "our loop rewards the wrong things in big-tech PMs." Either way, the fix is to the instruments, not to filter on background.
  • 12-month ratings are one manager's judgement, and they also reflect team and scope.
  • Interviewer effects. We don't know who scored which candidates, so some of this may come from individual raters rather than from the stages themselves.

What I am confident about is narrower. The take-home, presentation and VP Engineering scores give us no evidence that they help. Given what they cost candidates and the panel, that is reason enough to stop relying on them.

The loop I'd run next quarter

1. Recruiter screen. Unchanged.

2. Structured behavioural round, now the core of the loop (two interviewers, about 60 minutes). - Keep "a metric you moved," and add a second prompt: "a decision you got wrong and what you changed." - Use a written rubric that scores: - specificity about the candidate's own actions; - how they diagnosed the problem; - the trade-offs they made; - what they learned. - Score the substance, not the delivery. Each interviewer scores independently.

3. Replace the week-long take-home and presentation with a 90-minute live working session. - Give the candidate a real, anonymised Kestrel problem with messy data, and work through it with a PM and an engineer. - Score how they reason, prioritise and respond to pushback, not how finished the output looks. - This also cuts the candidate time we ask for. Ten of 52 finalists withdrew; I can't say the week-long take-home caused that, but it won't have helped.

4. Redesign the VP Engineering round as a structured "working with engineering" round. - Base it on a scenario such as a scope cut, a technical-debt trade-off or an estimate dispute, with a rubric agreed with Engineering. - Score it pass/concern only until we have evidence that it predicts anything, so it can flag real problems without rewarding big-tech vocabulary.

5. Change how we decide. - All scores and written evidence go in before the debrief, so nobody anchors on the loudest voice. - No hire with a behavioural score below 3. - The hiring manager still decides. If they hire against the behavioural score, they write down why.

6. Measure it this time. - Log every interviewer's scores, not just the panel's. - Collect 6-month and 12-month ratings against a consistent rubric. - With four more hires we'll have 13 to analyse. We should review the data in two quarters, not two years. - I'd also like a quick look back at the strong behavioural candidates we rejected, to see whether the debrief was systematically discounting that round.

The ask

I'd like your sign-off to:

  • pilot this loop for the four upcoming hires;
  • get two hours of interviewer calibration on the new rubrics;
  • secure the VP Engineering's agreement to redesign their round.

I can have the rubrics and the working-session problem ready within three weeks.

Claims the judge couldn’t find in the brief

  • Take-home: hires scoring 4 averaged 2.5, 3 averaged 4.0, 2 averaged 4.0.
  • Presentation: hires scoring 4 averaged 2.25, 3 averaged 4.0, 2 averaged 5.0.
  • Every VP Engineering score of 4 in the dataset went to a big-tech candidate (9 of 11 big-tech candidates).

Grades and run details

Decision model 92 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly35%
  • passAddresses the actual decision82%
  • passRespects explicit constraints53%
  • passIdentifies material uncertainty96%
  • partialAvoids unsupported claims26%
  • passProduces the required deliverable93%
  • passFinds which rounds line up with performance98%
  • passSays how little nine hires can show100%
  • passSpots the interviewer pattern100%
  • passJudges on evidence, not presence98%
  • partialDefines good for this role first44%
  • passKeeps each judgement independent100%
Run
Run
#1
Time to output
60 s
Submitted
5 Oct 2026

Usable with a quick edit12 of 12 checks passed

The memo is clear, evidence-based, and actionable; a PM could use it with minor wording tweaks and no need to re-verify the data or reasoning.

Output

To: Elif Demir, VP Product From: Director of Product Subject: PM interview loop: what the evidence supports and what to change next quarter

Recommendation

For the next four hires, make demonstrated product judgment, metric ownership, and cross-functional execution the center of the loop. Replace the week-long strategy deck and separate presentation with a short, standardized work sample; retain and strengthen the behavioral metric interview; replace the unstructured “technical credibility” conversation with a scored collaboration exercise.

These are evidence-informed changes, not a validated hiring formula. Our strongest signal comes from just nine hires.

What appears to predict performance

We have 52 finalists: nine hired, 33 rejected, and ten withdrawn. Only the nine hires have performance outcomes. Their average 12-month rating is 3.33, versus the company PM average of 3.4. Three received a 2; none received a 1. The similar averages do not negate the three struggling hires, but neither establish that the entire loop is failing.

Among hires:

Interview componentCorrelation with 12-month rating*Interpretation
------:---
Behavioral metric round+0.82Strongest promising signal
Take-home strategy deck−0.49No positive predictive evidence
Presentation−0.75High scores did not identify stronger performers
VP Engineering conversation−0.71“Technical credibility” scores did not identify stronger performers

Descriptive Pearson correlations on ordinal scores, based on nine hires—not validated predictive estimates.

The behavioral round shows a particularly clear pattern: the three hires scoring 2 all subsequently received performance ratings of 2; those scoring 3 averaged 3.33; those scoring 4 averaged 4.67. It plausibly tests something relevant to the job: whether candidates can explain their own contribution, reason about outcomes, and learn from results.

Conversely, three hires—C107, C132, and C148—received presentation and VP Engineering scores of 4, but later performance ratings of 2. C101 scored 2 in both those rounds and subsequently received a 5. We should stop treating polish or this version of technical credibility as sufficient evidence of PM effectiveness.

We have no scored evidence for the recruiter screen, and cannot assess the independent value of panel discussion or hiring-manager judgment from these files.

How confident should we be?

Confident about the observed pattern; low confidence that it will generalize.

  • Nine outcomes are too few to estimate stable predictive relationships or optimize weights and cutoffs. A few cases could materially change the results.
  • Selection limits the analysis. We observe performance only for people our existing process hired. Rejected and withdrawn candidates are not failed PMs; their missing outcomes prevent us from measuring false negatives or assessing validity across all finalists.
  • Background and scores are entangled. The four big-tech hires averaged 2.25 in performance; the three startup hires averaged 4.67. Background also overlaps with interview-score patterns. We cannot isolate interview effects from background, assignments, support, or other differences—and should not turn this into a background preference.
  • The rounds are not independent. The deck and presentation share content, while ratings may reflect overlapping impressions. We cannot establish which round adds value beyond another.
  • Performance ratings are an imperfect criterion. We lack role scope, manager, onboarding, and rating-calibration information. The company average is not a matched comparison group.

Negative correlations therefore do not mean we should prefer poor presenters or candidates with low technical scores. They mean the current versions have not demonstrated useful positive signal.

The loop I would run next quarter

1. Structured recruiter screen. Confirm role expectations, relevant experience, motivation, and logistics using consistent questions. Explain the process and preparation requirements upfront.

2. Standardized product work sample, replacing the deck and presentation. Use a fictional product brief and small dataset, with a roughly 60–90-minute total candidate commitment. Ask candidates to diagnose a problem, choose an outcome metric, prioritize options, and explain tradeoffs. Introduce a new constraint during discussion to test adaptation, not slide polish. Use equivalent prompts, consistent preparation conditions, and accommodations where needed.

3. Structured behavioral metric interview. Keep the strongest observed signal. Probe the baseline, candidate’s personal actions, alternatives considered, attribution versus correlation, unintended effects, and lessons learned. Score ownership, analytical judgment, and learning separately using behaviorally anchored rubrics.

4. Structured engineering collaboration interview. Replace the standalone credibility conversation with a realistic scenario involving delivery constraints, instrumentation, and competing priorities. An engineering interviewer scores clarification, tradeoffs, and collaboration—not jargon, pedigree, or conversational rapport.

5. Evidence-led decision. Set role-specific competencies and minimum requirements before interviewing. Interviewers submit independent scores and concrete evidence before discussion. The hiring manager documents the decision and any override. Presentation polish must not compensate for missing evidence on critical competencies. Do not derive mechanical cutoffs from these nine hires.

How we will know whether this is better

Preserve round-level scores, interviewer identities, decisions, overrides, withdrawals, and candidate feedback. Define consistent, role-adjusted performance measures at 3, 6, and 12 months; where feasible, have reviewers assess performance without seeing interview scores.

Review candidate burden and scoring consistency immediately. Treat outcomes from the next four hires as additional evidence—not proof of success. Build a larger cohort before claiming predictive validity or tuning numerical weights.

Grades and run details

Decision model 92 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly20%
  • passAddresses the actual decision88%
  • passRespects explicit constraints64%
  • passIdentifies material uncertainty96%
  • passAvoids unsupported claims69%
  • passProduces the required deliverable86%
  • passFinds which rounds line up with performance93%
  • passSays how little nine hires can show100%
  • partialSpots the interviewer pattern24%
  • passJudges on evidence, not presence100%
  • partialDefines good for this role first87%
  • passKeeps each judgement independent99%
Run
Run
#1
API response time
45 s
Submitted
6 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyMixedMixedRight
GPT-6 Luna · API

All factual claims are taken directly from the CSV or derived by arithmetic, with no invented facts.

Opus 5.5 · Claude

The take-home and presentation averages are miscalculated (e.g., take-home 4s average 2.75 not 2.5), and the VP Engineering 4 count is 8 of 11, not 9 of 11.

GPT-6.1 Sol · API

All claims about the current situation are directly supported by the supplied context or arithmetic on it.

Avoids unsupported claimsRightWrongRight
GPT-6 Luna · API

Labels correlations as exploratory, cautions against overinterpretation, and does not present hypotheses as fact.

Opus 5.5 · Claude

Presents incorrect averages and counts as established data without labelling them as estimates; the errors are factual, not interpretive.

GPT-6.1 Sol · API

Correlations are labelled as descriptive and not validated; interpretations are hedged with 'plausibly' and 'not demonstrated useful positive signal'.

Spots the interviewer patternWrongRightRight
GPT-6 Luna · API

Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.

Opus 5.5 · Claude

Notes VP Engineering scores big-tech candidates higher and those hires rated lower, and proposes a redesigned, rubric-based round with pass/concern scoring.

GPT-6.1 Sol · API

It notes VP Engineering scores big-tech candidates higher and those hires rated lower, and proposes a scored collaboration exercise with a defined bar.

Defines good for this role firstWrongWrongRight
GPT-6 Luna · API

Does not define what good looks like for this role (competencies, strong/weak signals) before designing the loop.

Opus 5.5 · Claude

Does not define the competencies that matter most for this PM role at Kestrel before designing the loop; the rubric is partial but not a role-level bar.

GPT-6.1 Sol · API

It names product judgment, metric ownership, and cross-functional execution as core competencies, and requires setting role-specific competencies and behaviorally anchored rubrics before interviewing.

All got right 8

Addresses the actual decisionRightRightRight
GPT-6 Luna · API

Commits to a specific loop for next quarter and frames it as a test, not a proven predictor.

Opus 5.5 · Claude

Commits to a specific new loop for next quarter and says to measure and review, framed for Elif.

GPT-6.1 Sol · API

The memo commits to a specific loop for next quarter, framed for Elif, and says to treat next hires as additional evidence, not proof.

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

Memo format, addressed to Elif, well under 1,000 words, and covers the requested points.

Opus 5.5 · Claude

Memo form, addressed to Elif, appears under 1,000 words, and proposes enforceable changes.

GPT-6.1 Sol · API

The output is a memo under 1,000 words addressed to Elif, respecting the form and length constraint.

Identifies material uncertaintyRightRightRight
GPT-6 Luna · API

Names small sample, range restriction, missing data, and proposes collecting outcomes from the next cohort to resolve.

Opus 5.5 · Claude

Names small sample, only-hired bias, background confound, single-rater ratings, and interviewer effects; says what would change the call (more data, measurement).

GPT-6.1 Sol · API

It names small sample, selection bias, background entanglement, and imperfect ratings as unknowns, and says to resolve by tracking future hires.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

Complete memo with analysis, confidence, and a concrete loop; a PM could act on it with light edits.

Opus 5.5 · Claude

Complete memo with summary, analysis, confidence, and actionable loop; a PM could act on the structure.

GPT-6.1 Sol · API

The memo answers which parts predict, how confident, and the loop to run, and is usable by the VP Product with light edits.

Finds which rounds line up with performanceRightRightRight
GPT-6 Luna · API

Works through each round's scores against 12-month ratings, identifying behavioural as positive and the others as negative.

Opus 5.5 · Claude

Compares each round's scores with 12-month ratings and correctly identifies behavioural as positive, take-home/presentation/VP Eng as negative in direction.

GPT-6.1 Sol · API

It compares each round's scores with 12-month ratings, finding behavioral lines up and take-home, presentation, VP Engineering do not.

Says how little nine hires can showRightRightRight
GPT-6 Luna · API

States plainly that eight hires is far too few, and proposes measuring the next four hires to test the changes.

Opus 5.5 · Claude

Explicitly states nine hires is thin, calls it a hint not proof, and proposes measuring the next hires.

GPT-6.1 Sol · API

It states plainly that nine hires is too few, frames changes as evidence-informed not proven, and proposes to keep measuring.

Judges on evidence, not presenceRightRightRight
GPT-6 Luna · API

Proposes probing specific metrics, evidence of impact, and structured scoring, not impressions or polish.

Opus 5.5 · Claude

Proposes behavioural prompts on specific actions and learning, a rubric on substance not delivery, and a working session scored on reasoning, not polish.

GPT-6.1 Sol · API

It asks for concrete evidence of personal contribution, learning, and tradeoffs, and treats impressions as weak signals.

Keeps each judgement independentRightRightRight
GPT-6 Luna · API

Requires interviewers to score independently before discussion, with disagreement recorded.

Opus 5.5 · Claude

Requires written scores and evidence before debrief, and independent scoring by each interviewer.

GPT-6.1 Sol · API

It requires interviewers to submit independent scores and evidence before discussion, and the hiring manager to document any override.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review