Tasks / Leadership

Hire a PM

Can the model design a hiring process that finds the right PM, and make the call on real candidates from the evidence?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 79% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    Complete recommendation memo with evidence, risks, and actionable next steps, usable as is.
    GPT-6 Luna · API · Two finalists, one Group PM role
  2. Judges on evidence, not presence100% pass
    Judges on concrete evidence (metric moved, team credit, learning from failure) and treats impressions like presence as weak signals.
    GPT-6 Luna · API · Two finalists, one Group PM role
  3. Tests what the last hire failed at100% pass
    The loop includes an influence simulation with the Head of Sales and a scorecard dimension on influence without authority, directly testing what the last hire failed at.
    GPT-6.1 Sol · API · A loop for the first growth PM

Where it slips

  1. Spots the interviewer pattern65% pass
    Does not notice that VP Engineering scores big-tech candidates higher and those hires were rated lower; only reports a negative correlation without the pattern.
    GPT-6 Luna · API · What our interviews predict
  2. Avoids unsupported claims70% pass
    The claim 'A growth PM in a senior role is likely to be exactly that kind of person' is presented as fact without evidence and is not labelled as an assumption.
    Opus 5.5 · Claude · A loop for the first growth PM
  3. Identifies material uncertainty72% pass
    The output does not name unknowns that could change the hiring decision or the loop design, nor does it say how they would be resolved.
    Sonnet 5.5 · API · A loop for the first growth PM

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Judges on evidence, not presence

    Does the output judge candidates, or design the process to judge them, on specific evidence of results they drove and how they worked with their team, rather than impressions such as presence, confidence or polish?

    Passes when Asks for or weighs concrete evidence: the metric a candidate moved, the part they played, how they shared credit, what they learned from a failure. Impressions are treated as weak signals.

  2. Defines good for this role first

    Does the output set out what good looks like for this specific role (the competencies that matter most at this level, and what strong and weak look like) before judging candidates or designing interviews?

    Passes when Names the competencies weighted for this role and level, with strong and weak signals, and the judgement or the process follows from them.

  3. Keeps each judgement independent

    Does the output keep each interviewer's judgement independent (written down before the group discusses it), and guard against the loudest voice or the first speaker deciding?

    Passes when Requires or relies on written, independent evaluations before discussion, and discounts views that changed under group pressure.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Director of Product at Kestrel. Elif Demir, our VP Product, wants to fix the PM interview loop before we hire four more PMs next quarter (her note is below). Using the data we have, write her a memo of no more than 1,000 words: which parts of the loop seem to predict how PMs do once hired and which don't, how confident we can be, and the loop you'd run next quarter.\n\nThe files are attached.

What the model was given3 items: elif_note.md, current_loop.md, pm_candidates_last_18_months.csv (every PM candidate who reached the final rounds)
elif_note.md1 lines · Download
From Elif (VP Product): "We've made three PM hires in 18 months who struggled. I want to know what in our loop is and isn't working, using the data we have, and what to change before we hire four more PMs next quarter."
current_loop.md10 lines · Download
# Current PM interview loop

1. Recruiter screen.
2. Take-home: a product strategy deck, one week to complete.
3. Presentation of the take-home to a panel.
4. Behavioural round: "Tell me about a metric you moved", with follow-ups on what you did and what you learned.
5. VP Engineering conversation ("technical credibility").
…
pm_candidates_last_18_months.csv (every PM candidate who reached the final rounds)candidate,background,takehome_score,presentation_score,behavioural_metric_round_score,vp_engineering_score,outcome,rating_after_12_months C100,Big tech,4,4,3,2,Withdrew, C101,Startup,3,2,4,2,Hired,5 C102,Big tech,1,2,3,4,Rejected, C103,Agency,1,3,2,3,Rejected, C104,Big tech,4,3,3,4,Hired,3 C105,Startup,4,2,1,1,Rejected, C106,Startup,4,2,1,3,Rejected, C107,Big tech,4,4,2,4,Hired,2 C108,Consulting,1,4,2,1,Rejected, C109,Agency,4,2,1,2,Rejected, C110,Consulting,3,1,2,2,Rejected, C111,Big tech,2,2,1,4,Rejected, C112,Agency,3,1,2,1,Rejected, C113,Startup,4,2,3,1,Withdrew, C114,Big tech,2,2,3,2,Withdrew, C115,Consulting,3,1,1,1,Rejected, C116,Startup,2,2,1,1,Rejected, C117,Startup,2,1,4,3,Withdrew, C118,Startup,2,2,1,2,Withdrew, C119,Consulting,1,1,3,3,Rejected, C120,Consulting,2,4,1,3,Rejected, C121,Consulting,2,3,4,1,Rejected, C122,Agency,4,1,3,3,Withdrew, C123,Agency,2,4,4,1,Rejected, C124,Big tech,2,2,2,3,Rejected, C125,Startup,2,2,1,3,Rejected, C126,Consulting,4,1,2,2,Rejected, C127,Consulting,1,3,3,1,Rejected, C128,Startup,3,3,2,3,Rejected, C129,Agency,3,3,3,1,Rejected, C130,Startup,1,3,2,2,Rejected, C131,Agency,1,2,4,2,Rejected, C132,Big tech,3,4,2,4,Hired,2 C133,Agency,2,4,2,1,Rejected, C134,Consulting,4,4,3,3,Hired,3 C135,Consulting,3,4,4,3,Rejected, C136,Consulting,2,2,3,2,Rejected, C137,Consulting,3,4,3,1,Withdrew, C138,Startup,2,4,4,2,Withdrew, C139,Startup,3,2,2,3,Rejected, C140,Startup,1,1,3,3,Rejected, C141,Startup,3,3,4,3,Hired,5 C142,Big tech,4,2,3,4,Withdrew, C143,Consulting,1,4,3,2,Rejected, C144,Consulting,1,1,1,2,Rejected, C145,Agency,4,1,2,1,Rejected, C146,Startup,3,1,4,1,Rejected, C147,Big tech,4,1,1,4,Withdrew, C148,Big tech,4,4,2,4,Hired,2 C149,Agency,2,3,4,2,Hired,4 C150,Consulting,4,1,1,2,Rejected, C151,Startup,2,3,3,3,Hired,4
What a strong answer doesThe answer key the graders mark against

Uses the nine hires to compare each round with 12-month ratings: the behavioural 'metric you moved' round lines up strongly with how hires performed (higher scores, better ratings), while the take-home, the presentation and the VP Engineering conversation line up the wrong way: the hires they scored highest are the ones rated 2. The VP Engineering scores big-tech candidates higher, and all four big-tech hires were rated 2 or 3. Is explicit that nine hires is far too few to conclude much: the pattern is a reason to change and then measure, not proof. Proposes a loop with a scorecard defined up front, the behavioural round strengthened (more than one, structured), the week-long take-home replaced or cut down (it costs candidates, and some withdraw), the VP Engineering round given a defined bar and calibrated, and every interviewer writing a score before the debrief. Sets up tracking so the next hires test the change.

Critical failures (cap the score at 40)
  • Recommends keeping the take-home and presentation as the main deciding rounds
Task

Staff-level · v1.0 · synthetic · hiring, data files, staff-level

The outputs

5 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit4 of 12 checks passed

The output contains factual errors (number of hires, number of withdrawals) and unsupported causal claims, so a PM would need to re-check the evidence and redo parts of the reasoning.

Output

MEMORANDUM

TO: Elif Demir, VP of Product FROM: Director of Product DATE: October 24, 2023 SUBJECT: PM Interview Loop Analysis & Recommendations for Q1 Hiring

*

Elif, you are entirely right to flag this. Looking back at our last 18 months of hiring data, out of our 6 hired PMs who completed their first year, half (C107, C132, C148) are struggling with 12-month performance ratings of 2 (well below our company average of 3.4). Meanwhile, our top performers (C101, C141 with ratings of 5, and C149 with a rating of 4) represent the standard we actually need.

To prep for hiring four more PMs next quarter, I’ve analyzed the data for all 52 candidates who reached our final rounds. Here is what works, what is completely broken, how confident we can be, and how we should redesign the loop.

---

1. What Predicts Performance? (And What Doesn't)

By analyzing the correlation between interview stage scores and 12-month post-hire ratings for our 6 completed hires, we see stark contrasts in signal quality:

* Take-Home Strategy Deck (Zero to Negative Correlation): Our take-home exam is a massive waste of candidate time and offers zero predictive value. Our struggling hires (C107, C132, C148) averaged a 3.7 out of 4 on the take-home. Conversely, our top performers averaged a 2.7. High scores on a 1-week unsupervised deck merely signal that a candidate has free time, a supportive current employer, or an agency background—not that they can execute at Kestrel. * Take-Home Presentation Panel (Weak / Misleading Signal): Similar to the take-home itself, presentation scores do not align with success. C107 and C148 both scored 4s on their presentation and turned out to be 12-month "2" ratings. * VP Engineering Conversation (Inverted Signal): Surprisingly, the VP Engineering score shows a negative relationship with success. Candidates who scored high with Engineering often struggled on the job, while our top performers (C101, C141, C149) mixed moderate-to-low engineering scores (2 to 3) with high marks elsewhere. While technical credibility matters, our current technical screen appears to favor academic/theoretical answers over pragmatic product-engineering partnership. * Behavioural Metric Round (Our Only Real Signal): The behavioural round ("Tell me about a metric you moved") is the only component in our current loop that correlates with on-the-job success. * Our three struggling hires averaged 2.0 on this round. * Our three successful hires averaged 4.0 on this round. This is our sole bright spot. It tests past execution, ownership, and learning—things that actually matter day-to-day at Kestrel.

---

2. Statistical Confidence

We must temper our confidence. We have only 6 data points of complete hire-and-tenure data. While the directional trend on the behavioural round is unmistakable (+2.0 delta between high and low performers), running statistical significance tests on $n=6$ is risky.

However, we can look at broader hiring funnel data across all 52 candidates to understand drop-offs. Notably, 13 candidates withdrew from the process, heavily skewed toward candidates who scored well on the take-home (e.g., C100, C113, C142, C147 all scored 4s on the take-home and walked away). Our heavy, one-week take-home is actively driving away top-tier passive talent who refuse to do free consulting work for a week, leaving us to select from a pool of people willing to endure it.

---

3. The Proposed Loop for Next Quarter

To scale our team to four new PMs without repeating past mistakes, we need to strip away vanity metrics (the take-home deck) and lean heavily into what actually predicts success: rigorous behavioural assessment of past execution, combined with a real-time collaborative product session.

Here is the 4-stage loop I propose we deploy immediately:

  1. Recruiter Screen (30 mins): Standard fit, compensation alignment, and basic background check.
  2. Product Craft & Execution Interview (45 mins - Live Video): Replaces the take-home. Instead of a week-long deck, we give the candidate a real Kestrel product challenge 48 hours in advance (requiring no more than 2 hours of prep) and spend the interview deep-diving into their hypotheses, prioritization trade-offs, and customer empathy.
  3. Behavioural & Metric Ownership Round (45 mins): Double down on what works. Retain our behavioural round focusing on metrics moved, but expand the rubric to explicitly test for accountability, cross-functional conflict resolution, and handling failure.
  4. Product-Engineering Partnership Round (45 mins): Refocused Engineering screen. Instead of "technical credibility" evaluated in a vacuum, this is a joint session with the VP of Engineering or a Senior Tech Lead to evaluate how the PM collaborates with engineering to ship product, scope down MVPs, and manage technical debt.

Decision Gate: Drop the consensus panel meeting where the loudest voice wins. The hiring manager makes the final call, but only if the candidate scores a predefined threshold on the Behavioural and Product Craft rounds.

Summary

By eliminating the 1-week take-home, we will stop scaring away top-tier talent, shorten our time-to-hire, and refocus our interview panel on real-world execution rather than polished slide decks. Let me know if you are aligned, and I will draft the new interview rubrics by Friday.

Claims the judge couldn’t find in the brief

  • There are 6 hired PMs who completed their first year.
  • 13 candidates withdrew from the process.
  • High scores on a 1-week unsupervised deck merely signal that a candidate has free time, a supportive current employer, or an agency background.
  • Our heavy, one-week take-home is actively driving away top-tier passive talent who refuse to do free consulting work for a week.
  • The VP Engineering conversation shows a negative relationship with success and appears to favor academic/theoretical answers over pragmatic product-engineering partnership.
  • The behavioural metric round is the only component that correlates with on-the-job success.
  • We have only 6 data points of complete hire-and-tenure data.

Grades and run details

Decision model 54 · LLM judge 4 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly68%
  • passAddresses the actual decision29%
  • partialRespects explicit constraints25%
  • partialIdentifies material uncertainty86%
  • failAvoids unsupported claims59%
  • passProduces the required deliverable28%
  • passFinds which rounds line up with performance78%
  • partialSays how little nine hires can show76%
  • partialSpots the interviewer pattern94%
  • passJudges on evidence, not presence66%
  • failDefines good for this role first26%
  • partialKeeps each judgement independent66%
Run
Run
#1
Time to output
7 s
Submitted
5 Oct 2026

Needs more than a quick edit9 of 12 checks passed

The take-home correlation is miscalculated, so a PM would need to re-check the data analysis before using the memo. The error undermines a key piece of evidence.

Output

To: Elif Demir, VP Product From: Director of Product Subject: PM interview loop: what the evidence supports and what to change

Recommendation

For next quarter’s four hires, retain and standardise the behavioural metric round; replace the week-long strategy deck and presentation with a bounded, realistic work sample; and replace “technical credibility” with a structured engineering-partnership interview.

The evidence supports changing what we reward—not claiming we have discovered a validated hiring formula.

What the data says

We have 52 finalists: nine hired, 33 rejected and ten withdrawn. Only the nine hires have performance outcomes. Their average 12-month rating is 3.33, close to the company PM average of 3.4. Three received a 2; none received a 1. That average masks the individual misses Elif is concerned about, but does not establish that the loop underperforms an alternative.

Among those nine hires:

Interview componentCorrelation with 12-month rating*Interpretation
------:---
Behavioural metric round+0.94Strongest positive signal in this sample
Take-home strategy deck−0.57Higher scores did not identify better performers
Presentation−0.87Strong negative association in this sample
VP Engineering conversation−0.82Higher “technical credibility” scores did not identify better performers

Descriptive Pearson correlations on ordinal scores; not validated predictive estimates.

The behavioural pattern is particularly clear: the three hires scoring 2 all received performance ratings of 2; those scoring 3 averaged 3.33; those scoring 4 averaged 4.67. Probing actual ownership, measurement and learning appears promising.

Conversely, all three hires rated 2 had scored 4 on both presentation and VP Engineering. Two also scored 4 on the take-home. These rounds clearly failed to screen out our weakest performers. They may reward polish, familiarity or pedigree rather than effective PM work—but the data does not establish that mechanism.

We cannot evaluate the recruiter screen: it has neither scores nor data on candidates eliminated before finals. Nor can we isolate the value of panel discussion or hiring-manager judgment without decision rationales.

How confident should we be?

Enough to change the process; not enough to treat these correlations as durable prediction.

There are four important limitations:

  • Nine outcomes are very few. One or two different hires could materially change the estimates. We cannot credibly estimate each round’s independent contribution.
  • We observe performance only for selected candidates. Hiring already depended on these interviews. Conditioning on selection can distort associations; we do not know how rejected or withdrawn candidates would have performed.
  • Background and scores are entangled. All four big-tech hires received VP Engineering scores of 4 and averaged 2.25 in performance; the three startup hires averaged 4.67. That does not establish a background effect—or justify preferring startup candidates. It prevents us from separating interview signal, background and role fit.
  • Performance ratings are imperfect outcomes. We lack information on manager calibration, assignment difficulty, onboarding and opportunities to deliver results.

Accordingly, I would prioritise the behavioural round, but would not reverse-score presentations or engineering interviews, adopt background preferences, or derive numerical weights from nine people. The negative correlations show that those scores were not useful positive signals here; they do not prove the underlying skills are unimportant.

The loop I would run next quarter

1. Structured recruiter screen — 25 minutes. Check role requirements, relevant scope, motivation and practical constraints against a common checklist. Do not use employer prestige as a proxy for capability. Explain the process and time commitment.

2. Realistic work sample — 75 minutes, replacing both deck and presentation. Give every candidate the same bounded, fictional product problem, customer evidence, metric definitions and engineering constraints. Ask them to diagnose the problem, prioritise a response, define success and revise their recommendation when new evidence arrives. No slide deck or advance homework.

Score problem framing, evidence use, prioritisation, measurement and adaptation. Assess communication through clarity of reasoning, not presentation polish. Offer reasonable accommodations and equivalent administration conditions.

3. Structured behavioural metric interview — 50 minutes. Retain the current core question, with consistent probes: baseline, personal contribution, alternatives considered, causal evidence, trade-offs, setbacks and learning. Use a second example to distinguish repeatable judgment from one well-rehearsed story. Anchor scores to demonstrated ownership and reasoning, not the absolute size of a metric increase.

4. Engineering partnership interview — 40 minutes. An engineering leader uses a common scenario involving feasibility, scope, reliability and disagreement. Score how candidates uncover constraints and negotiate trade-offs—not jargon, coding knowledge or perceived “credibility.”

Decision process. Before interviewing, agree on competency definitions, anchored 1–4 ratings and minimum acceptable evidence for critical competencies. Interviewers submit scores and supporting observations independently before debriefing. Discuss discrepancies against the rubric; do not mechanically average scores. The hiring manager records the decision, unresolved risks and any exception to the agreed bar.

How we will learn

Record scores, assessor identity, decision reasons, candidate time and withdrawal stage for every finalist. Audit scoring consistency and progression by background and, where appropriately collected, demographic group.

For hires, collect consistently defined 3-, 6- and 12-month outcomes, with role context and performance assessment kept independent of interview scores where feasible. Review process quality after the next four hires, but accumulate more cohorts before claiming predictive validity. Four additional hires will test whether we can run a better loop—not prove that it works.

Claims the judge couldn’t find in the brief

  • Take-home strategy deck correlation with 12-month rating is -0.57.

Grades and run details

Decision model 96 · LLM judge 9 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly28%
  • passAddresses the actual decision86%
  • passRespects explicit constraints64%
  • passIdentifies material uncertainty92%
  • passAvoids unsupported claims74%
  • passProduces the required deliverable91%
  • passFinds which rounds line up with performance96%
  • passSays how little nine hires can show100%
  • passSpots the interviewer pattern96%
  • passJudges on evidence, not presence100%
  • partialDefines good for this role first76%
  • passKeeps each judgement independent99%
Run
Run
#1
Time to output
47 s
Submitted
5 Oct 2026

Usable with a quick edit12 of 12 checks passed

The memo is clear, evidence-based, and actionable; a PM could use it with minor wording tweaks and no need to re-verify the data or reasoning.

Output

To: Elif Demir, VP Product From: Director of Product Subject: PM interview loop: what the evidence supports and what to change next quarter

Recommendation

For the next four hires, make demonstrated product judgment, metric ownership, and cross-functional execution the center of the loop. Replace the week-long strategy deck and separate presentation with a short, standardized work sample; retain and strengthen the behavioral metric interview; replace the unstructured “technical credibility” conversation with a scored collaboration exercise.

These are evidence-informed changes, not a validated hiring formula. Our strongest signal comes from just nine hires.

What appears to predict performance

We have 52 finalists: nine hired, 33 rejected, and ten withdrawn. Only the nine hires have performance outcomes. Their average 12-month rating is 3.33, versus the company PM average of 3.4. Three received a 2; none received a 1. The similar averages do not negate the three struggling hires, but neither establish that the entire loop is failing.

Among hires:

Interview componentCorrelation with 12-month rating*Interpretation
------:---
Behavioral metric round+0.82Strongest promising signal
Take-home strategy deck−0.49No positive predictive evidence
Presentation−0.75High scores did not identify stronger performers
VP Engineering conversation−0.71“Technical credibility” scores did not identify stronger performers

Descriptive Pearson correlations on ordinal scores, based on nine hires—not validated predictive estimates.

The behavioral round shows a particularly clear pattern: the three hires scoring 2 all subsequently received performance ratings of 2; those scoring 3 averaged 3.33; those scoring 4 averaged 4.67. It plausibly tests something relevant to the job: whether candidates can explain their own contribution, reason about outcomes, and learn from results.

Conversely, three hires—C107, C132, and C148—received presentation and VP Engineering scores of 4, but later performance ratings of 2. C101 scored 2 in both those rounds and subsequently received a 5. We should stop treating polish or this version of technical credibility as sufficient evidence of PM effectiveness.

We have no scored evidence for the recruiter screen, and cannot assess the independent value of panel discussion or hiring-manager judgment from these files.

How confident should we be?

Confident about the observed pattern; low confidence that it will generalize.

  • Nine outcomes are too few to estimate stable predictive relationships or optimize weights and cutoffs. A few cases could materially change the results.
  • Selection limits the analysis. We observe performance only for people our existing process hired. Rejected and withdrawn candidates are not failed PMs; their missing outcomes prevent us from measuring false negatives or assessing validity across all finalists.
  • Background and scores are entangled. The four big-tech hires averaged 2.25 in performance; the three startup hires averaged 4.67. Background also overlaps with interview-score patterns. We cannot isolate interview effects from background, assignments, support, or other differences—and should not turn this into a background preference.
  • The rounds are not independent. The deck and presentation share content, while ratings may reflect overlapping impressions. We cannot establish which round adds value beyond another.
  • Performance ratings are an imperfect criterion. We lack role scope, manager, onboarding, and rating-calibration information. The company average is not a matched comparison group.

Negative correlations therefore do not mean we should prefer poor presenters or candidates with low technical scores. They mean the current versions have not demonstrated useful positive signal.

The loop I would run next quarter

1. Structured recruiter screen. Confirm role expectations, relevant experience, motivation, and logistics using consistent questions. Explain the process and preparation requirements upfront.

2. Standardized product work sample, replacing the deck and presentation. Use a fictional product brief and small dataset, with a roughly 60–90-minute total candidate commitment. Ask candidates to diagnose a problem, choose an outcome metric, prioritize options, and explain tradeoffs. Introduce a new constraint during discussion to test adaptation, not slide polish. Use equivalent prompts, consistent preparation conditions, and accommodations where needed.

3. Structured behavioral metric interview. Keep the strongest observed signal. Probe the baseline, candidate’s personal actions, alternatives considered, attribution versus correlation, unintended effects, and lessons learned. Score ownership, analytical judgment, and learning separately using behaviorally anchored rubrics.

4. Structured engineering collaboration interview. Replace the standalone credibility conversation with a realistic scenario involving delivery constraints, instrumentation, and competing priorities. An engineering interviewer scores clarification, tradeoffs, and collaboration—not jargon, pedigree, or conversational rapport.

5. Evidence-led decision. Set role-specific competencies and minimum requirements before interviewing. Interviewers submit independent scores and concrete evidence before discussion. The hiring manager documents the decision and any override. Presentation polish must not compensate for missing evidence on critical competencies. Do not derive mechanical cutoffs from these nine hires.

How we will know whether this is better

Preserve round-level scores, interviewer identities, decisions, overrides, withdrawals, and candidate feedback. Define consistent, role-adjusted performance measures at 3, 6, and 12 months; where feasible, have reviewers assess performance without seeing interview scores.

Review candidate burden and scoring consistency immediately. Treat outcomes from the next four hires as additional evidence—not proof of success. Build a larger cohort before claiming predictive validity or tuning numerical weights.

Grades and run details

Decision model 92 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly20%
  • passAddresses the actual decision88%
  • passRespects explicit constraints64%
  • passIdentifies material uncertainty96%
  • passAvoids unsupported claims69%
  • passProduces the required deliverable86%
  • passFinds which rounds line up with performance93%
  • passSays how little nine hires can show100%
  • partialSpots the interviewer pattern24%
  • passJudges on evidence, not presence100%
  • partialDefines good for this role first87%
  • passKeeps each judgement independent99%
Run
Run
#1
API response time
45 s
Submitted
6 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 8

Uses the supplied evidence correctlyWrongMixedRight
Gemini 3.5 Flash-Lite · Gemini

The output misstates the number of hires (6 instead of 9) and the number of withdrawals (13 instead of 10), and presents unsupported causal claims as fact.

GPT-6 Astra · ChatGPT

The take-home correlation is stated as -0.57, but the data yields approximately -0.33; this fact about the current situation is incorrect.

GPT-6.1 Sol · API

All claims about the current situation are directly supported by the supplied context or arithmetic on it.

Addresses the actual decisionMixedRightRight
Gemini 3.5 Flash-Lite · Gemini

The output does not say what result or condition would change the recommended loop, as required for a committed answer.

GPT-6 Astra · ChatGPT

The memo commits to a clear recommendation (retain behavioural, replace take-home/presentation and VP Eng round) early and unambiguously for Elif.

GPT-6.1 Sol · API

The memo commits to a specific loop for next quarter, framed for Elif, and says to treat next hires as additional evidence, not proof.

Identifies material uncertaintyWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It does not name specific unknowns that could change the decision or say how they would be resolved; it only notes the small sample size (incorrectly) without bounding uncertainty.

GPT-6 Astra · ChatGPT

It names specific unknowns (small sample, selection bias, background entanglement, imperfect ratings) and says more cohorts are needed to validate.

GPT-6.1 Sol · API

It names small sample, selection bias, background entanglement, and imperfect ratings as unknowns, and says to resolve by tracking future hires.

Avoids unsupported claimsWrongMixedRight
Gemini 3.5 Flash-Lite · Gemini

It presents interpretations (e.g., what take-home scores signal, why candidates withdraw) as established fact without labelling them as hypotheses.

GPT-6 Astra · ChatGPT

The take-home correlation of -0.57 is presented as a descriptive fact but is not supported by the data; this is an unsupported claim.

GPT-6.1 Sol · API

Correlations are labelled as descriptive and not validated; interpretations are hedged with 'plausibly' and 'not demonstrated useful positive signal'.

Says how little nine hires can showWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It misstates the sample size as 6 instead of 9 and does not propose how to keep measuring the new loop to test the changes.

GPT-6 Astra · ChatGPT

It states plainly that nine hires is very few, frames changes as a test, and proposes continued measurement.

GPT-6.1 Sol · API

It states plainly that nine hires is too few, frames changes as evidence-informed not proven, and proposes to keep measuring.

Spots the interviewer patternWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It does not name the pattern that VP Engineering scores big-tech candidates higher and those hires were rated lower; it only offers a vague interpretation.

GPT-6 Astra · ChatGPT

It notes VP Engineering scores big-tech candidates higher and those hires rated lower, and proposes a structured engineering-partnership interview with a defined bar.

GPT-6.1 Sol · API

It notes VP Engineering scores big-tech candidates higher and those hires rated lower, and proposes a scored collaboration exercise with a defined bar.

Defines good for this role firstWrongWrongRight
Gemini 3.5 Flash-Lite · Gemini

It does not define the competencies that matter most for this role and level before designing the interview loop.

GPT-6 Astra · ChatGPT

The memo says to agree on competency definitions but does not name the specific competencies or what good looks like for this role before designing the interviews.

GPT-6.1 Sol · API

It names product judgment, metric ownership, and cross-functional execution as core competencies, and requires setting role-specific competencies and behaviorally anchored rubrics before interviewing.

Keeps each judgement independentWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

It drops the panel meeting but does not require written, independent evaluations before any discussion to guard against group influence.

GPT-6 Astra · ChatGPT

It requires interviewers to submit scores and observations independently before debriefing and to discuss discrepancies against a rubric.

GPT-6.1 Sol · API

It requires interviewers to submit independent scores and evidence before discussion, and the hiring manager to document any override.

All got right 4

Respects explicit constraintsRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The output is a memo to Elif, from the Director of Product, and is well under 1,000 words.

GPT-6 Astra · ChatGPT

The output is a memo to Elif, under 1,000 words, and addresses the requested decision.

GPT-6.1 Sol · API

The output is a memo under 1,000 words addressed to Elif, respecting the form and length constraint.

Produces the required deliverableRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The memo is complete, in the requested form, for the named reader, and could be acted on with light edits.

GPT-6 Astra · ChatGPT

The memo is complete, in the correct form, for the named reader, and within the word limit.

GPT-6.1 Sol · API

The memo answers which parts predict, how confident, and the loop to run, and is usable by the VP Product with light edits.

Finds which rounds line up with performanceRightRightRight
Gemini 3.5 Flash-Lite · Gemini

It compares each round's scores with 12-month ratings and correctly identifies that the behavioural round lines up while the others do not, for the hires it examines.

GPT-6 Astra · ChatGPT

It compares each round's scores with 12-month ratings and correctly identifies behavioural as lining up and the others as not.

GPT-6.1 Sol · API

It compares each round's scores with 12-month ratings, finding behavioral lines up and take-home, presentation, VP Engineering do not.

Judges on evidence, not presenceRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The proposed behavioural round asks for concrete evidence of metrics moved, accountability, and handling failure, not impressions.

GPT-6 Astra · ChatGPT

The proposed process scores concrete evidence (problem framing, evidence use, ownership, reasoning) and explicitly avoids rewarding polish or jargon.

GPT-6.1 Sol · API

It asks for concrete evidence of personal contribution, learning, and tradeoffs, and treats impressions as weak signals.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6.1 SolwithAPI94.4100.03None
2Sonnet 5.5withAPI95.896.22None
3Gemini 3.8 FlashwithAPI89.692.32None
4GPT-6 AstrawithChatGPT95.887.23None
5Opus 5.5withClaude91.787.23None
6GPT-6 LunawithAPI93.184.63None
7Gemini 3.5 Flash-LitewithGemini75.064.13None

About the task

The PM job

Hiring product managers.

Why it matters

A bad PM hire costs a team a year. Most loops reward polish and presence, and the debrief goes to whoever speaks first.

What good looks like

  • Defines what good looks like for this role before judging anyone
  • Judges on evidence of results and how they work with a team
  • Tests the real job, not a rehearsed framework
  • Keeps each interviewer's judgement independent

Deliberately not measured

  • Employment law
  • Compensation benchmarking
Capability tested

Hiring judgement

The failure we’re looking for

Hiring for presence and polish over evidence

Grading

Decision model and LLM judge, calibrated against a blind PM review