Needs more than a quick edit9 of 12 checks passed
The memo contains arithmetic errors in the take-home and presentation averages and the VP Engineering count, so a PM would need to re-check the data before using it.
Output
To: Elif Demir, VP Product From: Director of Product Re: What our PM loop predicts, and the loop for next quarter
Summary
Of the five stages in our loop, only the behavioural "metric you moved" round tracks how hires perform after 12 months. The take-home, the presentation and the VP Engineering conversation all point the wrong way: hires who scored higher on them performed worse. The evidence is thin (nine hires), so I'd treat this as a strong hint rather than proof. Even so, it is consistent enough to change the loop now and to measure the new loop properly.
What the data shows
Over 18 months, 52 candidates reached the final rounds. We hired 9, rejected 33, and 10 withdrew. Our hires averaged 3.3 at 12 months, slightly below the PM average of 3.4. The three who struggled are C107, C132 and C148, all rated 2.
Average 12-month rating of hires, by interview score
| Stage | Hires who scored 4 | Hires who scored 3 | Hires who scored 2 | Correlation with rating |
|---|---|---|---|---|
| Behavioural (metric moved) | 4.7 (n=3) | 3.3 (n=3) | 2.0 (n=3) | +0.94 |
| Take-home | 2.5 (n=4) | 4.0 (n=3) | 4.0 (n=2) | −0.57 |
| Presentation | 2.25 (n=4) | 4.0 (n=4) | 5.0 (n=1) | −0.87 |
| VP Engineering | 2.25 (n=4) | 4.0 (n=3) | 4.5 (n=2) | −0.82 |
The three strugglers share an identical profile:
- 4 on the presentation
- 4 on VP Engineering
- 3 or 4 on the take-home
- 2 on the behavioural round
- Big-tech background
Our two strongest hires, C101 and C141 (both rated 5), had middling presentation and VP Engineering scores but scored 4 on the behavioural round.
In other words, we have been hiring for polish and for technical fluency, and overriding the one signal that worked.
Background is tangled up in this. All four big-tech hires were rated 2 or 3, averaging 2.25. The five others averaged 4.2. The VP Engineering round looks especially exposed to this:
- Every VP Engineering score of 4 in the dataset went to a big-tech candidate (9 of 11 big-tech candidates).
- No other candidate scored 4.
The round may be measuring familiarity with big-tech engineering culture rather than how well someone will work with our engineers.
The behavioural result holds within each group. This is the most reassuring part of the analysis.
- Among non-big-tech hires, the behavioural score still tracks performance. Those who scored 4 were rated 5, 5 and 4. Those who scored 3 were rated 3 and 4.
- Among big-tech hires, the one who scored 3 was rated 3 and the three who scored 2 were rated 2.
So the behavioural signal isn't simply a stand-in for background. Within non-big-tech hires, the presentation and take-home scores still lean negative.
How confident we can be
I'm only moderately confident, for five reasons:
- Nine data points. One or two different outcomes could change the picture, and none of these correlations would survive a strict significance test.
- We only see outcomes for people we hired. Hires were selected partly on high presentation and VP Engineering scores, so we can't see how low scorers on those stages would have done. We also don't know about strong behavioural candidates we rejected (C121, C123, C131, C135, C146), or about withdrawals. C135, for example, was rejected despite scoring 3, 4, 4 and 3, among the strongest sets of scores in the pool.
- Background is confounded. With four big-tech hires, I can't fully separate "big-tech PMs struggle here" from "our loop rewards the wrong things in big-tech PMs." Either way, the fix is to the instruments, not to filter on background.
- 12-month ratings are one manager's judgement, and they also reflect team and scope.
- Interviewer effects. We don't know who scored which candidates, so some of this may come from individual raters rather than from the stages themselves.
What I am confident about is narrower. The take-home, presentation and VP Engineering scores give us no evidence that they help. Given what they cost candidates and the panel, that is reason enough to stop relying on them.
The loop I'd run next quarter
1. Recruiter screen. Unchanged.
2. Structured behavioural round, now the core of the loop (two interviewers, about 60 minutes). - Keep "a metric you moved," and add a second prompt: "a decision you got wrong and what you changed." - Use a written rubric that scores: - specificity about the candidate's own actions; - how they diagnosed the problem; - the trade-offs they made; - what they learned. - Score the substance, not the delivery. Each interviewer scores independently.
3. Replace the week-long take-home and presentation with a 90-minute live working session. - Give the candidate a real, anonymised Kestrel problem with messy data, and work through it with a PM and an engineer. - Score how they reason, prioritise and respond to pushback, not how finished the output looks. - This also cuts the candidate time we ask for. Ten of 52 finalists withdrew; I can't say the week-long take-home caused that, but it won't have helped.
4. Redesign the VP Engineering round as a structured "working with engineering" round. - Base it on a scenario such as a scope cut, a technical-debt trade-off or an estimate dispute, with a rubric agreed with Engineering. - Score it pass/concern only until we have evidence that it predicts anything, so it can flag real problems without rewarding big-tech vocabulary.
5. Change how we decide. - All scores and written evidence go in before the debrief, so nobody anchors on the loudest voice. - No hire with a behavioural score below 3. - The hiring manager still decides. If they hire against the behavioural score, they write down why.
6. Measure it this time. - Log every interviewer's scores, not just the panel's. - Collect 6-month and 12-month ratings against a consistent rubric. - With four more hires we'll have 13 to analyse. We should review the data in two quarters, not two years. - I'd also like a quick look back at the strong behavioural candidates we rejected, to see whether the debrief was systematically discounting that round.
The ask
I'd like your sign-off to:
- pilot this loop for the four upcoming hires;
- get two hours of interviewer calibration on the new rubrics;
- secure the VP Engineering's agreement to redesign their round.
I can have the rubrics and the working-session problem ready within three weeks.
Claims the judge couldn’t find in the brief
- Take-home: hires scoring 4 averaged 2.5, 3 averaged 4.0, 2 averaged 4.0.
- Presentation: hires scoring 4 averaged 2.25, 3 averaged 4.0, 2 averaged 5.0.
- Every VP Engineering score of 4 in the dataset went to a big-tech candidate (9 of 11 big-tech candidates).
Grades and run details
Decision model 92 · LLM judge 9 of 13 checks
Decision model checks
- passUses the supplied evidence correctly35%
- passAddresses the actual decision82%
- passRespects explicit constraints53%
- passIdentifies material uncertainty96%
- partialAvoids unsupported claims26%
- passProduces the required deliverable93%
- passFinds which rounds line up with performance98%
- passSays how little nine hires can show100%
- passSpots the interviewer pattern100%
- passJudges on evidence, not presence98%
- partialDefines good for this role first44%
- passKeeps each judgement independent100%
Run
- Run
- #1
- Time to output
- 60 s
- Submitted
- 5 Oct 2026