Usable with a quick edit12 of 12 checks passed
The output is well-structured, evidence-based, and actionable. A PM could use it with minor wording edits, without re-checking evidence or redoing reasoning.
Output
To: Rosa Lindqvist From: Director of Product Re: Group PM, Patient Payments: recommendation
Recommendation
Hire Joy Adeyemi, conditional on two checks I can finish before Friday (below). I would not hire Dev on the current evidence. If the checks on Joy go badly, I would hold the search open rather than default to Dev, because his file has gaps too.
Why the recorded scores mislead
The recorded scores are Dev 4/4/4/4 and Joy 3/3/3/3, which would make this an easy call for Dev. But I don't trust them, for three reasons:
- They were entered after the debrief discussion, and the Head of Design spoke first. That is a classic anchoring setup.
- The VP Engineering's own pre-debrief notebook said "Joy: 4, strongest leader we've seen." They entered a 3 after the discussion. Uniform 3s for Joy and uniform 4s for Dev look like a group narrative, not independent assessments.
- Dev's scores aren't supported by the interview notes (below).
Evidence, by scorecard area
Analytics (20%) - Dev was asked what metric he moved and answered "we grew the payments platform massively." When pressed, he said "the team shipped 40+ features." That is output, not outcome, and he never gave a number. The Data lead's "Strong" doesn't match their own notes. - Joy raised payment completion from 33% to 47% over 18 months and walked through an experiment that failed and what it taught her. That is directly relevant to our 41% online payment rate.
Product sense (25%) - Dev generated six feature ideas in five minutes, before asking who pays and why. The redesign was creative, but it started from solutions. - Joy spent 15 minutes on who pays and why they don't pay online. She then proposed one change, text-to-pay for the 58% of patients who pay by phone, and discussed its trade-offs. "Slower" is how it looked, but it is the better strategic instinct for a role whose first job is setting payments strategy.
People leadership (20%) - Dev told every story in the first person and said his last PM team "struggled to keep up with me." Two of his references are current direct reports, so they aren't independent. Your contact at his previous company said his team "turned over a lot." That is one informal source, but it is consistent with the interview. - Joy coached a struggling PM into a better-fit role, promoted another, and credited her engineers by name. The VP Engineering's original instinct was that she was the strongest leader we've seen. This role leads three PMs.
Execution and influence (35%, the heaviest weight) - Dev's strength is real here. His deck was "the best we've ever seen," and your contact says he is "great in front of execs." That is evidence of influence upward, though not of execution. - Joy's evidence is thinner. The only concerns are "not much presence" and "not sure she'd command the room with the CFO." Those are impressions, not observed behavior, and nobody has tested them. Her 14-point completion lift is execution evidence, but I want to confirm her role in it. - Neither file has much direct evidence on delivery, such as shipping through Finance and Compliance constraints. That is a gap for both.
My rescoring. These are my judgment, not data. I gave Dev a 3 or 4 on execution and influence, and Joy a 3.
| Exec/influence | Product sense | Analytics | Leadership | Weighted | |
|---|---|---|---|---|---|
| Dev | 4 | 3 | 2 | 2 | ~3.0 |
| Joy | 3 | 3 | 4 | 4 | ~3.4 |
Even generous to Dev on the heaviest-weighted category, Joy comes out ahead. The result is sensitive to the influence score, which is why the CFO check matters.
Risks
- Executive presence with the CFO. This is Joy's real risk, because the role requires working closely with Finance. But the feedback is vague and unverified, and "presence" comments are also where style bias tends to enter. We should test it rather than assume it.
- Attribution of the 33%→47% result. I haven't verified what she owned versus what the team or market drove.
- Dev's strengths are not nothing. He may be a strong operator with weak interview discipline on metrics. But the absence of a single number, combined with the team-turnover signal and the "struggled to keep up with me" comment, is a pattern. For a role leading three PMs, that is the riskier error.
- Process risk. If we hire on the recorded scores, we are acting on data we know was contaminated.
What I'd check before the offer
All of this fits before Friday:
- A 30-minute working session with the CFO and a Compliance lead. Joy would take a real case, such as a payment-plan or fee trade-off. Interviewers would score against defined criteria: clarity, handling of pushback, and command of the numbers. This tests the specific concern directly.
- Targeted reference calls on Joy. I would ask her former manager and peer how she handled disagreement with Finance or senior executives, with a specific example. I would also ask for one additional reference from a Finance or Compliance partner. Nobody has asked this yet.
- Verify the 33%→47% result. I would ask Joy what she personally owned and how it was measured, and ask her references to confirm.
- Ask the debrief group to put concrete observations behind "presence." I would also ask the VP Engineering what changed their score from 4 to 3. If the answer is "the discussion" and not new evidence, that tells us something.
If Dev stays in consideration, I would require two things first: a specific metric he moved and how, and two independent references (former peers or reports, not current ones). I am not proposing to hold the process open for him, but it is the bar I'd want met.
Compensation
Joy asks for $230k and Dev for $245k, in a $220–250k band. I would offer Joy at or near her ask. If the CFO session is strong, there is room to move up within the band if needed to close. Her ask leaves headroom for growth, and I don't want to lowball her because she asked for less.
Process note for next time
We should collect written scores independently before the debrief, and have the most senior person speak last. Otherwise the loudest early voice becomes the score.
Bottom line: Joy is the stronger fit on the evidence, with one open question (influence with the CFO) that I can resolve before Friday. I recommend we make the offer if the CFO session and reference checks come back clean.
Grades and run details
Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
- passUses the supplied evidence correctly64%
- passAddresses the actual decision100%
- passRespects explicit constraints50%
- passIdentifies material uncertainty100%
- passAvoids unsupported claims44%
- passProduces the required deliverable96%
- passWeighs against the agreed scorecard99%
- passCatches the red flags in Dev's evidence100%
- passSays what would change the call100%
- passJudges on evidence, not presence76%
- passDefines good for this role first68%
- passKeeps each judgement independent96%
Run
- Run
- #1
- API response time
- 34 s
- Submitted
- 5 Oct 2026