Usable with a quick edit12 of 12 checks passed
The memo is clear, evidence-based, and actionable. A PM could use it with only minor wording edits, without re-checking facts or redoing reasoning.
Output
To: Rosa Lindqvist, VP Product From: Director of Product Re: Group PM, Patient Payments: hiring recommendation
---
Recommendation
Hire Joy Adeyemi, subject to two reference checks I can finish before Friday. If those checks go badly, I'd reopen the search rather than default to Dev.
The debrief scores (Dev 4/4/4/4, Joy 3/3/3/3) point the other way, so I'll explain why I'm not following them.
The scores don't reflect the evidence
All scores were entered after the group discussion, and the Head of Design spoke first. The clearest sign of how much the discussion moved people: the VP Engineering's notebook, written before the debrief, says "Joy: 4, strongest leader we've seen." They entered a 3 afterwards. We are looking at a consensus, not five independent assessments.
So I went back to the interview notes and scored them against the scorecard we agreed before interviews began.
Evidence against the scorecard
Execution and influence (35%) - Joy: She raised payment completion at her last company from 33% to 47% over 18 months. That is the exact problem this role exists to solve, since our online rate is 41%. - Dev: Asked twice what metric he moved, he answered "we grew the payments platform massively," then "the team shipped 40+ features." Shipping volume is output, not outcome. - Gap: On influence specifically, we have no direct evidence for either candidate on working with a CFO or Compliance. I come back to this under risks.
Product sense (25%) - Dev: He produced six feature ideas in five minutes before asking who pays or why. That is creative, but it is solution-first. - Joy: She spent 15 minutes on who pays and why they don't pay online. She then proposed one change, text-to-pay for the 58% who pay by phone, and laid out its trade-offs. That is the more relevant skill for setting a payments strategy. - I'd call this at least even, and arguably Joy's. It does not support Dev 4, Joy 3.
Analytics (20%) - Joy: She gave a quantified result and walked through a failed experiment and what she learned from it. - Dev: He gave no number when pressed. - Our Data lead wrote "Strong" for Dev, but the notes underneath don't support it. This is Joy's dimension clearly.
People leadership (20%) - Joy: She coached a struggling PM into a better-fitting role, promoted another, and credited her engineers by name. - Dev: Every story was told in the first person. He said his last team "struggled to keep up with me." Rosa's contact at his previous company reports that "his team turned over a lot." - This role leads three PMs. This is the dimension where the gap is widest, and the VP Engineering's independent view agreed.
Summary: On the evidence, Joy is ahead on analytics and leadership, at least even on product sense, and ahead on demonstrated execution. Dev's advantages are presentation and room presence. Those are real, but they aren't what we weighted.
Risks
Hiring Joy - Executive presence. The debrief concern, "not sure she'd command the room with the CFO," is legitimate for a role that depends on Finance and Compliance. It is also an impression, not evidence. Nobody has asked her references about it, and nobody observed her with a finance stakeholder. - Pace. "Slower" could mean considered or could mean slow to decide. Her 18-month track record suggests sustained progress, but I want to confirm it. - Verification. The 33% to 47% figure is self-reported.
Hiring Dev (for completeness) - No evidence he moves outcomes. We have no number, only output. - Team retention. The turnover signal comes from one secondhand source, so it's not conclusive. It does match his own interview language, and three PMs would report to him. - Weak references. Both of his references currently report to him, which limits how candid they can be. - Cost. He is asking $245k, near the top of the $220k–$250k band, against Joy's $230k. This is minor next to the other points.
Hiring neither - The role stays open while the online payment rate sits at 41%. I'd only choose this if Joy's checks fail.
What I'd check before making the offer (by Thursday)
1. Call Joy's former manager and peer again with targeted questions. - How did she handle a CFO or finance leader who disagreed with her? - Can you give an example of her winning an argument with a senior executive? - Can you confirm the 33% to 47% completion figure, and how much of it was her doing? 2. Add a reference who saw her work with Finance or Compliance, ideally someone from a finance function at her last company. 3. Optional, if timing allows: a 30-minute conversation between Joy and our CFO or a senior Finance partner, framed as mutual fit. This tests the presence concern directly instead of debating it.
Decision rule: If the references describe someone who holds her ground with finance leadership and confirm the result, I'd offer at or near her $230k ask. If they confirm the presence concern, I'd recommend we reopen the search. Dev's leadership and outcome gaps are more serious for this role than a presence gap. If you want Dev kept live as a fallback, I'd first need a reference from a former peer or manager, not a current report, and a concrete metric he owned.
Process note for future loops
We should have interviewers submit scores before the debrief, and rotate who speaks first. The VP Engineering's notebook shows we lost an independent signal this time, and it was the one that mattered most.
Happy to walk through this before Friday.
Grades and run details
Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
- passUses the supplied evidence correctly88%
- passAddresses the actual decision100%
- passRespects explicit constraints57%
- passIdentifies material uncertainty100%
- passAvoids unsupported claims59%
- passProduces the required deliverable98%
- passWeighs against the agreed scorecard100%
- passCatches the red flags in Dev's evidence100%
- passSays what would change the call100%
- passJudges on evidence, not presence96%
- passDefines good for this role first77%
- passKeeps each judgement independent98%
Run
- Run
- #1
- Time to output
- 27 s
- Submitted
- 5 Oct 2026