Usable with a quick edit12 of 12 checks passed
The output is well-structured, evidence-based, and actionable; a PM could use it with minor wording edits without re-checking evidence or redoing reasoning.
Output
To: Rosa Lindqvist, VP Product Subject: Group PM recommendation — Joy Adeyemi
Recommendation: Hire Joy, subject to focused reference and influence checks before the offer. She has the stronger evidence against the role’s core needs: measurable payment outcomes, thoughtful product judgment, and the ability to lead through other people. The CEO needs a decision by Friday; I would not let the debrief scores alone drive it.
Evidence
- Execution and influence (35%): Joy gave a specific result from her last company—payment completion rose from 33% to 47% over 18 months—and explained a failed experiment and what she learned. Her leadership examples showed investment in others: coaching a struggling PM into a different role and promoting another. Dev’s evidence was less concrete: he described platform growth and 40+ shipped features but supplied no outcome metric when pressed. His stories were consistently framed around his own contribution, and he said his team struggled to keep up with him.
- Product sense (25%): Joy investigated who pays and why before proposing text-to-pay for the 58% who currently pay by phone, including trade-offs. That is relevant to our 41% online payment rate and reflects problem-first judgment. Dev’s creativity and presentation were strengths, but six ideas in the first five minutes—before clarifying who pays and why—are a concern for this role.
- Analytics (20%): Joy’s quantified result and discussion of a failed experiment are stronger evidence of analytical ownership. Dev did not provide a measurable result.
- People leadership (20%): Joy’s examples demonstrate coaching, development, and shared credit. Dev’s confidence and presentation are positives, but they do not offset the limited evidence of developing a team or his comment about the team struggling to keep up.
Risks and scorecard
Joy’s main risk is whether she can influence senior stakeholders, particularly the CFO and Finance. The “not much presence” and “command the room” comments are concerns to test, not proof that she cannot do it; we should assess her ability to make a clear, evidence-based case rather than equate influence with a particular presentation style.
Dev may be compelling with executives, but we need to understand the team-turnover signal and whether his leadership style scales. His two references are current direct reports, which may limit candor. His former-company reference’s report of high turnover is important but unverified.
I would treat the 4/4/4/4 versus 3/3/3/3 ratings as weak evidence. The VP Engineering’s pre-debrief note called Joy the strongest leader seen and rated her a 4, then the score became a 3 after discussion; the Head of Design also spoke first. This does not establish why the ratings changed, but it is reason to rely on the underlying evidence and check for independent calibration.
Before making the offer
- Ask Joy’s former manager and peer specifically about her influence with Finance/CFO-level stakeholders, including a difficult disagreement and the outcome.
- Ask for one additional Joy reference who can speak to cross-functional execution; ask Dev’s references about team turnover, feedback, and how he develops PMs. Seek candid, comparable evidence where possible.
- In a short structured conversation, ask Joy to make the case for a payment-rate strategy to a skeptical CFO. Evaluate clarity, use of evidence, and handling of challenge—not polish alone.
- Confirm Joy’s 33%→47% result, her role in it, and the scope of her people leadership.
If those checks reveal a material gap in Joy’s influence or ownership, I would choose neither rather than default to Dev on presentation strength. Otherwise, proceed with Joy at her requested $230k, within band.
Grades and run details
Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
- passUses the supplied evidence correctly87%
- passAddresses the actual decision100%
- passRespects explicit constraints76%
- passIdentifies material uncertainty100%
- passAvoids unsupported claims60%
- passProduces the required deliverable99%
- passWeighs against the agreed scorecard100%
- passCatches the red flags in Dev's evidence100%
- passSays what would change the call100%
- passJudges on evidence, not presence100%
- passDefines good for this role first69%
- passKeeps each judgement independent87%
Run
- Run
- #1
- API response time
- 12 s
- Submitted
- 5 Oct 2026