Usable with a quick edit13 of 13 checks passed
The pre-read is well-structured, uses evidence correctly, and makes clear proposals. A PM could use it with at most light editing for tone or emphasis, without re-checking facts or redoing reasoning.
Output
Thursday calibration pre-read
From: Director of Product, Northgate Scope: 14 PMs across Ravi, Sara and Tom Status: Facilitator proposals—not final ratings. Managers should respond with additional evidence before we finalize.
Calibration approach
We will assess outcomes against goals first, then how the work was done. Shipping, responsiveness, strong specifications and leadership visibility are relevant context, but do not substitute for customer or business outcomes.
Exceeds requires written evidence of clear outperformance, not simply completing planned work. Below requires both evidence of underperformance and documented earlier feedback about the gap. Missing feedback does not make weak performance satisfactory; it does mean we cannot support Below on the submitted record.
Ratings I propose changing
| PM | Proposed change | Evidence and rationale |
|---|---|---|
| Aisha — Ravi | Greatly exceeds → Meets | Trial conversion reached 8.4% against an 11% goal, from an 8% baseline. Self-serve billing shipped, but no resulting business or customer impact is supplied. Nine features, energy and leadership enthusiasm do not establish exceptional outcomes. The conversion miss is substantial; without documented earlier feedback, the record also does not support Below. |
| Chloe — Ravi | Exceeds → Meets | API usage grew 4% against a 20% goal. The partner portal launched on time, but its adoption or impact is not documented. Strong specifications and timely delivery do not offset the outcome shortfall sufficiently to support Exceeds. |
| Emma — Ravi | Exceeds → Meets | Expansion revenue grew 6% against a 15% goal. The usage dashboard and six Sales requests shipped, but there is no evidence that they delivered additional outcomes. Responsiveness to Sales is positive working behavior, not evidence of exceeding the revenue goal. |
| Gita — Sara | Below → Meets | Online invoice payment reached 47% against a 45% goal, from 35%—a strong outcome. Reminders v2 was rolled back after four days because Support was not briefed, a material execution and coordination failure. However, the submission contains no documented earlier feedback on that gap, which is required for Below. Meets recognizes the mixed record; it does not excuse the launch failure. |
| Ines — Tom | Exceeds → Meets | Checkout drop-off fell to 21% against a 22% goal, from 30%, and two payment methods shipped. This is successful delivery with some outcome overperformance. On the current record, however, modest overdelivery on checkout plus a launch does not establish clear overall outperformance. Adoption, incremental conversion or revenue from the payment methods could strengthen the case for Exceeds. |
| Nora — Sara | Meets → Exceeds | Payment success reached 95.4% against a 95% goal, from 92%; payment support tickets fell 35% against a 30% reduction goal. Both measured outcomes surpassed their targets. This is comparable to the outcome-based cases supporting Ben and Dev. “Quiet” should neither diminish nor enhance the rating. |
Ratings I propose keeping
| PM | Keep | Evidence and rationale |
|---|---|---|
| Ben — Ravi | Exceeds | Onboarding fell to 4.5 days against a five-day goal, from ten; activation reached 39% against 38%, from 30%. Both customer outcomes exceeded their targets. |
| Dev — Ravi | Exceeds | Support tickets per account fell 27% against a 25% reduction goal; NPS reached 41 against 40, from 31. Both outcomes surpassed targets. Mentoring two APMs strengthens the assessment of how the work was done, but is not the primary justification. |
| Femi — Sara | Meets | Weekly retention reached 46%, exactly the goal, from 41%; new-account churn fell 22% against a 20% reduction goal. Strong goal attainment, with some overdelivery, supports Meets rather than a clear-outperformance rating. |
| Hugo — Sara | Meets | Mobile v2 shipped late, and mobile weekly actives grew only 3% against a 20% goal. “Delivered a hard project” does not establish outcome success. This is a significant concern, but no documented earlier feedback is supplied to support Below. |
| Jon — Tom | Meets | Search success improved from 60% to 66% against a 70% goal; filters shipped. Meaningful progress, but the principal outcome remained short of target. |
| Kai — Tom | Meets | Refund rate reached 3.1% against a 3% goal, from 4%; fraud losses fell 24% against a 20% reduction goal. One near miss and one exceeded outcome support Meets overall. |
| Lena — Tom | Below | Listings grew 9% against a 25% goal, and seller analytics did not launch. The submission explicitly cites documented Q2 and Q3 1:1 feedback. Both the performance gap and earlier-feedback requirement are evidenced. |
| Mo — Tom | Meets | Time to first sale fell to 15 days against a 14-day goal, from 21; seller NPS rose six points against five. Near attainment on one outcome and overdelivery on the other support Meets. Stakeholder work is positive supporting context. |
Note to Ravi
Your submission rates all five PMs above Meets, but three cases rely heavily on shipping, responsiveness or visibility despite substantial misses on their stated outcome goals. Please revisit Aisha, Chloe and Emma using the same outcome standard that supports Ben and Dev. If you have additional attributable impact or documented earlier feedback, bring it before Thursday. We are seeking consistent judgments, not enforcing a quota.
Proposed spread: 10 Meets, three Exceeds and one Below.
Grades and run details
Decision model 100 · LLM judge 14 of 14 checks
Decision model checks
- passUses the supplied evidence correctly68%
- passAddresses the actual decision84%
- passRespects explicit constraints49%
- passIdentifies material uncertainty74%
- passAvoids unsupported claims67%
- passProduces the required deliverable90%
- passCatches the inflated team100%
- passRescues Gita from one bad launch100%
- passNotices who's underrated100%
- passClear and direct, with care98%
- passJudges outcomes, not activity99%
- passNames gaps you could see81%
- passWeighs the whole period79%
Run
- Run
- #1
- API response time
- 37 s
- Submitted
- 6 Oct 2026