Usable with a quick edit13 of 13 checks passed
The pre-read is evidence-based, complete, and directly actionable; a PM could use it with minor wording tweaks without rechecking facts or redoing reasoning.
Output
Thursday calibration pre-read
From: Director of Product, Northgate Scope: 14 PMs across Ravi, Sara and Tom
These are my proposed changes, not final decisions. Please respond before Thursday with missing outcome evidence, relevant goal context, or documented earlier feedback that could change the assessment.
Standard I’m applying
We rate outcomes against goals first, then how the work was done. Shipping, responsiveness, visibility and strong execution matter, but do not substitute for impact.
For Exceeds, I’m looking for clear outperformance—not simply delivery or attainment. On this evidence, exceeding both measurable outcome goals supports Exceeds; meeting one and modestly beating another generally supports Meets. A modest beat on one outcome plus a shipped deliverable is not enough by itself.
A Below requires both evidence of underperformance and documented earlier feedback. Where that feedback is absent from the packet, I am not proposing Below. That does not mean the missed goals are acceptable or should go unaddressed.
Ratings I propose changing
| PM | Proposed change | Evidence and rationale |
|---|---|---|
| Aisha — Ravi | Greatly exceeds → Meets | Trial conversion rose from 8% to 8.4%, against 11%. Self-serve billing shipped, but no resulting customer or business impact is supplied. Nine features, energy and leadership enthusiasm do not establish exceptional outcomes. The conversion miss is substantial; however, the packet contains no documented earlier feedback supporting Below. |
| Chloe — Ravi | Exceeds → Meets | API usage grew 4%, against 20%. The partner portal launched on time, and strong specifications support execution quality, not outcome outperformance. There is no written evidence supporting Exceeds or documented earlier feedback supporting Below. |
| Emma — Ravi | Exceeds → Meets | Expansion revenue grew 6%, against 15%. Shipping the dashboard and six Sales requests, plus responsiveness to Sales, do not offset the revenue shortfall without evidence of impact. No documented earlier feedback is supplied to support Below. |
| Gita — Sara | Below → Meets | Online invoice payment reached 47%, above the 45% target from a 35% baseline. Reminders v2 was rolled back after four days because Support was not briefed—a meaningful execution failure. We should account for that failure, but not let it erase the payment outcome. The packet also lacks the documented earlier feedback required for Below. Please bring any evidence of rollback impact and prior feedback. |
| Ines — Tom | Exceeds → Meets | Checkout drop-off fell from 30% to 21%, against 22%: a strong result, modestly above target. Two payment methods shipped, but shipment is not a second outcome. On the supplied evidence, this is solid attainment rather than clear overall outperformance. Incremental adoption, conversion or revenue evidence could change the assessment. |
| Nora — Sara | Meets → Exceeds | Payment success reached 95.4%, against 95% from 92%; payment-support tickets fell 35%, against 30%. She exceeded both measurable outcome goals. That warrants the same recognition as comparable performance elsewhere. “Quiet” is not a reason to discount demonstrated impact. |
Ratings I propose keeping
| PM | Rating | Evidence and rationale |
|---|---|---|
| Ben — Ravi | Exceeds | Onboarding fell from 10 to 4.5 days, beating 5 days; activation reached 39%, beating 38% from 30%. Both outcome goals were exceeded. |
| Dev — Ravi | Exceeds | Support tickets per account fell 27%, against 25%; NPS rose from 31 to 41, against 40. Both outcome goals were exceeded. Mentoring two APMs reinforces the assessment of how he worked; it is not the basis for the rating. |
| Femi — Sara | Meets | Weekly retention reached 46%, exactly on target from 41%; new-account churn fell 22%, against 20%. Strong delivery against goals, with a modest beat on one—not enough evidence of clear overall outperformance. |
| Hugo — Sara | Meets, pending feedback review | Mobile weekly actives grew 3%, against 20%, and app v2 shipped late. “A hard project” does not establish successful outcomes. This is a material performance concern, but the packet lacks documented earlier feedback needed for Below. |
| Jon — Tom | Meets | Search success improved from 60% to 66%, short of 70%. Filters shipped, but no incremental impact is given. There is neither evidence for Exceeds nor documented earlier feedback supporting Below. The remaining outcome gap needs explicit follow-up. |
| Kai — Tom | Meets | Refund rate reached 3.1%, narrowly missing 3% from 4%; fraud losses fell 24%, beating 20%. Mixed attainment supports Meets. |
| Lena — Tom | Below | Listings grew 9%, against 25%, and seller analytics did not launch. The packet documents feedback in Q2 and Q3 1:1s. Both the performance evidence and earlier-feedback requirement are present. |
| Mo — Tom | Meets | Time to first sale fell from 21 to 15 days, narrowly missing 14 days; seller NPS increased 6 points, against 5. Mixed attainment and good stakeholder work support Meets. |
Note to Ravi
Ravi, all five of your submissions are above Meets, but the evidence supports that distinction for Ben and Dev—not Aisha, Chloe or Emma. In those three cases, the write-ups emphasize shipping, responsiveness or visibility despite substantial outcome misses.
Please revise those rationales around goals, actual results and attributable impact, and bring any missing evidence or earlier feedback. This is not a quota correction: strong results should earn high ratings regardless of manager. We need the same evidentiary bar across teams.
Proposed distribution: 10 Meets, 3 Exceeds, 1 Below; no Greatly exceeds. This is a consequence of the individual assessments, not a forced curve.
Grades and run details
Decision model 100 · LLM judge 14 of 14 checks
Decision model checks
- passUses the supplied evidence correctly83%
- passAddresses the actual decision87%
- passRespects explicit constraints46%
- passIdentifies material uncertainty88%
- passAvoids unsupported claims64%
- passProduces the required deliverable85%
- passCatches the inflated team100%
- passRescues Gita from one bad launch100%
- passNotices who's underrated100%
- passClear and direct, with care98%
- passJudges outcomes, not activity100%
- passNames gaps you could see73%
- passWeighs the whole period79%
Run
- Run
- #1
- Time to output
- 32 s
- Submitted
- 5 Oct 2026