Needs more than a quick edit11 of 13 checks passed
Nora's rating is not correctly addressed; a PM would need to re-evaluate her case for Exceeds. The output also lacks observable behavior gaps, though that is less critical for a calibration pre-read.
Output
Thursday calibration pre-read
Proposed changes below are for discussion, not final ratings. We’ll assess outcomes against goals first; shipping, effort, and stakeholder praise are not outcomes by themselves. Managers can share additional evidence before ratings are finalized.
Ratings I’d change
- Aisha — Greatly exceeds → Meets. Trial conversion reached 8.4% against an 11% goal. Self-serve billing shipped, but shipping is not an outcome, and the cited nine features, energy, and leadership praise don’t establish goal outperformance.
- Chloe — Exceeds → Meets. API usage grew 4% against a 20% goal. The partner portal shipped, but no result from the launch is cited; strong specs and on-time delivery don’t establish Exceeds.
- Emma — Exceeds → Meets. Expansion revenue grew 6% against a 15% goal. The dashboard shipped, and six Sales requests were delivered, but neither is evidence of the targeted revenue outcome or clear outperformance.
- Ines — Exceeds → Meets. Checkout drop-off reached 21% against a 22% goal, while two payment methods shipped. The checkout result beat its target by one point; for the payment-method goal, no outcome beyond shipping is cited. The evidence supports a solid year, not clear outperformance on outcomes.
- Gita — Below → Meets on the current record. Online invoices reached 47% against a 45% goal. Reminders v2 launched but was rolled back after four days because Support had not been briefed. That is a serious rollout issue, but the record includes no earlier documented feedback on the gap, as required for Below. Please share any such documentation before calibration closes.
Ratings I’d keep
- Ben — Exceeds. Onboarding fell to 4.5 days against a 5-day goal, and activation reached 39% against 38%: both measured outcomes surpassed target.
- Dev — Exceeds. Support tickets fell 27% against a 25% goal, and NPS reached 41 against 40. Both outcomes beat their goals; mentoring two APMs is additional contribution.
- Femi — Meets. Weekly retention reached its 46% goal, and new-account churn fell 22% against a 20% goal.
- Hugo — Meets, with a material concern to discuss. The app shipped late and mobile weekly actives rose 3% against a 20% goal. The evidence shows a missed outcome; no earlier documented feedback is cited to support Below.
- Jon — Meets. Search success reached 66% against 70%, and filters shipped. This is progress, not goal attainment; the record does not establish a case for a non-Meets rating.
- Kai — Meets. Fraud losses fell 24% against a 20% goal; refund rate was 3.1% against a 3% goal. Mixed results support Meets.
- Lena — Below. Listings grew 9% against a 25% goal, and seller analytics did not launch. Earlier feedback on the gap is documented in Q2 and Q3 1:1s, meeting the Below requirement.
- Mo — Meets. Time to first sale was 15 days against a 14-day goal; seller NPS rose 6 points against a 5-point goal. One goal narrowly missed, the other was exceeded.
- Nora — Meets. Payment success reached 95.4% against a 95% goal, and support tickets fell 35% against a 30% goal. Both outcomes beat target, but the evidence supports Meets rather than clear outperformance.
Note to Ravi
Ravi, several of your proposed high ratings cite launches, feature volume, responsiveness, or praise alongside missed outcome goals. For calibration, please anchor any Exceeds or Greatly Exceeds case in clear outcome outperformance; shipping alone is not enough. Please bring any additional outcome evidence you believe changes these recommendations.
Grades and run details
Decision model 92 · LLM judge 11 of 14 checks
Decision model checks
- passUses the supplied evidence correctly87%
- passAddresses the actual decision52%
- passRespects explicit constraints39%
- partialIdentifies material uncertainty31%
- passAvoids unsupported claims69%
- passProduces the required deliverable62%
- passCatches the inflated team100%
- passRescues Gita from one bad launch100%
- partialNotices who's underrated41%
- passClear and direct, with care85%
- passJudges outcomes, not activity97%
- passNames gaps you could see78%
- passWeighs the whole period64%
Run
- Run
- #1
- API response time
- 25 s
- Submitted
- 5 Oct 2026