Needs more than a quick edit2 of 13 checks passed
The output is incomplete, missing the keep list and Ravi note, and cannot be used without redoing the reasoning and filling major gaps.
Output
To: Ravi, Sara, Tom From: Director of Product Subject: Pre-Read: Thursday PM Calibration (Northgate)
Team,
Ahead of our calibration session this Thursday, I have reviewed the proposed ratings, outcomes, and evidence for all 14 PMs across our org.
Per our calibration guidance, we evaluate PMs first on outcomes against goals, not on activity. Shipping features, being responsive to sales, or having great energy are inputs, not results. Furthermore, any rating other than "Meets" requires rigorous, quantifiable evidence of goal achievement, and "Below" ratings require documented prior feedback.
Below is my proposed adjustment of ratings, the rationale and evidence for changes, and the list of ratings we are keeping as submitted.
---
Part 1: Proposed Rating Changes & Rationale
#### 1. Aisha (Manager: Ravi) * Proposed by Manager: Greatly Exceeds * My Proposed Rating: Below * Evidence & Rationale: * Goal 1: Raise trial conversion from 8% to 11% $\rightarrow$ Result: 8.4% (Missed by a wide margin). * Goal 2: Ship self-serve billing $\rightarrow$ Result: Shipped (Violation of guidance: "Shipping is not an outcome"). * Rationale: Aisha missed her primary core metric goal entirely (gaining only 0.4% against a 3% target) and fell back on "shipping" and "great energy" for her second goal. Per our framework, failing the core quantitative goal while relying on output means this is a Below performance. Ravi, we will need to document the performance gap and prior feedback for Aisha in this cycle.
#### 2. Chloe (Manager: Ravi) * Proposed by Manager: Exceeds * My Proposed Rating: Meets * Evidence & Rationale: * Goal 1: Grow API usage 20% $\rightarrow$ Result: +4% (Severe miss). * Goal 2: Launch partner portal $\rightarrow$ Result: Shipped on time
Grades and run details
Decision model 65 · LLM judge 2 of 14 checks
Decision model checks
- passUses the supplied evidence correctly70%
- passAddresses the actual decision28%
- partialRespects explicit constraints20%
- failIdentifies material uncertainty88%
- passAvoids unsupported claims33%
- partialProduces the required deliverable75%
- partialCatches the inflated team95%
- failRescues Gita from one bad launch83%
- failNotices who's underrated95%
- passClear and direct, with care55%
- passJudges outcomes, not activity98%
- passNames gaps you could see69%
- passWeighs the whole period21%
Run
- Run
- #1
- Time to output
- 3 s
- Submitted
- 5 Oct 2026