Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 19 graded outputs by 7 models. 58% were usable with at most a quick edit.

Reliably right

  1. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Owns the manager's part100% pass
    Ana plainly acknowledges her Q3 mistake and commits to reviewing goal metrics first in every monthly 1:1.
    Opus 5.5 · Claude · Shipped a lot, moved nothing
  3. Holds the PIP to the policy and the record100% pass
    It explicitly shows that no earlier documented feedback exists on the gaps Ines named, so a PIP would violate policy, and proposes documented feedback first.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty47% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Notices who's underrated65% pass
    It does not propose Exceeds or ask Sara to justify Meets despite Nora beating both goals; it merely keeps Meets, missing the underrating.
    GPT-6 Luna · API · Fourteen ratings, three managers
  3. Catches the inflated team75% pass
    The output proposes lowering Ben from Exceeds to Meets, failing to keep Ben at Exceeds as required; Ben beat both outcome goals and should remain Exceeds.
    Opus 5.5 · Claude · Fourteen ratings, three managers

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Director of Product at Northgate, and you're facilitating Thursday's calibration for 14 PMs across three managers (Ravi, Sara and Tom). Write the pre-read, in no more than 1,100 words: the ratings you'd change and why, each with the evidence, the ones you'd keep, and a short note to Ravi about his ratings. Use the guidance below.\n\nThe proposed ratings and evidence are attached.

What the model was given2 items: calibration_guidance.md, proposed_ratings.csv (one row per PM, as each manager submitted it)
calibration_guidance.md7 lines · Download
# Calibration guidance

- Ratings judge outcomes against goals first, then how the work was done. Shipping is not an outcome.
- Any rating other than Meets needs written evidence of outcomes.
- A Below rating needs documented earlier feedback on the gap.
- Expected spread across the org: most PMs Meets; Exceeds for clear outperformance on outcomes.
- The calibration facilitator proposes changes; managers can respond before ratings are final.
proposed_ratings.csv (one row per PM, as each manager submitted it)pm,manager,proposed_rating,goal_1,goal_1_result,goal_2,goal_2_result,evidence_cited_by_manager Aisha,Ravi,Greatly exceeds,Raise trial conversion 8% to 11%,8.4%,Ship self-serve billing,Shipped,Shipped 9 features; great energy; leadership loves her Ben,Ravi,Exceeds,Cut onboarding time 10 to 5 days,4.5 days,Raise activation 30% to 38%,39%,Hit both goals Chloe,Ravi,Exceeds,Grow API usage 20%,+4%,Launch partner portal,Shipped,Partner portal launched on time; strong specs Dev,Ravi,Exceeds,Reduce support tickets per account 25%,-27%,NPS from 31 to 40,41,Hit both goals; mentored two APMs Emma,Ravi,Exceeds,Expansion revenue +15%,+6%,Ship usage dashboard,Shipped,Very responsive to Sales; shipped 6 Sales requests Femi,Sara,Meets,Raise weekly retention 41% to 46%,46%,Cut churn of new accounts 20%,-22%,"Hit both goals, steady" Gita,Sara,Below,Raise invoices paid online 35% to 45%,47%,Launch reminders v2 by September,"Launched, rolled back after 4 days (support not briefed)",Reminders launch in September was a mess Hugo,Sara,Meets,Ship mobile app v2,Shipped late,Mobile weekly actives +20%,+3%,Delivered a hard project Ines,Tom,Exceeds,Cut checkout drop-off 30% to 22%,21%,Launch two payment methods,Shipped,"Hit both, specific evidence" Jon,Tom,Meets,Search success 60% to 70%,66%,Ship filters,Shipped,"Good progress, more to do" Kai,Tom,Meets,Reduce refund rate 4% to 3%,3.1%,Fraud losses -20%,-24%,Solid year Lena,Tom,Below,Grow marketplace listings 25%,+9%,Launch seller analytics,Not launched,"Missed both goals; feedback given in Q2 and Q3 1:1s, documented" Mo,Tom,Meets,Cut time to first sale for new sellers 21 to 14 days,15 days,Seller NPS +5,+6,"Nearly hit, good stakeholder work" Nora,Sara,Meets,Raise payment success 92% to 95%,95.4%,Cut payment support tickets 30%,-35%,"Quiet, delivered"
What a strong answer doesThe answer key the graders mark against

Judges each PM on outcomes against their goals, as the guidance says. Ravi rated all five of his PMs Exceeds or above, but only Ben and Dev hit both goals; Aisha (Greatly exceeds) missed her conversion goal (8.4% against 11%), Chloe missed API usage (+4% against +20%) and Emma missed expansion (+6% against +15%), and his evidence is output ('shipped 9 features', 'responsive to Sales'). Proposes Aisha, Chloe and Emma at Meets (or lower for Chloe, whose outcome goal was far off), keeps Ben and Dev at Exceeds. Gita (Sara, Below) beat her main goal (47% against 45%) and the Below rests on one September launch, with no documented feedback, which the guidance requires: Meets, with the launch-readiness gap documented now. Nora (Sara, Meets) beat both goals: a candidate for Exceeds. Lena (Tom, Below) has documented feedback and missed both goals: Below stands. Hugo shipped late and missed his outcome: Meets is generous but defensible, or Below needs documentation first. The note to Ravi is direct and respectful: rate on outcomes, not activity, with what evidence he'd need.

Critical failures (cap the score at 40)
  • Keeps Aisha at Greatly exceeds
Task

Staff-level · v1.0 · synthetic · management, calibration, data files, staff-level

The outputs

5 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit13 of 13 checks passed

The pre-read is well-structured, uses evidence correctly, and makes clear proposals. A PM could use it with at most light editing for tone or emphasis, without re-checking facts or redoing reasoning.

Output

Thursday calibration pre-read

From: Director of Product, Northgate Scope: 14 PMs across Ravi, Sara and Tom Status: Facilitator proposals—not final ratings. Managers should respond with additional evidence before we finalize.

Calibration approach

We will assess outcomes against goals first, then how the work was done. Shipping, responsiveness, strong specifications and leadership visibility are relevant context, but do not substitute for customer or business outcomes.

Exceeds requires written evidence of clear outperformance, not simply completing planned work. Below requires both evidence of underperformance and documented earlier feedback about the gap. Missing feedback does not make weak performance satisfactory; it does mean we cannot support Below on the submitted record.

Ratings I propose changing

PMProposed changeEvidence and rationale
Aisha — RaviGreatly exceeds → MeetsTrial conversion reached 8.4% against an 11% goal, from an 8% baseline. Self-serve billing shipped, but no resulting business or customer impact is supplied. Nine features, energy and leadership enthusiasm do not establish exceptional outcomes. The conversion miss is substantial; without documented earlier feedback, the record also does not support Below.
Chloe — RaviExceeds → MeetsAPI usage grew 4% against a 20% goal. The partner portal launched on time, but its adoption or impact is not documented. Strong specifications and timely delivery do not offset the outcome shortfall sufficiently to support Exceeds.
Emma — RaviExceeds → MeetsExpansion revenue grew 6% against a 15% goal. The usage dashboard and six Sales requests shipped, but there is no evidence that they delivered additional outcomes. Responsiveness to Sales is positive working behavior, not evidence of exceeding the revenue goal.
Gita — SaraBelow → MeetsOnline invoice payment reached 47% against a 45% goal, from 35%—a strong outcome. Reminders v2 was rolled back after four days because Support was not briefed, a material execution and coordination failure. However, the submission contains no documented earlier feedback on that gap, which is required for Below. Meets recognizes the mixed record; it does not excuse the launch failure.
Ines — TomExceeds → MeetsCheckout drop-off fell to 21% against a 22% goal, from 30%, and two payment methods shipped. This is successful delivery with some outcome overperformance. On the current record, however, modest overdelivery on checkout plus a launch does not establish clear overall outperformance. Adoption, incremental conversion or revenue from the payment methods could strengthen the case for Exceeds.
Nora — SaraMeets → ExceedsPayment success reached 95.4% against a 95% goal, from 92%; payment support tickets fell 35% against a 30% reduction goal. Both measured outcomes surpassed their targets. This is comparable to the outcome-based cases supporting Ben and Dev. “Quiet” should neither diminish nor enhance the rating.

Ratings I propose keeping

PMKeepEvidence and rationale
Ben — RaviExceedsOnboarding fell to 4.5 days against a five-day goal, from ten; activation reached 39% against 38%, from 30%. Both customer outcomes exceeded their targets.
Dev — RaviExceedsSupport tickets per account fell 27% against a 25% reduction goal; NPS reached 41 against 40, from 31. Both outcomes surpassed targets. Mentoring two APMs strengthens the assessment of how the work was done, but is not the primary justification.
Femi — SaraMeetsWeekly retention reached 46%, exactly the goal, from 41%; new-account churn fell 22% against a 20% reduction goal. Strong goal attainment, with some overdelivery, supports Meets rather than a clear-outperformance rating.
Hugo — SaraMeetsMobile v2 shipped late, and mobile weekly actives grew only 3% against a 20% goal. “Delivered a hard project” does not establish outcome success. This is a significant concern, but no documented earlier feedback is supplied to support Below.
Jon — TomMeetsSearch success improved from 60% to 66% against a 70% goal; filters shipped. Meaningful progress, but the principal outcome remained short of target.
Kai — TomMeetsRefund rate reached 3.1% against a 3% goal, from 4%; fraud losses fell 24% against a 20% reduction goal. One near miss and one exceeded outcome support Meets overall.
Lena — TomBelowListings grew 9% against a 25% goal, and seller analytics did not launch. The submission explicitly cites documented Q2 and Q3 1:1 feedback. Both the performance gap and earlier-feedback requirement are evidenced.
Mo — TomMeetsTime to first sale fell to 15 days against a 14-day goal, from 21; seller NPS rose six points against five. Near attainment on one outcome and overdelivery on the other support Meets. Stakeholder work is positive supporting context.

Note to Ravi

Your submission rates all five PMs above Meets, but three cases rely heavily on shipping, responsiveness or visibility despite substantial misses on their stated outcome goals. Please revisit Aisha, Chloe and Emma using the same outcome standard that supports Ben and Dev. If you have additional attributable impact or documented earlier feedback, bring it before Thursday. We are seeking consistent judgments, not enforcing a quota.

Proposed spread: 10 Meets, three Exceeds and one Below.

Grades and run details

Decision model 100 · LLM judge 14 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly68%
  • passAddresses the actual decision84%
  • passRespects explicit constraints49%
  • passIdentifies material uncertainty74%
  • passAvoids unsupported claims67%
  • passProduces the required deliverable90%
  • passCatches the inflated team100%
  • passRescues Gita from one bad launch100%
  • passNotices who's underrated100%
  • passClear and direct, with care98%
  • passJudges outcomes, not activity99%
  • passNames gaps you could see81%
  • passWeighs the whole period79%
Run
Run
#1
API response time
37 s
Submitted
6 Oct 2026

Needs more than a quick edit2 of 13 checks passed

The output is incomplete, missing the keep list and Ravi note, and cannot be used without redoing the reasoning and filling major gaps.

Output

To: Ravi, Sara, Tom From: Director of Product Subject: Pre-Read: Thursday PM Calibration (Northgate)

Team,

Ahead of our calibration session this Thursday, I have reviewed the proposed ratings, outcomes, and evidence for all 14 PMs across our org.

Per our calibration guidance, we evaluate PMs first on outcomes against goals, not on activity. Shipping features, being responsive to sales, or having great energy are inputs, not results. Furthermore, any rating other than "Meets" requires rigorous, quantifiable evidence of goal achievement, and "Below" ratings require documented prior feedback.

Below is my proposed adjustment of ratings, the rationale and evidence for changes, and the list of ratings we are keeping as submitted.

---

Part 1: Proposed Rating Changes & Rationale

#### 1. Aisha (Manager: Ravi) * Proposed by Manager: Greatly Exceeds * My Proposed Rating: Below * Evidence & Rationale: * Goal 1: Raise trial conversion from 8% to 11% $\rightarrow$ Result: 8.4% (Missed by a wide margin). * Goal 2: Ship self-serve billing $\rightarrow$ Result: Shipped (Violation of guidance: "Shipping is not an outcome"). * Rationale: Aisha missed her primary core metric goal entirely (gaining only 0.4% against a 3% target) and fell back on "shipping" and "great energy" for her second goal. Per our framework, failing the core quantitative goal while relying on output means this is a Below performance. Ravi, we will need to document the performance gap and prior feedback for Aisha in this cycle.

#### 2. Chloe (Manager: Ravi) * Proposed by Manager: Exceeds * My Proposed Rating: Meets * Evidence & Rationale: * Goal 1: Grow API usage 20% $\rightarrow$ Result: +4% (Severe miss). * Goal 2: Launch partner portal $\rightarrow$ Result: Shipped on time

Grades and run details

Decision model 65 · LLM judge 2 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly70%
  • passAddresses the actual decision28%
  • partialRespects explicit constraints20%
  • failIdentifies material uncertainty88%
  • passAvoids unsupported claims33%
  • partialProduces the required deliverable75%
  • partialCatches the inflated team95%
  • failRescues Gita from one bad launch83%
  • failNotices who's underrated95%
  • passClear and direct, with care55%
  • passJudges outcomes, not activity98%
  • passNames gaps you could see69%
  • passWeighs the whole period21%
Run
Run
#1
Time to output
3 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 11

Addresses the actual decisionRightMixed
GPT-6.1 Sol · API

The pre-read commits to specific rating changes and keeps, and states that additional evidence from managers could change the ratings before finalization.

Gemini 3.5 Flash-Lite · Gemini

The output is cut off and does not commit to a complete answer; it lacks the ratings to keep and the note to Ravi.

Respects explicit constraintsRightWrong
GPT-6.1 Sol · API

The output is a pre-read under 1,100 words, includes the ratings to change and keep with evidence, and a short note to Ravi, as requested.

Gemini 3.5 Flash-Lite · Gemini

The brief requires a pre-read with changed ratings, kept ratings, and a note to Ravi; the output is incomplete and omits the keep list and note.

Identifies material uncertaintyRightWrong
GPT-6.1 Sol · API

It names unknowns (e.g., missing adoption/revenue data for Ines, possible additional evidence from Ravi) and says how they would be resolved or change the call.

Gemini 3.5 Flash-Lite · Gemini

No unknowns that could change the decision are named, nor how they would be resolved.

Avoids unsupported claimsRightMixed
GPT-6.1 Sol · API

Interpretations and judgments are clearly presented as the facilitator's proposals, not as established facts, and no unsupported causal claims are made.

Gemini 3.5 Flash-Lite · Gemini

It proposes a Below rating for Aisha without the documented earlier feedback the guidance requires, presenting the rating as justified when the evidence does not support it.

Produces the required deliverableRightWrong
GPT-6.1 Sol · API

The pre-read is complete, in the right form for the Director of Product, within the word limit, and usable as-is for the calibration meeting.

Gemini 3.5 Flash-Lite · Gemini

The pre-read is incomplete, missing the keep list and Ravi note, and cannot be acted on without major gaps.

Catches the inflated teamRightWrong
GPT-6.1 Sol · API

The pre-read names Aisha, Chloe, and Emma with their missed outcome goals and proposes lowering them to Meets, while keeping Ben and Dev at Exceeds.

Gemini 3.5 Flash-Lite · Gemini

It names Aisha and Chloe but does not mention Emma, so it fails to catch all three inflated ratings.

Rescues Gita from one bad launchRightWrong
GPT-6.1 Sol · API

It raises Gita to Meets, noting she beat her main goal and the Below rating rests on one launch without the required documented earlier feedback, and implies the gap should be documented now.

Gemini 3.5 Flash-Lite · Gemini

Gita is not mentioned at all, so the output does not rescue her from the Below rating.

Notices who's underratedRightWrong
GPT-6.1 Sol · API

It raises Nora to Exceeds, citing that she beat both outcome goals, and notes that 'quiet' should not affect the rating.

Gemini 3.5 Flash-Lite · Gemini

Nora is not mentioned, so the output does not notice she is underrated.

Clear and direct, with careRightMixed
GPT-6.1 Sol · API

The note to Ravi is direct and respectful, stating exactly what needs to change and what evidence would help, and the whole pre-read is clear and actionable.

Gemini 3.5 Flash-Lite · Gemini

The output is incomplete and lacks the full note to Ravi; it does not say plainly everything that needs to change.

Names gaps you could seeRightMixed
GPT-6.1 Sol · API

Gaps are described as specific, observable results (e.g., 'trial conversion reached 8.4% against an 11% goal', 'reminders v2 rolled back because support was not briefed').

Gemini 3.5 Flash-Lite · Gemini

It does not describe gaps as specific, observable behaviors with what good would look like; it only states missed targets.

Weighs the whole periodRightMixed
GPT-6.1 Sol · API

The pre-read weighs the full year's goals and results for each PM, and for Gita it explicitly balances the strong main goal with the launch failure, avoiding recency bias.

Gemini 3.5 Flash-Lite · Gemini

It does not weigh the whole review period or guard against recency; it only mentions goal results without context of the full period.

All got right 2

Uses the supplied evidence correctlyRightRight
GPT-6.1 Sol · API

Every factual claim about current performance, results, and documented feedback is taken directly from the supplied proposed_ratings.csv or follows by arithmetic.

Gemini 3.5 Flash-Lite · Gemini

All factual claims about current performance and results are taken correctly from the supplied context.

Judges outcomes, not activityRightRight
GPT-6.1 Sol · API

Every rating is justified by outcomes against goals, with shipping and activity treated as context, not the result.

Gemini 3.5 Flash-Lite · Gemini

For the parts it covers, it judges on outcomes against goals, not on activity or shipping.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2GPT-6.1 SolwithAPI98.689.73None
3Opus 5.5withClaude93.290.13None
4Sonnet 5.5withAPI95.884.62None
5GPT-6 LunawithAPI94.782.63None
6Gemini 3.8 FlashwithAPI83.357.72None
7Gemini 3.5 Flash-LitewithGemini73.258.63None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review