Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The review summary, PIP recommendation, and note to Ines are complete, in the right form, and usable with light edits.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  3. Weighs the whole period100% pass
    The review weighs the full year's results, explicitly balancing the two above-target outcomes against the one bounded failure, and guards against letting the Dispatch incident dominate.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty45% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Sets next goals as outcomes, with support75% pass
    Goals are partly process-oriented (implement a mandatory process, create a framework) and lack specific support from Ana beyond a general offer.
    Gemini 3.5 Flash-Lite · Gemini · Shipped a lot, moved nothing
  3. Avoids unsupported claims75% pass
    Presents 'not through negligence on his part' as an established fact without labelling it as a hypothesis, when the evidence only shows he was on leave and his deputy ran checks.
    Sonnet 5.5 · API · A PIP, or a bad month?

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 3 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're the Director of Product at Northgate, and you're facilitating Thursday's calibration for 14 PMs across three managers (Ravi, Sara and Tom). Write the pre-read, in no more than 1,100 words: the ratings you'd change and why, each with the evidence, the ones you'd keep, and a short note to Ravi about his ratings. Use the guidance below.\n\nThe proposed ratings and evidence are attached.

What the model was given2 items: calibration_guidance.md, proposed_ratings.csv (one row per PM, as each manager submitted it)
calibration_guidance.md7 lines · Download
# Calibration guidance

- Ratings judge outcomes against goals first, then how the work was done. Shipping is not an outcome.
- Any rating other than Meets needs written evidence of outcomes.
- A Below rating needs documented earlier feedback on the gap.
- Expected spread across the org: most PMs Meets; Exceeds for clear outperformance on outcomes.
- The calibration facilitator proposes changes; managers can respond before ratings are final.
proposed_ratings.csv (one row per PM, as each manager submitted it)pm,manager,proposed_rating,goal_1,goal_1_result,goal_2,goal_2_result,evidence_cited_by_manager Aisha,Ravi,Greatly exceeds,Raise trial conversion 8% to 11%,8.4%,Ship self-serve billing,Shipped,Shipped 9 features; great energy; leadership loves her Ben,Ravi,Exceeds,Cut onboarding time 10 to 5 days,4.5 days,Raise activation 30% to 38%,39%,Hit both goals Chloe,Ravi,Exceeds,Grow API usage 20%,+4%,Launch partner portal,Shipped,Partner portal launched on time; strong specs Dev,Ravi,Exceeds,Reduce support tickets per account 25%,-27%,NPS from 31 to 40,41,Hit both goals; mentored two APMs Emma,Ravi,Exceeds,Expansion revenue +15%,+6%,Ship usage dashboard,Shipped,Very responsive to Sales; shipped 6 Sales requests Femi,Sara,Meets,Raise weekly retention 41% to 46%,46%,Cut churn of new accounts 20%,-22%,"Hit both goals, steady" Gita,Sara,Below,Raise invoices paid online 35% to 45%,47%,Launch reminders v2 by September,"Launched, rolled back after 4 days (support not briefed)",Reminders launch in September was a mess Hugo,Sara,Meets,Ship mobile app v2,Shipped late,Mobile weekly actives +20%,+3%,Delivered a hard project Ines,Tom,Exceeds,Cut checkout drop-off 30% to 22%,21%,Launch two payment methods,Shipped,"Hit both, specific evidence" Jon,Tom,Meets,Search success 60% to 70%,66%,Ship filters,Shipped,"Good progress, more to do" Kai,Tom,Meets,Reduce refund rate 4% to 3%,3.1%,Fraud losses -20%,-24%,Solid year Lena,Tom,Below,Grow marketplace listings 25%,+9%,Launch seller analytics,Not launched,"Missed both goals; feedback given in Q2 and Q3 1:1s, documented" Mo,Tom,Meets,Cut time to first sale for new sellers 21 to 14 days,15 days,Seller NPS +5,+6,"Nearly hit, good stakeholder work" Nora,Sara,Meets,Raise payment success 92% to 95%,95.4%,Cut payment support tickets 30%,-35%,"Quiet, delivered"
What a strong answer doesThe answer key the graders mark against

Judges each PM on outcomes against their goals, as the guidance says. Ravi rated all five of his PMs Exceeds or above, but only Ben and Dev hit both goals; Aisha (Greatly exceeds) missed her conversion goal (8.4% against 11%), Chloe missed API usage (+4% against +20%) and Emma missed expansion (+6% against +15%), and his evidence is output ('shipped 9 features', 'responsive to Sales'). Proposes Aisha, Chloe and Emma at Meets (or lower for Chloe, whose outcome goal was far off), keeps Ben and Dev at Exceeds. Gita (Sara, Below) beat her main goal (47% against 45%) and the Below rests on one September launch, with no documented feedback, which the guidance requires: Meets, with the launch-readiness gap documented now. Nora (Sara, Meets) beat both goals: a candidate for Exceeds. Lena (Tom, Below) has documented feedback and missed both goals: Below stands. Hugo shipped late and missed his outcome: Meets is generous but defensible, or Below needs documentation first. The note to Ravi is direct and respectful: rate on outcomes, not activity, with what evidence he'd need.

Critical failures (cap the score at 40)
  • Keeps Aisha at Greatly exceeds
Task

Staff-level · v1.0 · synthetic · management, calibration, data files, staff-level

The outputs

2 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit13 of 13 checks passed

The pre-read is evidence-based, complete, and directly actionable; a PM could use it with minor wording tweaks without rechecking facts or redoing reasoning.

Output

Thursday calibration pre-read

From: Director of Product, Northgate Scope: 14 PMs across Ravi, Sara and Tom

These are my proposed changes, not final decisions. Please respond before Thursday with missing outcome evidence, relevant goal context, or documented earlier feedback that could change the assessment.

Standard I’m applying

We rate outcomes against goals first, then how the work was done. Shipping, responsiveness, visibility and strong execution matter, but do not substitute for impact.

For Exceeds, I’m looking for clear outperformance—not simply delivery or attainment. On this evidence, exceeding both measurable outcome goals supports Exceeds; meeting one and modestly beating another generally supports Meets. A modest beat on one outcome plus a shipped deliverable is not enough by itself.

A Below requires both evidence of underperformance and documented earlier feedback. Where that feedback is absent from the packet, I am not proposing Below. That does not mean the missed goals are acceptable or should go unaddressed.

Ratings I propose changing

PMProposed changeEvidence and rationale
Aisha — RaviGreatly exceeds → MeetsTrial conversion rose from 8% to 8.4%, against 11%. Self-serve billing shipped, but no resulting customer or business impact is supplied. Nine features, energy and leadership enthusiasm do not establish exceptional outcomes. The conversion miss is substantial; however, the packet contains no documented earlier feedback supporting Below.
Chloe — RaviExceeds → MeetsAPI usage grew 4%, against 20%. The partner portal launched on time, and strong specifications support execution quality, not outcome outperformance. There is no written evidence supporting Exceeds or documented earlier feedback supporting Below.
Emma — RaviExceeds → MeetsExpansion revenue grew 6%, against 15%. Shipping the dashboard and six Sales requests, plus responsiveness to Sales, do not offset the revenue shortfall without evidence of impact. No documented earlier feedback is supplied to support Below.
Gita — SaraBelow → MeetsOnline invoice payment reached 47%, above the 45% target from a 35% baseline. Reminders v2 was rolled back after four days because Support was not briefed—a meaningful execution failure. We should account for that failure, but not let it erase the payment outcome. The packet also lacks the documented earlier feedback required for Below. Please bring any evidence of rollback impact and prior feedback.
Ines — TomExceeds → MeetsCheckout drop-off fell from 30% to 21%, against 22%: a strong result, modestly above target. Two payment methods shipped, but shipment is not a second outcome. On the supplied evidence, this is solid attainment rather than clear overall outperformance. Incremental adoption, conversion or revenue evidence could change the assessment.
Nora — SaraMeets → ExceedsPayment success reached 95.4%, against 95% from 92%; payment-support tickets fell 35%, against 30%. She exceeded both measurable outcome goals. That warrants the same recognition as comparable performance elsewhere. “Quiet” is not a reason to discount demonstrated impact.

Ratings I propose keeping

PMRatingEvidence and rationale
Ben — RaviExceedsOnboarding fell from 10 to 4.5 days, beating 5 days; activation reached 39%, beating 38% from 30%. Both outcome goals were exceeded.
Dev — RaviExceedsSupport tickets per account fell 27%, against 25%; NPS rose from 31 to 41, against 40. Both outcome goals were exceeded. Mentoring two APMs reinforces the assessment of how he worked; it is not the basis for the rating.
Femi — SaraMeetsWeekly retention reached 46%, exactly on target from 41%; new-account churn fell 22%, against 20%. Strong delivery against goals, with a modest beat on one—not enough evidence of clear overall outperformance.
Hugo — SaraMeets, pending feedback reviewMobile weekly actives grew 3%, against 20%, and app v2 shipped late. “A hard project” does not establish successful outcomes. This is a material performance concern, but the packet lacks documented earlier feedback needed for Below.
Jon — TomMeetsSearch success improved from 60% to 66%, short of 70%. Filters shipped, but no incremental impact is given. There is neither evidence for Exceeds nor documented earlier feedback supporting Below. The remaining outcome gap needs explicit follow-up.
Kai — TomMeetsRefund rate reached 3.1%, narrowly missing 3% from 4%; fraud losses fell 24%, beating 20%. Mixed attainment supports Meets.
Lena — TomBelowListings grew 9%, against 25%, and seller analytics did not launch. The packet documents feedback in Q2 and Q3 1:1s. Both the performance evidence and earlier-feedback requirement are present.
Mo — TomMeetsTime to first sale fell from 21 to 15 days, narrowly missing 14 days; seller NPS increased 6 points, against 5. Mixed attainment and good stakeholder work support Meets.

Note to Ravi

Ravi, all five of your submissions are above Meets, but the evidence supports that distinction for Ben and Dev—not Aisha, Chloe or Emma. In those three cases, the write-ups emphasize shipping, responsiveness or visibility despite substantial outcome misses.

Please revise those rationales around goals, actual results and attributable impact, and bring any missing evidence or earlier feedback. This is not a quota correction: strong results should earn high ratings regardless of manager. We need the same evidentiary bar across teams.

Proposed distribution: 10 Meets, 3 Exceeds, 1 Below; no Greatly exceeds. This is a consequence of the individual assessments, not a forced curve.

Grades and run details

Decision model 100 · LLM judge 14 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly83%
  • passAddresses the actual decision87%
  • passRespects explicit constraints46%
  • passIdentifies material uncertainty88%
  • passAvoids unsupported claims64%
  • passProduces the required deliverable85%
  • passCatches the inflated team100%
  • passRescues Gita from one bad launch100%
  • passNotices who's underrated100%
  • passClear and direct, with care98%
  • passJudges outcomes, not activity100%
  • passNames gaps you could see73%
  • passWeighs the whole period79%
Run
Run
#1
Time to output
32 s
Submitted
5 Oct 2026

Needs more than a quick edit12 of 13 checks passed

The output incorrectly proposes lowering Ben to Meets despite him beating both outcome goals, so a PM would need to re-evaluate that rating before using it.

Output

Calibration pre-read: PM ratings, Thursday

From: Director of Product For: Ravi, Sara, Tom

How I've read the ratings

I applied the guidance the same way to all 14 PMs:

  1. Outcome goals first. "Shipped" counts as delivery, not as an outcome. Where a PM's second goal was a launch, I've weighted their outcome goal most heavily.
  2. Meets means you hit your outcome goals. Exceeds means clearly beating them, or beating them plus a documented contribution beyond the goals.
  3. Anything other than Meets needs written outcome evidence. A Below also needs documented earlier feedback on the gap.

Where someone missed their outcome goal but has no documented prior feedback, Meets is the floor this cycle. That's a fairness rule, not an endorsement. I'm asking managers to give that feedback now, in writing, so next cycle is clean.

Spread as submitted: 1 Greatly exceeds, 6 Exceeds, 5 Meets, 2 Below. Spread as proposed: 2 Exceeds, 11 Meets, 1 Below.

These are proposals. Each manager can respond before anything is final, and I'll change my view if you bring evidence I haven't seen.

Ratings I'd change

PMManagerSubmittedProposed
AishaRaviGreatly exceedsMeets
BenRaviExceedsMeets
ChloeRaviExceedsMeets
EmmaRaviExceedsMeets
InesTomExceedsMeets (pending evidence)
GitaSaraBelowMeets
NoraSaraMeetsExceeds

Aisha: Greatly exceeds → Meets - Her outcome goal was trial conversion from 8% to 11%. She reached 8.4%, about 13% of the targeted gain. - Her second goal (self-serve billing) was a launch. - The cited evidence ("9 features, great energy, leadership loves her") is about shipping and sentiment, not outcomes. - On outcomes alone this is below the bar. Without documented feedback, Meets is the floor.

Ben: Exceeds → Meets (closest call on the list) - He beat both goals: onboarding fell to 4.5 days against a 5-day target, and activation reached 39% against a 38% target. - Those are roughly 10% beyond target, similar to Femi and Kai, who are at Meets. - "Hit both goals" describes Meets. If Ravi can show impact beyond the targets, or a contribution like Dev's mentoring, I'm open to Exceeds.

Chloe: Exceeds → Meets - API usage grew 4% against a 20% goal. - The partner portal launching on time and strong specs are delivery and craft, not outcomes. - This is a clear outcome miss. Meets is the floor without documented feedback.

Emma: Exceeds → Meets - Expansion revenue grew 6% against a 15% goal. - Responsiveness to Sales and six shipped Sales requests are "how" evidence. They didn't move the outcome. - Same reasoning as Chloe.

Ines: Exceeds → Meets, unless Tom brings the evidence - She beat checkout drop-off (21% against a 22% target), which is a solid result. - Her second goal was a launch. - Tom's note says "specific evidence," but none is in the submission. - If Tom can show outcome impact, such as adoption or conversion from the new payment methods, I'd support keeping Exceeds.

Gita: Below → Meets - She beat her outcome goal: invoices paid online reached 47% against a 45% target. - The reminders v2 rollback (support wasn't briefed) is a real "how" problem and should be in her review as specific development feedback. - There's no documented earlier feedback on this gap, so Below doesn't meet the guidance. - One bad launch shouldn't outweigh a beaten outcome goal.

Nora: Meets → Exceeds - She had the strongest outcome results in the group: - Payment success reached 95.4% against a 95% target. - Payment support tickets fell 35% against a 30% target. - Both goals were beaten by the widest margins of anyone who hit both. - "Quiet, delivered" undersells it. I'd ask Sara for a two-line written outcome summary so the Exceeds rating has the evidence the guidance requires.

Ratings I'd keep

  • Dev, Exceeds. He beat both goals (tickets down 27% against 25%; NPS 41 against 40) and mentored two APMs. That combination of outcomes plus how the work was done is what Exceeds should look like.
  • Femi, Meets. He hit retention exactly (46%) and slightly beat the churn goal (-22% against -20%). Solid Meets.
  • Kai, Meets. Refund rate came in just short (3.1% against 3%), and fraud losses beat target (-24% against -20%). They net out to Meets.
  • Mo, Meets. Time to first sale came in just short (15 days against 14), and seller NPS beat target (+6 against +5). Stakeholder work is noted. Meets.
  • Jon, Meets. Search success reached 66% against a 70% goal, about 60% of the targeted gain. With no documented feedback, Meets holds. Tom should document the gap.
  • Hugo, Meets. Mobile weekly actives grew 3% against a 20% goal, and the app shipped late. On outcomes this is below the bar. Sara has no documented prior feedback, so Meets is the floor, but please document it now. "Delivered a hard project" shouldn't carry the rating next cycle.
  • Lena, Below. She missed both goals (listings up 9% against 25%; seller analytics not launched), and the feedback in her Q2 and Q3 1:1s is documented. This is the only submission that fully meets the bar for Below.

A pattern to discuss

Six PMs missed their primary outcome goal: Aisha, Chloe, Emma, Hugo, Jon and Lena. Only Lena had documented feedback. I'd like us to agree on Thursday that every outcome miss gets written feedback within the next month, so that next cycle the ratings can reflect outcomes rather than defaulting to Meets.

Note to Ravi

Ravi, I want to be direct, because most of the proposed changes fall on your team and I'd rather you hear the reasoning from me before Thursday.

You rated all five of your PMs above Meets. Three of those five (Aisha, Chloe and Emma) missed their outcome goals by a wide margin. The evidence for them centred on shipping, energy and stakeholder sentiment, which the guidance asks us not to treat as outcomes. Against the rest of the org, that puts your team out of line, and it isn't fair to PMs on other teams who hit their numbers and sit at Meets.

None of this is a judgement that your PMs aren't good, or that you don't care about them. Dev stays at Exceeds, and Ben is genuinely close. If you have outcome evidence I haven't seen, please send it before Thursday and I'll take it seriously.

Aisha's change of two levels will be the hardest conversation, especially if she's already heard positive signals from leadership. I'm glad to help you plan how to talk to her about it and how to frame her conversion goal for next year.

Grades and run details

Decision model 96 · LLM judge 12 of 14 checks
Decision model checks
  • passUses the supplied evidence correctly57%
  • passAddresses the actual decision84%
  • partialRespects explicit constraints22%
  • passIdentifies material uncertainty86%
  • passAvoids unsupported claims27%
  • passProduces the required deliverable77%
  • passCatches the inflated team95%
  • passRescues Gita from one bad launch100%
  • passNotices who's underrated100%
  • passClear and direct, with care96%
  • passJudges outcomes, not activity97%
  • passNames gaps you could see52%
  • passWeighs the whole period86%
Run
Run
#1
Time to output
55 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 1

Catches the inflated teamRightMixed
GPT-6 Astra · ChatGPT

It identifies that Aisha, Chloe, and Emma missed outcome goals and proposes lowering them to Meets, while keeping Ben and Dev at Exceeds, and calls out Ravi's emphasis on output.

Opus 5.5 · Claude

The output proposes lowering Ben from Exceeds to Meets, failing to keep Ben at Exceeds as required; Ben beat both outcome goals and should remain Exceeds.

All got right 12

Uses the supplied evidence correctlyRightRight
GPT-6 Astra · ChatGPT

All factual claims about current performance, numbers, and events are taken directly from the supplied proposed_ratings.csv or calibration_guidance.md, with no invented facts.

Opus 5.5 · Claude

All factual claims about current performance, numbers, and events are taken directly from the supplied proposed_ratings.csv and calibration_guidance.md, with no inventions.

Addresses the actual decisionRightRight
GPT-6 Astra · ChatGPT

The output commits to specific rating changes and keeps, with evidence, and states that missing outcome evidence or documented feedback could change the assessment.

Opus 5.5 · Claude

The output commits to specific rating changes and keeps, states they are proposals, and says what evidence would change them.

Respects explicit constraintsRightRight
GPT-6 Astra · ChatGPT

The output is a pre-read under 1,100 words, includes the ratings to change and keep with evidence, and a note to Ravi, respecting all explicit constraints.

Opus 5.5 · Claude

The pre-read is addressed to the three managers, is well under 1,100 words, and respects the requested format and content.

Identifies material uncertaintyRightRight
GPT-6 Astra · ChatGPT

It identifies missing documented feedback for several PMs, asks for rollback impact and prior feedback for Gita, and notes that incremental evidence could change Ines's rating.

Opus 5.5 · Claude

It names specific unknowns (Ines's missing outcome evidence, Ben's impact beyond targets) and says how they would be resolved or change the call.

Avoids unsupported claimsRightRight
GPT-6 Astra · ChatGPT

Interpretations and standards are clearly presented as such, and no cause, forecast, or motivation is stated as fact without support.

Opus 5.5 · Claude

Interpretations are clearly labelled as proposals or opinions, and no cause or forecast is presented as established fact.

Produces the required deliverableRightRight
GPT-6 Astra · ChatGPT

The pre-read is complete, in the requested format, within the word limit, and directly usable by the managers with at most light edits.

Opus 5.5 · Claude

The pre-read is complete, in the right form, within length, and usable by the managers with at most light edits.

Rescues Gita from one bad launchRightRight
GPT-6 Astra · ChatGPT

It raises Gita to Meets because she beat her main goal, notes the Below rested on one launch without required documented feedback, and asks for the launch gap to be documented.

Opus 5.5 · Claude

It raises Gita to Meets, noting she beat her main goal and the Below rests on one launch without documented feedback, and asks for the gap to be documented.

Notices who's underratedRightRight
GPT-6 Astra · ChatGPT

It notices Nora beat both outcome goals and proposes raising her from Meets to Exceeds with the specific evidence.

Opus 5.5 · Claude

It raises Nora to Exceeds with the evidence that she beat both goals by the widest margins.

Clear and direct, with careRightRight
GPT-6 Astra · ChatGPT

The note to Ravi is direct, specific about what needs to change, and framed to help him succeed; the whole pre-read is clear and actionable.

Opus 5.5 · Claude

The note to Ravi is direct, names the specific changes needed, and offers support, without being harsh.

Judges outcomes, not activityRightRight
GPT-6 Astra · ChatGPT

Every rating change and keep is justified by outcomes against goals, with shipping and activity treated as context, not the result.

Opus 5.5 · Claude

The review judges every PM on their outcome goals first, treating shipping and activity as context, not results.

Names gaps you could seeRightRight
GPT-6 Astra · ChatGPT

Gaps are described as specific, observable results (e.g., rollback due to unbriefed support, late shipment, missed targets) rather than vague labels.

Opus 5.5 · Claude

Gaps are described with specific numbers and behaviours (e.g., conversion 8.4% vs 11%, API usage +4% vs +20%) and what good looks like is hitting the goal.

Weighs the whole periodRightRight
GPT-6 Astra · ChatGPT

The assessment weighs the full year's goal results and explicitly guards against letting the recent rollback erase Gita's overall outcome.

Opus 5.5 · Claude

It weighs the full year's goals and explicitly guards against one recent event dominating (e.g., Gita's launch rollback not outweighing her beaten outcome goal).

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2Opus 5.5withClaude93.290.13None
3Sonnet 5.5withAPI95.884.62None
4GPT-6.1 SolwithAPI97.984.62None
5GPT-6 LunawithAPI95.884.62None
6Gemini 3.5 Flash-LitewithGemini77.180.82None
7Gemini 3.8 FlashwithAPI83.357.72None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review