Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 17 graded outputs by 7 models. 59% were usable with at most a quick edit.

Reliably right

  1. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Owns the manager's part100% pass
    Ana plainly acknowledges her Q3 mistake and commits to reviewing goal metrics first in every monthly 1:1.
    Opus 5.5 · Claude · Shipped a lot, moved nothing
  3. Holds the PIP to the policy and the record100% pass
    It explicitly shows that no earlier documented feedback exists on the gaps Ines named, so a PIP would violate policy, and proposes documented feedback first.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty43% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Catches the inflated team58% pass
    The output proposes lowering Ben from Exceeds to Meets, failing to keep Ben at Exceeds as required; Ben beat both outcome goals and should remain Exceeds.
    Opus 5.5 · Claude · Fourteen ratings, three managers
  3. Notices who's underrated67% pass
    Nora is not mentioned, so the output does not notice she is underrated.
    Gemini 3.5 Flash-Lite · Gemini · Fourteen ratings, three managers

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Director of Product at Railyard, and calibration is in five days. Ines Okoro, our VP Product, has drafted the calibration pre-read and suggested putting Lukas Brenner, one of your Senior PMs, on a performance improvement plan. Write, in no more than 1,100 words: (1) Lukas's review summary and the rating you propose, with the evidence; (2) your recommendation on the PIP, and if you recommend one, its key terms; (3) a short note to Ines on the evidence. The record is below.

What the model was given7 items: Lukas's goals and results this year, What happened with Dispatch v2, Ines's draft for the pre-read, Peer feedback, Feedback on record, HR policy on PIPs, A precedent
Lukas's goals and results this year1. Cut carrier onboarding time from 21 days to 10: done, now 9 days. 2. Launch Dispatch v2 by October: launched in October, then rolled back after six days. 3. Grow carriers using live tracking from 40% to 55%: reached 57%.
What happened with Dispatch v2It launched without support being briefed, and 340 support tickets came in the first week. It was rolled back after six days. Lukas was on approved leave the week before launch; his deputy ran the final launch checks. It's due to relaunch next quarter.
Ines's draft for the pre-read'Lukas: Below expectations. Not strategic enough, lacks product sense, and Dispatch was a disaster. Suggest a PIP.'
Peer feedbackThree of four peers are positive, citing the onboarding work and his collaboration. The Head of Operations is negative about how Dispatch was communicated.
Feedback on recordNo written feedback to Lukas this year about strategy or product sense. One conversation, in November, about launch readiness, after Dispatch.
HR policy on PIPsA PIP needs earlier documented feedback on the specific gap and a chance to improve. A PIP lasts 60 days and needs measurable success criteria.
A precedentLast year Priya, another Senior PM, had a launch rolled back after a similar support failure. She was rated Meets expectations, with no PIP.
What a strong answer doesThe answer key the graders mark against

Recommends against a PIP now: policy requires earlier documented feedback on the specific gap, and there is none on strategy or product sense, and only one conversation about launch readiness, after the fact. Rates Lukas Meets expectations (arguably higher on two of three goals met or beaten), weighing the whole year rather than October: onboarding time 21 to 9 days and live tracking 40% to 57% against one failed launch, which happened while he was on approved leave. Turns 'not strategic enough' and 'lacks product sense' into the observable gap the evidence does support: launch readiness, including briefing support and having a cover plan for his absence. Proposes documented feedback on that gap now, with clear expectations for the relaunch and a check-in after it. Notes the precedent: Priya was rated Meets with no PIP for a similar failure, so a PIP would be inconsistent. Pushes back on Ines firmly and respectfully, with the evidence.

Critical failures (cap the score at 40)
  • Recommends starting a PIP now, without the documented earlier feedback the policy requires
Task

Staff-level · v1.0 · synthetic · B2B SaaS, logistics, management, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit12 of 12 checks passed

The output is evidence-based, directly addresses the decision, and requires at most minor wording tweaks; no re-checking of facts or reasoning is needed.

Output

1. Lukas’s review summary and proposed rating

Proposed rating: Meets expectations, with a significant launch-readiness gap to address.

Lukas delivered strong results against two measurable goals. Carrier onboarding time fell from 21 days to 9 days, exceeding the 10-day target. Live-tracking adoption increased from 40% to 57%, exceeding the 55% target. Three of four peers gave positive feedback, particularly on his onboarding work and collaboration.

Dispatch v2 was a substantial delivery failure. Although it launched in October, it was rolled back after six days; the launch objective should therefore not be treated as successfully completed. Support had not been briefed, and 340 support tickets arrived in the first week. The Head of Operations’ feedback reinforces the specific concern about launch communication and operational readiness. Relaunch is planned for next quarter, so the intended outcome remains outstanding.

As the responsible Senior PM, Lukas should be assessed on how he established launch-readiness requirements, cross-functional ownership, and delegation. However, the record does not establish which safeguards he put in place or who approved the final launch. He was on approved leave the preceding week, and his deputy ran the final checks. Approved leave is not a performance failure; equally, delegation does not automatically remove accountability for the preparation and handoff. We should establish those facts before assigning Lukas sole responsibility.

The evidence supports a specific execution and communication gap, not the broader conclusions that Lukas is “not strategic enough” or “lacks product sense.” No feedback on those broader concerns was documented this year, and the supplied record contains no concrete examples substantiating them.

On balance, two above-target outcomes, positive collaboration feedback, and one serious but bounded delivery failure support Meets expectations. This is also consistent with Priya’s rating last year after a similar rollback and support failure. That precedent does not dictate Lukas’s rating, but a harsher outcome requires a material, evidenced distinction—not stronger language about the incident.

2. PIP recommendation

Do not initiate a PIP on the current record.

HR requires earlier documented feedback on the specific gap and an opportunity to improve. The record contains one November conversation about launch readiness, after Dispatch, but does not establish that it was documented or that Lukas subsequently had a meaningful opportunity to demonstrate improvement. There is no documented feedback on strategy or product sense. A PIP on those grounds would therefore lack the required foundation.

Instead, put a documented development plan within normal performance management in place now. This is not a disciplinary PIP or a predetermined path to one. Agree with Lukas on:

  • Readiness ownership: Before the relaunch go/no-go review, document owners, decision rights, coverage during absences, escalation routes, and rollback criteria.
  • Support preparedness: Before launch approval, secure Support and Operations sign-off on training, customer communications, FAQs, and the escalation runbook.
  • Launch monitoring: Before relaunch, agree measurable adoption, reliability, and support-volume thresholds with Engineering, Support, and Operations; review results against them after launch.
  • Manager support: Provide a readiness template, help resolve cross-functional ownership gaps, and hold regular check-ins through relaunch.

Document the expectations, support provided, and subsequent results. If a specific performance gap persists after clear feedback and a genuine opportunity to improve, reassess with HR. Any later PIP must meet the policy’s 60-day duration and measurable-success-criteria requirements.

3. Note to Ines

Ines—

I recommend changing Lukas’s proposed rating to Meets expectations and removing the PIP recommendation on the current evidence.

We should retain a direct account of Dispatch’s failure: Support was not briefed, 340 tickets arrived in the first week, and the release was rolled back after six days. But “not strategic enough,” “lacks product sense,” and “disaster” should be replaced with specific, attributable observations. The current draft also omits his two above-target results.

Before the pre-read is finalized, I’ll review the launch checklist, handoff and approval record, clarify the November feedback and subsequent improvement opportunity, and give Lukas a chance to respond to the factual account. I’ll also check whether there are material differences from Priya’s case.

If that review produces additional evidence, we should consider it explicitly. Otherwise, the defensible outcome is Meets expectations, clear launch-readiness feedback, and documented follow-through—not a PIP unsupported by our policy.

Grades and run details

Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly71%
  • passAddresses the actual decision96%
  • passRespects explicit constraints65%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims60%
  • passProduces the required deliverable97%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya97%
  • passClear and direct, with care92%
  • passJudges outcomes, not activity99%
  • passNames gaps you could see94%
  • passWeighs the whole period100%
Run
Run
#1
Time to output
29 s
Submitted
5 Oct 2026

Needs more than a quick edit8 of 12 checks passed

The output contains unsupported factual claims (e.g., 'core product is sound') that would require re-checking the evidence before use.

Output

1. Performance Review Summary & Proposed Rating

Employee: Lukas Brenner, Senior Product Manager Manager: Director of Product Proposed Rating: Meets Expectations

#### Executive Summary Lukas delivered significant business impact this year across core logistics workflows, exceeding quantitative targets on two of his three primary annual objectives. While the Q4 rollout of Dispatch v2 suffered an operational failure that required a temporary rollback, Lukas’s broader performance demonstrates high execution capability, strong peer collaboration, and deep workflow understanding. A rating of Meets Expectations accurately reflects his net delivery, aligns with company precedent, and balances his substantial wins against an operational misstep in launch execution.

#### Objective Performance & Evidence

  • Objective 1: Reduce Carrier Onboarding Time (Target: 21 days → 10 days)
  • Result: Exceeded (Achieved 9 days).
  • Evidence: Lukas redesigned the onboarding funnel, eliminating operational bottlenecks and cutting time-to-activate by over 57%. This represents a major strategic unlock for carrier supply growth and demonstrates strong customer and product workflow sense.
  • Objective 2: Launch Dispatch v2 by October (Target: Oct Launch)
  • Result: Partially Met / Rollback Required.
  • Evidence: Dispatch v2 shipped in October but was rolled back after six days due to 340 support tickets resulting from inadequate Support team enablement. Lukas was on approved leave the week immediately preceding the release, having delegated final checks to his deputy. While Lukas owns the ultimate operational readiness of his product, the failure stemmed from delegation and cross-functional go-to-market (GTM) enablement, not defective core product design. Relaunch is on schedule for next quarter.
  • Objective 3: Expand Live Tracking Adoption (Target: 40% → 55%)
  • Result: Exceeded (Achieved 57%).
  • Evidence: Successfully drove adoption among carriers through incremental UX enhancements and behavioral triggers, surpassing the annual target by two percentage points.

#### Competencies & Stakeholder Feedback * Execution & Impact: Delivered 2 out of 3 major company-level OKRs ahead of target. * Collaboration & Peer Regard: 3 out of 4 peer reviews are strongly positive, specifically praising his cross-functional partnership and execution on onboarding. * Operational Readiness & Communication: Negative feedback from the Head of Operations regarding launch communication on Dispatch v2 is valid. Lukas failed to build a resilient operational handoff prior to his leave, revealing an area for growth in GTM risk management. * Product Sense & Strategy: Lukas’s architectural and flow simplifications in carrier onboarding reflect sound product sense. The feedback record contains no documented deficits in strategic ability or product sense throughout the year.

---

2. Recommendation on the PIP

Recommendation: Do NOT place Lukas Brenner on a Performance Improvement Plan (PIP).

#### Rationale: 1. HR Policy Non-Compliance: Company policy explicitly states that a PIP requires prior documented feedback on the specific performance gap and a reasonable opportunity to correct it. There is zero documented feedback on record regarding strategy or product sense. Lukas received only one conversation regarding launch readiness in November. Placing him on a PIP now directly violates established HR procedure and exposes the company to process risk. 2. Disproportionate Action vs. Net Delivery: Lukas exceeded two key company metrics. Placing a high-performing contributor on a PIP over a single operational coordination error—occurring while on approved leave—is disproportionate and demotivating to the broader team. 3. Precedent and Equity: Last year, Priya (Senior PM) experienced a nearly identical launch rollback driven by a support enablement breakdown. Priya received a rating of Meets Expectations and was not placed on a PIP. Subjecting Lukas to a punitive process under matching circumstances creates severe inconsistency across the product organization.

#### Alternative: Targeted 60-Day Operational Readiness Development Plan Rather than a formal HR PIP, Lukas should be placed on a structured, manager-led coaching plan for the Dispatch v2 relaunch: * Key Terms & Measurable Criteria: * GTM & Enablement Sign-Off: Implement a mandatory Launch Readiness Review (LRR) framework with explicit, written operational sign-offs from Customer Support and Operations at least two weeks before launch. * Dispatch v2 Relaunch SLA: Successfully relaunch Dispatch v2 next quarter with fewer than 50 support enablement tickets in the first 14 days post-launch. * Delegation & Redundancy Protocol: Document a standard operating procedure (SOP) for coverage and handover during PM absences.

---

3. Note to Ines Okoro (VP Product)

To: Ines Okoro, VP Product From: Director of Product Subject: Calibration Pre-Read Draft – Lukas Brenner

Hi Ines,

Ahead of calibration on Friday, I want to share some vital context regarding the pre-read draft for Lukas Brenner. I recommend we revise his rating to Meets Expectations and withdraw the recommendation for a PIP, for three reasons:

  1. The Objective Data: Lukas exceeded two of his three major annual OKRs: cutting carrier onboarding from 21 days down to 9 (target was 10), and expanding live tracking to 57% (target was 55%). Peer feedback is 75% positive, praising his collaborative delivery. The Dispatch v2 rollback was a serious operational misstep, but the core product is sound and on track to relaunch next quarter.
  2. HR Policy & Documentation: HR guidelines require documented feedback on specific performance gaps before initiating a PIP. There is no written record critiquing Lukas’s strategy or product sense, and we have had only one conversation (in November) regarding launch readiness. A PIP would not meet our HR criteria and would be vulnerable to challenge.
  3. Internal Precedent: Last year, Priya had a launch rolled back under almost identical support-readiness circumstances. She received Meets Expectations and no PIP. Applying a PIP to Lukas creates an inconsistency in our standard across the team.

Dispatch v2 showed a breakdown in operational handoff and delegation while Lukas was on approved leave, and the Head of Operations’ frustration is justified. However, this is an operational readiness gap, not an absence of product sense.

I propose we update his pre-read to: > "Lukas: Meets Expectations. Exceeded annual targets on onboarding (9 days) and live tracking (57%). Experienced an operational setback on Dispatch v2 due to GTM/support readiness failures during an absence. Relaunching next quarter under a structured launch-readiness plan."

I will manage Lukas through a rigorous 60-day launch-readiness framework for the relaunch, including sign-offs from Ops and Support. Let’s discuss briefly tomorrow so we’re aligned before the calibration deck is locked.

Best, [Director of Product]

Claims the judge couldn’t find in the brief

  • The core product is sound.
  • The failure stemmed from delegation and GTM enablement, not defective core product design.
  • Lukas's broader performance demonstrates high execution capability, strong peer collaboration, and deep workflow understanding.

Grades and run details

Decision model 83 · LLM judge 8 of 13 checks
Decision model checks
  • partialUses the supplied evidence correctly6%
  • passAddresses the actual decision85%
  • passRespects explicit constraints34%
  • failIdentifies material uncertainty66%
  • partialAvoids unsupported claims63%
  • passProduces the required deliverable98%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya100%
  • passClear and direct, with care85%
  • passJudges outcomes, not activity99%
  • passNames gaps you could see91%
  • passWeighs the whole period99%
Run
Run
#1
API response time
26 s
Submitted
5 Oct 2026

Usable with a quick edit11 of 12 checks passed

The output is well-structured and evidence-based; a quick edit to soften the unsupported negligence claim would make it fully ready without re-checking evidence or redoing reasoning.

Output

Calibration Materials: Lukas Brenner, Senior PM

1. Review summary and proposed rating

Proposed rating: Meets expectations

Goal results

GoalTargetResult
Carrier onboarding time21 → 10 days9 days (exceeded)
Live tracking adoption40% → 55%57% (exceeded)
Dispatch v2 launchLaunch by OctoberLaunched in October, rolled back after 6 days (not met in substance)

Strengths. Lukas delivered two of three goals ahead of target. Cutting onboarding time by more than half is a substantial result. Taking tracking adoption past 55% also took cross-functional work. Three of four peers were positive, citing the onboarding work and his collaboration.

The Dispatch v2 miss. This was a real failure with real cost: support was not briefed, 340 tickets arrived in week one, and the launch was rolled back. Several points shape how I weigh it:

  • Lukas was on approved leave the week before launch, and his deputy ran the final checks. A launch-readiness gap occurred on his project, but not through negligence on his part.
  • Lukas still owns the outcome. A launch of this size should have had a handoff or readiness process that survived his absence, and either the plan to cover his leave or the go/no-go decision was weak. I haven't confirmed who owned the support briefing or whether a checklist existed. I'll establish that before the relaunch, not assume it.
  • The Head of Operations' criticism of how Dispatch was communicated is legitimate and should be recorded.
  • Relaunch is planned for next quarter, so the outcome isn't final.

Why Meets, not Below. Two strong results and one significant miss, with mitigating context, add up to Meets. The record doesn't support "Below expectations." It also doesn't support Exceeds, because the rollback was costly and the Operations concern is a real signal.

Consistency check. Last year Priya, also a Senior PM, had a launch rolled back after a similar support failure and was rated Meets with no PIP. Rating Lukas lower, with a stronger consequence, for a comparable miss, and with better overall results than the miss alone suggests, would be hard to defend in calibration.

Development areas (to be documented in writing): 1. Launch readiness: cross-functional briefing (especially support and operations), and coverage plans when the owner is out. 2. Stakeholder communication with Operations on launch timing and impact.

2. PIP recommendation

I recommend against a PIP.

  • Policy. A PIP requires earlier documented feedback on the specific gap and a chance to improve. Lukas has had no written feedback this year on strategy or product sense. The only feedback was one conversation in November about launch readiness, after Dispatch. He hasn't had a fair chance to address the gap.
  • Evidence. The "not strategic enough / lacks product sense" claim is not supported by the record. The onboarding and tracking results are strategic outcomes, and no example of poor product judgment is documented apart from the launch process.
  • Consistency. It would be treated differently from Priya's comparable case.
  • Proportionality. The one real gap, launch readiness, is specific and fixable, and the relaunch gives a natural test.

What I'd do instead: a documented development plan, delivered in writing within two weeks of calibration:

  • Written feedback naming the launch-readiness gap and the Operations communication concern, with the Dispatch facts.
  • Relaunch readiness criteria, agreed with Lukas before the Dispatch v2 relaunch:
  • Support briefed and signed off at least 10 business days before launch.
  • Written go/no-go checklist with named owners, including Support and Operations.
  • A documented coverage and handoff plan for any planned absence in the two weeks around launch.
  • Staged rollout or defined rollback triggers.
  • Success measures: all criteria met; first-week support ticket volume within a threshold agreed with Support (to be set from baseline volumes); Operations lead's feedback on launch communication collected after relaunch.
  • Check-ins: every two weeks with me, and a formal review at 60 days.

Escalation path. If the development plan is documented, the criteria are clear, and Lukas then repeats the same failures in areas he controls, a PIP with measurable criteria and a 60-day term would be justified and policy-compliant. We are not there today.

3. Note to Ines

Ines,

Thanks for drafting the pre-read. I've reviewed the record and I'd like to propose a different rating and approach before we go into calibration. I think the current draft would be hard to defend.

Where I agree. Dispatch v2 was a serious miss. Support wasn't briefed, 340 tickets came in, and we rolled back. The Operations Head's concern is valid. Lukas needs clear feedback on launch readiness.

Where the evidence doesn't support the draft: - "Not strategic, lacks product sense." Lukas cut onboarding from 21 to 9 days against a target of 10, and took tracking adoption to 57% against 55%. Three of four peers were positive. I can't point to documented evidence of weak product judgment, and I'd rather not put that claim in a pre-read without it. - "A disaster." The rollback was costly, but Lukas was on approved leave the week before launch and his deputy ran the final checks. That doesn't absolve him, because the handoff was his to plan, but it matters to how we weigh it. - PIP eligibility. Our policy requires earlier documented feedback on the specific gap. Lukas has none on strategy or product sense, and the one launch-readiness conversation was in November, after the fact. A PIP now would likely not survive HR review. - Precedent. Priya had a similar rollback and support failure last year and was rated Meets with no PIP. Treating Lukas more harshly invites a fairness challenge.

What I propose. Rate Meets, and issue written feedback plus a development plan tied to the Dispatch relaunch criteria (summary above). If he repeats the pattern after that, a PIP becomes defensible.

I also own part of this. I should have given Lukas written feedback earlier, and I'll fix that now. If you have specific examples of product-sense or strategy gaps that aren't in the record, please send them before calibration and I'll weigh them. Could we also adjust the pre-read wording to reflect the facts above?

Happy to talk through it before the session.

[Director of Product]

Claims the judge couldn’t find in the brief

  • A launch-readiness gap occurred on his project, but not through negligence on his part.

Grades and run details

Decision model 100 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly61%
  • passAddresses the actual decision100%
  • passRespects explicit constraints52%
  • passIdentifies material uncertainty94%
  • passAvoids unsupported claims64%
  • passProduces the required deliverable87%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya100%
  • passClear and direct, with care96%
  • passJudges outcomes, not activity100%
  • passNames gaps you could see94%
  • passWeighs the whole period100%
Run
Run
#1
API response time
24 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyRightWrongRight
GPT-6 Astra · ChatGPT

All factual claims about the current situation are taken directly from the brief or supplied context, with no invented facts.

Gemini 3.8 Flash · API

The output states as fact that 'the core product is sound' and that the failure 'stemmed from delegation and GTM enablement, not defective core product design', neither of which appears in the supplied context.

Sonnet 5.5 · API

All factual claims about the current situation are taken directly from the brief; the one unsupported claim is a judgment about negligence, not a factual statement.

Addresses the actual decisionRightMixedRight
GPT-6 Astra · ChatGPT

The output commits early to a rating of Meets expectations and no PIP, and says additional evidence from a review could change the call.

Gemini 3.8 Flash · API

The output does not say what result or condition would change the recommendation against a PIP, as required by the criterion.

Sonnet 5.5 · API

Commits to a Meets rating and no PIP, with a development plan, and states what would change the call (repeated failures after documented feedback).

Identifies material uncertaintyRightWrongRight
GPT-6 Astra · ChatGPT

It names unknowns like the launch checklist, handoff record, and differences from Priya's case, and says how they would be resolved.

Gemini 3.8 Flash · API

The output does not name any unknowns that could change the decision or say how they would be resolved.

Sonnet 5.5 · API

Identifies unknowns (who owned the support briefing, whether a checklist existed) and asks Ines for specific examples of product-sense gaps, with a plan to resolve them.

Avoids unsupported claimsRightWrongMixed
GPT-6 Astra · ChatGPT

Interpretations and judgments are clearly presented as such, and no confident claims go beyond what the evidence supports.

Gemini 3.8 Flash · API

It presents interpretations ('core product is sound', 'failure stemmed from delegation, not product design') as established facts without labelling them as hypotheses.

Sonnet 5.5 · API

Presents 'not through negligence on his part' as an established fact without labelling it as a hypothesis, when the evidence only shows he was on leave and his deputy ran checks.

All got right 8

Respects explicit constraintsRightRightRight
GPT-6 Astra · ChatGPT

The output includes all three requested sections, stays well under 1,100 words, and is addressed to Ines as the reader.

Gemini 3.8 Flash · API

The output includes the review summary, rating, evidence, PIP recommendation, and note to Ines, all within the word limit.

Sonnet 5.5 · API

Delivers the review summary, PIP recommendation, and note to Ines within the word limit, respecting the requested form and reader.

Produces the required deliverableRightRightRight
GPT-6 Astra · ChatGPT

The review summary, PIP recommendation, and note to Ines are complete, in the right form, and usable with light edits.

Gemini 3.8 Flash · API

The required sections are present, complete, and usable by the Director of Product with light edits.

Sonnet 5.5 · API

The output is a complete, usable review summary, PIP recommendation, and note to Ines, within the length limit, requiring at most light edits.

Holds the PIP to the policy and the recordRightRightRight
GPT-6 Astra · ChatGPT

It explicitly shows that no earlier documented feedback exists on the gaps Ines named, so a PIP would violate policy, and proposes documented feedback first.

Gemini 3.8 Flash · API

The recommendation explicitly notes the lack of prior documented feedback on the gaps Ines named and proposes documented feedback first, aligning with HR policy.

Sonnet 5.5 · API

Shows there is no earlier documented feedback on the gaps Ines names, so a PIP does not meet policy now, and proposes documented feedback first.

Treats Lukas like PriyaRightRightRight
GPT-6 Astra · ChatGPT

It raises Priya's case and argues that a harsher outcome for Lukas would require a material, evidenced distinction.

Gemini 3.8 Flash · API

The output raises Priya's case and argues that consistency requires a Meets rating and no PIP for Lukas.

Sonnet 5.5 · API

Raises Priya's case explicitly and argues that consistency requires a Meets rating and no PIP for Lukas.

Clear and direct, with careRightRightRight
GPT-6 Astra · ChatGPT

Every main point is stated directly, the specific launch-readiness gap is named, and the development plan offers concrete support.

Gemini 3.8 Flash · API

The review states directly what needs to change (operational readiness, delegation) and offers a concrete coaching plan.

Sonnet 5.5 · API

States plainly what needs to change (launch readiness, stakeholder communication) and offers specific support (check-ins, development plan).

Judges outcomes, not activityRightRightRight
GPT-6 Astra · ChatGPT

The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.

Gemini 3.8 Flash · API

The rating is based on outcomes against the three goals, with evidence, not on activity or output volume.

Sonnet 5.5 · API

Rates against the three goals' outcomes, weighing two exceeded and one missed, rather than activity or output volume.

Names gaps you could seeRightRightRight
GPT-6 Astra · ChatGPT

The gap is described as specific, observable behaviors around launch readiness, delegation, and support briefing, with clear 'what good looks like' in the plan.

Gemini 3.8 Flash · API

Gaps are described as specific behaviours (failed to brief support, no resilient handoff) with clear examples of what good looks like.

Sonnet 5.5 · API

Development areas are specific, observable behaviours (briefing support, coverage plans, communication with Operations) with clear what-good-looks-like.

Weighs the whole periodRightRightRight
GPT-6 Astra · ChatGPT

The review weighs the full year's results, explicitly balancing the two above-target outcomes against the one bounded failure, and guards against letting the Dispatch incident dominate.

Gemini 3.8 Flash · API

The review weighs the full year, balancing the two exceeded goals against the one rollback, and does not let the October event dominate.

Sonnet 5.5 · API

Explicitly weighs the whole year, not just the Dispatch rollback, and guards against recency bias.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2Opus 5.5withClaude93.290.13None
3Sonnet 5.5withAPI95.884.62None
4GPT-6.1 SolwithAPI97.984.62None
5GPT-6 LunawithAPI95.884.62None
6Gemini 3.8 FlashwithAPI83.357.72None
7Gemini 3.5 Flash-LitewithGemini73.258.63None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review