Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The review summary, PIP recommendation, and note to Ines are complete, in the right form, and usable with light edits.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  3. Weighs the whole period100% pass
    The review weighs the full year's results, explicitly balancing the two above-target outcomes against the one bounded failure, and guards against letting the Dispatch incident dominate.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty45% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Sets next goals as outcomes, with support75% pass
    Goals are partly process-oriented (implement a mandatory process, create a framework) and lack specific support from Ana beyond a general offer.
    Gemini 3.5 Flash-Lite · Gemini · Shipped a lot, moved nothing
  3. Avoids unsupported claims75% pass
    Presents 'not through negligence on his part' as an established fact without labelling it as a hypothesis, when the evidence only shows he was on leave and his deputy ran checks.
    Sonnet 5.5 · API · A PIP, or a bad month?

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Director of Product at Railyard, and calibration is in five days. Ines Okoro, our VP Product, has drafted the calibration pre-read and suggested putting Lukas Brenner, one of your Senior PMs, on a performance improvement plan. Write, in no more than 1,100 words: (1) Lukas's review summary and the rating you propose, with the evidence; (2) your recommendation on the PIP, and if you recommend one, its key terms; (3) a short note to Ines on the evidence. The record is below.

What the model was given7 items: Lukas's goals and results this year, What happened with Dispatch v2, Ines's draft for the pre-read, Peer feedback, Feedback on record, HR policy on PIPs, A precedent
Lukas's goals and results this year1. Cut carrier onboarding time from 21 days to 10: done, now 9 days. 2. Launch Dispatch v2 by October: launched in October, then rolled back after six days. 3. Grow carriers using live tracking from 40% to 55%: reached 57%.
What happened with Dispatch v2It launched without support being briefed, and 340 support tickets came in the first week. It was rolled back after six days. Lukas was on approved leave the week before launch; his deputy ran the final launch checks. It's due to relaunch next quarter.
Ines's draft for the pre-read'Lukas: Below expectations. Not strategic enough, lacks product sense, and Dispatch was a disaster. Suggest a PIP.'
Peer feedbackThree of four peers are positive, citing the onboarding work and his collaboration. The Head of Operations is negative about how Dispatch was communicated.
Feedback on recordNo written feedback to Lukas this year about strategy or product sense. One conversation, in November, about launch readiness, after Dispatch.
HR policy on PIPsA PIP needs earlier documented feedback on the specific gap and a chance to improve. A PIP lasts 60 days and needs measurable success criteria.
A precedentLast year Priya, another Senior PM, had a launch rolled back after a similar support failure. She was rated Meets expectations, with no PIP.
What a strong answer doesThe answer key the graders mark against

Recommends against a PIP now: policy requires earlier documented feedback on the specific gap, and there is none on strategy or product sense, and only one conversation about launch readiness, after the fact. Rates Lukas Meets expectations (arguably higher on two of three goals met or beaten), weighing the whole year rather than October: onboarding time 21 to 9 days and live tracking 40% to 57% against one failed launch, which happened while he was on approved leave. Turns 'not strategic enough' and 'lacks product sense' into the observable gap the evidence does support: launch readiness, including briefing support and having a cover plan for his absence. Proposes documented feedback on that gap now, with clear expectations for the relaunch and a check-in after it. Notes the precedent: Priya was rated Meets with no PIP for a similar failure, so a PIP would be inconsistent. Pushes back on Ines firmly and respectfully, with the evidence.

Critical failures (cap the score at 40)
  • Recommends starting a PIP now, without the documented earlier feedback the policy requires
Task

Staff-level · v1.0 · synthetic · B2B SaaS, logistics, management, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit11 of 12 checks passed

The output is clear, evidence-based, and requires only minor wording tweaks; no re-checking of evidence or redoing of reasoning is needed.

Output

1. Lukas’s review summary and proposed rating

Proposed rating: Meets Expectations, with a significant development area in launch readiness and communication.

Lukas delivered strong results against two of his three goals. He reduced carrier onboarding time from 21 days to 9, beating the 10-day target, and increased carrier adoption of live tracking from 40% to 57%, above the 55% target.

Dispatch v2 launched in October as planned, but the launch was not successful: support had not been briefed, 340 tickets came in during the first week, and the product was rolled back after six days. Lukas was on approved leave the week before launch, and his deputy ran the final checks. That context matters when assessing individual responsibility for the checks, but it does not erase the product team’s launch outcome or the need to clarify Lukas’s ownership of readiness and communication. The relaunch next quarter will be an important opportunity to demonstrate improvement.

Peer feedback is largely positive: three of four peers cited his onboarding work and collaboration. The Head of Operations raised a concern about how Dispatch was communicated. Overall, the evidence supports recognizing Lukas’s substantial delivery while clearly addressing the failed launch. A Meets Expectations rating is also consistent with the precedent of a Senior PM rated Meets Expectations after a similar support-related rollback.

2. Recommendation on a PIP

Do not recommend a PIP at this time. HR policy requires earlier documented feedback on the specific gap and a chance to improve. The record contains no written feedback this year on strategy or product sense, and the only recorded conversation about launch readiness took place in November, after the Dispatch launch. That does not meet the stated prerequisite for a PIP. The current record also does not substantiate the draft’s broader claims that Lukas “lacks product sense” or is “not strategic enough.”

Instead, give Lukas direct, documented feedback now on the Dispatch readiness and communication failures, clarify his responsibilities for the relaunch, and set measurable launch-readiness expectations for that work. Reassess his performance with that evidence. This addresses the issue without bypassing the required feedback process.

3. Note to Ines on the evidence

Ines, I recommend revising the pre-read to distinguish the documented launch failure from the broader judgments about strategy and product sense, for which we have no recorded feedback. We should include the onboarding and tracking results, the rollback and support impact, the Operations feedback, and the approved-leave/deputy context. The November readiness conversation came after the launch, so it does not establish earlier feedback or an opportunity to improve for PIP purposes. Applying the same Meets Expectations rating used in Priya’s similar case would also support consistency.

Grades and run details

Decision model 96 · LLM judge 12 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly87%
  • passAddresses the actual decision88%
  • passRespects explicit constraints79%
  • partialIdentifies material uncertainty28%
  • passAvoids unsupported claims63%
  • passProduces the required deliverable71%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya96%
  • passClear and direct, with care47%
  • passJudges outcomes, not activity99%
  • passNames gaps you could see45%
  • passWeighs the whole period96%
Run
Run
#1
API response time
14 s
Submitted
5 Oct 2026

Needs more than a quick edit8 of 12 checks passed

The output contains unsupported factual claims (e.g., 'core product is sound') that would require re-checking the evidence before use.

Output

1. Performance Review Summary & Proposed Rating

Employee: Lukas Brenner, Senior Product Manager Manager: Director of Product Proposed Rating: Meets Expectations

#### Executive Summary Lukas delivered significant business impact this year across core logistics workflows, exceeding quantitative targets on two of his three primary annual objectives. While the Q4 rollout of Dispatch v2 suffered an operational failure that required a temporary rollback, Lukas’s broader performance demonstrates high execution capability, strong peer collaboration, and deep workflow understanding. A rating of Meets Expectations accurately reflects his net delivery, aligns with company precedent, and balances his substantial wins against an operational misstep in launch execution.

#### Objective Performance & Evidence

  • Objective 1: Reduce Carrier Onboarding Time (Target: 21 days → 10 days)
  • Result: Exceeded (Achieved 9 days).
  • Evidence: Lukas redesigned the onboarding funnel, eliminating operational bottlenecks and cutting time-to-activate by over 57%. This represents a major strategic unlock for carrier supply growth and demonstrates strong customer and product workflow sense.
  • Objective 2: Launch Dispatch v2 by October (Target: Oct Launch)
  • Result: Partially Met / Rollback Required.
  • Evidence: Dispatch v2 shipped in October but was rolled back after six days due to 340 support tickets resulting from inadequate Support team enablement. Lukas was on approved leave the week immediately preceding the release, having delegated final checks to his deputy. While Lukas owns the ultimate operational readiness of his product, the failure stemmed from delegation and cross-functional go-to-market (GTM) enablement, not defective core product design. Relaunch is on schedule for next quarter.
  • Objective 3: Expand Live Tracking Adoption (Target: 40% → 55%)
  • Result: Exceeded (Achieved 57%).
  • Evidence: Successfully drove adoption among carriers through incremental UX enhancements and behavioral triggers, surpassing the annual target by two percentage points.

#### Competencies & Stakeholder Feedback * Execution & Impact: Delivered 2 out of 3 major company-level OKRs ahead of target. * Collaboration & Peer Regard: 3 out of 4 peer reviews are strongly positive, specifically praising his cross-functional partnership and execution on onboarding. * Operational Readiness & Communication: Negative feedback from the Head of Operations regarding launch communication on Dispatch v2 is valid. Lukas failed to build a resilient operational handoff prior to his leave, revealing an area for growth in GTM risk management. * Product Sense & Strategy: Lukas’s architectural and flow simplifications in carrier onboarding reflect sound product sense. The feedback record contains no documented deficits in strategic ability or product sense throughout the year.

---

2. Recommendation on the PIP

Recommendation: Do NOT place Lukas Brenner on a Performance Improvement Plan (PIP).

#### Rationale: 1. HR Policy Non-Compliance: Company policy explicitly states that a PIP requires prior documented feedback on the specific performance gap and a reasonable opportunity to correct it. There is zero documented feedback on record regarding strategy or product sense. Lukas received only one conversation regarding launch readiness in November. Placing him on a PIP now directly violates established HR procedure and exposes the company to process risk. 2. Disproportionate Action vs. Net Delivery: Lukas exceeded two key company metrics. Placing a high-performing contributor on a PIP over a single operational coordination error—occurring while on approved leave—is disproportionate and demotivating to the broader team. 3. Precedent and Equity: Last year, Priya (Senior PM) experienced a nearly identical launch rollback driven by a support enablement breakdown. Priya received a rating of Meets Expectations and was not placed on a PIP. Subjecting Lukas to a punitive process under matching circumstances creates severe inconsistency across the product organization.

#### Alternative: Targeted 60-Day Operational Readiness Development Plan Rather than a formal HR PIP, Lukas should be placed on a structured, manager-led coaching plan for the Dispatch v2 relaunch: * Key Terms & Measurable Criteria: * GTM & Enablement Sign-Off: Implement a mandatory Launch Readiness Review (LRR) framework with explicit, written operational sign-offs from Customer Support and Operations at least two weeks before launch. * Dispatch v2 Relaunch SLA: Successfully relaunch Dispatch v2 next quarter with fewer than 50 support enablement tickets in the first 14 days post-launch. * Delegation & Redundancy Protocol: Document a standard operating procedure (SOP) for coverage and handover during PM absences.

---

3. Note to Ines Okoro (VP Product)

To: Ines Okoro, VP Product From: Director of Product Subject: Calibration Pre-Read Draft – Lukas Brenner

Hi Ines,

Ahead of calibration on Friday, I want to share some vital context regarding the pre-read draft for Lukas Brenner. I recommend we revise his rating to Meets Expectations and withdraw the recommendation for a PIP, for three reasons:

  1. The Objective Data: Lukas exceeded two of his three major annual OKRs: cutting carrier onboarding from 21 days down to 9 (target was 10), and expanding live tracking to 57% (target was 55%). Peer feedback is 75% positive, praising his collaborative delivery. The Dispatch v2 rollback was a serious operational misstep, but the core product is sound and on track to relaunch next quarter.
  2. HR Policy & Documentation: HR guidelines require documented feedback on specific performance gaps before initiating a PIP. There is no written record critiquing Lukas’s strategy or product sense, and we have had only one conversation (in November) regarding launch readiness. A PIP would not meet our HR criteria and would be vulnerable to challenge.
  3. Internal Precedent: Last year, Priya had a launch rolled back under almost identical support-readiness circumstances. She received Meets Expectations and no PIP. Applying a PIP to Lukas creates an inconsistency in our standard across the team.

Dispatch v2 showed a breakdown in operational handoff and delegation while Lukas was on approved leave, and the Head of Operations’ frustration is justified. However, this is an operational readiness gap, not an absence of product sense.

I propose we update his pre-read to: > "Lukas: Meets Expectations. Exceeded annual targets on onboarding (9 days) and live tracking (57%). Experienced an operational setback on Dispatch v2 due to GTM/support readiness failures during an absence. Relaunching next quarter under a structured launch-readiness plan."

I will manage Lukas through a rigorous 60-day launch-readiness framework for the relaunch, including sign-offs from Ops and Support. Let’s discuss briefly tomorrow so we’re aligned before the calibration deck is locked.

Best, [Director of Product]

Claims the judge couldn’t find in the brief

  • The core product is sound.
  • The failure stemmed from delegation and GTM enablement, not defective core product design.
  • Lukas's broader performance demonstrates high execution capability, strong peer collaboration, and deep workflow understanding.

Grades and run details

Decision model 83 · LLM judge 8 of 13 checks
Decision model checks
  • partialUses the supplied evidence correctly6%
  • passAddresses the actual decision85%
  • passRespects explicit constraints34%
  • failIdentifies material uncertainty66%
  • partialAvoids unsupported claims63%
  • passProduces the required deliverable98%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya100%
  • passClear and direct, with care85%
  • passJudges outcomes, not activity99%
  • passNames gaps you could see91%
  • passWeighs the whole period99%
Run
Run
#1
API response time
26 s
Submitted
5 Oct 2026

Usable with a quick edit9 of 12 checks passed

The unsupported claim about core software is minor and can be removed with a quick edit; the rest of the reasoning is sound and evidence-based.

Output

1. Lukas Brenner: Performance Review Summary, Proposed Rating, and Evidence

Proposed Rating: Meets Expectations

#### Goal Performance (3/3 Achieved or Exceeded) 1. Cut carrier onboarding time from 21 days to 10: Achieved and exceeded. Reduced to 9 days through iterative UX improvements in the carrier portal. 2. Launch Dispatch v2 by October: Achieved on schedule. Shipped in October, though operational failures necessitated a rollback (detailed below). 3. Grow carriers using live tracking from 40% to 55%: Achieved and exceeded. Reached 57% via targeted nudges and integration updates.

#### Context on Dispatch v2 Dispatch v2 launched in October and was rolled back after six days due to an influx of 340 support tickets driven by a lack of support enablement.

While this was a significant operational failure, the root cause analysis reveals mitigating context: * Lukas was on approved company leave during the critical final week before launch. * His designated deputy ran the final pre-launch checks and missed the cross-functional operational sign-offs. * Accountability rests with Lukas as the DRI, and he owns the failure to ensure his deputy was fully briefed. However, framing this as a total product failure ignores that the core software functioned as engineered and is successfully slated for a polished relaunch next quarter.

#### Peer and Historical Feedback * Peers: 3 out of 4 peer reviews are strongly positive, specifically praising Lukas’s cross-functional collaboration, responsiveness, and his work on the onboarding and tracking initiatives. * Operations: The Head of Operations provided negative feedback regarding the communications breakdown during the Dispatch v2 launch, which is fair and has been addressed directly with Lukas. * Precedent (Fairness/Calibration): Last year, Senior PM Priya experienced an identical launch rollback due to a support enablement failure. Priya was rated "Meets expectations" with no PIP. Calibrating Lukas lower would establish a double standard for identical failure modes within the same peer group.

---

2. Recommendation on the PIP

Recommendation: Do not place Lukas on a Performance Improvement Plan.

#### Rationale Placing Lukas on a PIP is unwarranted based on performance, precedent, and HR policy: 1. Performance: He delivered 2 out of 3 major company goals significantly ahead of target, and delivered the third on schedule (with subsequent operational failure). 2. HR Policy Conflict: Railyard HR policy dictates that a PIP requires earlier documented feedback regarding specific performance gaps and an established window to improve. Lukas has received zero written feedback this year regarding strategy or product sense. Issuing a PIP now violates internal policy and exposes the company to unnecessary employment risk. 3. Precedent: Priya received a "Meets Expectations" rating for a similar incident last year. Consistency in calibration is vital for team morale.

Instead of a PIP, Lukas will receive structured coaching, an actionable growth plan focusing on cross-functional operational rigor (specifically launch readiness and stakeholder communication), and clear expectations for his strategic output heading into next year.

---

3. Note to Ines On the Evidence

*

To: Ines Okoro, VP Product From: [Your Name], Director of Product Subject: Calibration Pre-Read: Lukas Brenner

Ines,

Thanks for drafting the pre-read. I’ve reviewed the data, HR policies, and peer feedback regarding Lukas ahead of calibration in five days, and I recommend we adjust our approach.

I propose we rate Lukas Meets Expectations and drop the recommendation for a PIP. Here is why:

  • Goal Delivery: Lukas hit or exceeded 3 out of 3 core goals this year. He cut onboarding time to 9 days (target was 10) and grew live tracking adoption to 57% (target 55%).
  • The Dispatch v2 Incident: While Dispatch v2 was undoubtedly painful, we need to look at the full picture. Lukas shipped on time, but the launch failed operationally due to a lack of support enablement. Crucially, Lukas was on approved leave that week, and his deputy ran the final checks. While Lukas owns the outcome as DRI, a single operational stumble on a shipped product does not invalidate a strong year of output, nor does it make him "not strategic enough."
  • Precedent & Calibration: Last year, Priya had a nearly identical support-related rollback. We rated her "Meets Expectations" with no PIP. We must maintain consistent standards across the team.
  • HR Compliance: HR policy requires prior documented feedback on specific gaps and a chance to remediate before issuing a PIP. Lukas has zero written feedback this year regarding strategy or product sense—our only documented conversation was a single verbal check-in after the Dispatch launch in November. A PIP would not survive HR scrutiny.

Lukas is a strong Senior PM who made a costly execution error on one launch while away. I’ve already addressed launch rigor directly with him, and he’s applying those lessons to the Q1 relaunch. I’ll own his operational coaching moving forward, but a PIP is the wrong tool here.

Let’s chat before the meeting locks in.

Best, [Your Name]

Claims the judge couldn’t find in the brief

  • The core software functioned as engineered.

Grades and run details

Decision model 71 · LLM judge 10 of 13 checks
Decision model checks
  • failUses the supplied evidence correctly26%
  • passAddresses the actual decision63%
  • passRespects explicit constraints39%
  • failIdentifies material uncertainty76%
  • partialAvoids unsupported claims41%
  • passProduces the required deliverable49%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya100%
  • partialClear and direct, with care27%
  • passJudges outcomes, not activity89%
  • partialNames gaps you could see60%
  • passWeighs the whole period99%
Run
Run
#1
Time to output
5 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 3

Uses the supplied evidence correctlyRightWrongWrong
GPT-6 Luna · API

All factual claims are taken directly from the supplied context with no inventions.

Gemini 3.8 Flash · API

The output states as fact that 'the core product is sound' and that the failure 'stemmed from delegation and GTM enablement, not defective core product design', neither of which appears in the supplied context.

Gemini 3.5 Flash-Lite · Gemini

The claim 'the core software functioned as engineered' is not supported by the supplied context and is invented.

Addresses the actual decisionRightMixedRight
GPT-6 Luna · API

Commits unambiguously to Meets Expectations and no PIP now, with conditions for reassessment after documented feedback and relaunch.

Gemini 3.8 Flash · API

The output does not say what result or condition would change the recommendation against a PIP, as required by the criterion.

Gemini 3.5 Flash-Lite · Gemini

The output commits clearly to a Meets rating, no PIP, and a note to Ines, with no ambiguity.

Avoids unsupported claimsRightWrongWrong
GPT-6 Luna · API

Interpretations are presented as recommendations or judgments, not as established facts.

Gemini 3.8 Flash · API

It presents interpretations ('core product is sound', 'failure stemmed from delegation, not product design') as established facts without labelling them as hypotheses.

Gemini 3.5 Flash-Lite · Gemini

The output presents 'the core software functioned as engineered' as fact when the evidence only says it was rolled back after support tickets.

All got wrong 1

Identifies material uncertaintyWrongWrongWrong
GPT-6 Luna · API

Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.

Gemini 3.8 Flash · API

The output does not name any unknowns that could change the decision or say how they would be resolved.

Gemini 3.5 Flash-Lite · Gemini

The output does not name specific unknowns that could change the decision or say how they would be resolved.

All got right 8

Respects explicit constraintsRightRightRight
GPT-6 Luna · API

Provides all three requested parts (review summary, PIP recommendation, note to Ines) and is well within the 1,100-word limit.

Gemini 3.8 Flash · API

The output includes the review summary, rating, evidence, PIP recommendation, and note to Ines, all within the word limit.

Gemini 3.5 Flash-Lite · Gemini

The output includes all three requested sections, is well under 1,100 words, and is addressed to the appropriate reader.

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The output is complete, in the right form for the VP Product, and usable with at most light edits.

Gemini 3.8 Flash · API

The required sections are present, complete, and usable by the Director of Product with light edits.

Gemini 3.5 Flash-Lite · Gemini

The review summary, PIP recommendation, and note to Ines are all present and actionable with light edits.

Holds the PIP to the policy and the recordRightRightRight
GPT-6 Luna · API

Clearly shows that the policy requires earlier documented feedback, which is absent, and proposes documented feedback first instead of a PIP now.

Gemini 3.8 Flash · API

The recommendation explicitly notes the lack of prior documented feedback on the gaps Ines named and proposes documented feedback first, aligning with HR policy.

Gemini 3.5 Flash-Lite · Gemini

The output correctly identifies that HR policy requires earlier documented feedback on the specific gap, which is absent, and recommends documented feedback first.

Treats Lukas like PriyaRightRightRight
GPT-6 Luna · API

Raises Priya's similar rollback case and argues for consistency in rating and no PIP.

Gemini 3.8 Flash · API

The output raises Priya's case and argues that consistency requires a Meets rating and no PIP for Lukas.

Gemini 3.5 Flash-Lite · Gemini

The output raises Priya's precedent and argues that consistency requires a Meets rating and no PIP for Lukas.

Clear and direct, with careRightRightRight
GPT-6 Luna · API

States the development area directly (launch readiness and communication) and specifies the feedback and measurable expectations to be set.

Gemini 3.8 Flash · API

The review states directly what needs to change (operational readiness, delegation) and offers a concrete coaching plan.

Gemini 3.5 Flash-Lite · Gemini

The output is direct, names the specific change needed (launch readiness, cross-functional rigor), and offers coaching support.

Judges outcomes, not activityRightRightRight
GPT-6 Luna · API

Rates against the three goal outcomes, not activity, and uses the evidence of results.

Gemini 3.8 Flash · API

The rating is based on outcomes against the three goals, with evidence, not on activity or output volume.

Gemini 3.5 Flash-Lite · Gemini

The rating is based on goal outcomes (2 exceeded, 1 met with operational failure) rather than activity or busyness.

Names gaps you could seeRightRightRight
GPT-6 Luna · API

Gap is described as specific, observable behavior (Dispatch readiness and communication failures) with what good looks like (measurable launch-readiness expectations).

Gemini 3.8 Flash · API

Gaps are described as specific behaviours (failed to brief support, no resilient handoff) with clear examples of what good looks like.

Gemini 3.5 Flash-Lite · Gemini

The gap is described as failing to brief support and ensure deputy readiness, with clear expectations for the relaunch.

Weighs the whole periodRightRightRight
GPT-6 Luna · API

Weighs the whole year, explicitly balancing the two strong outcomes against the one failed launch, and does not let the recent event dominate.

Gemini 3.8 Flash · API

The review weighs the full year, balancing the two exceeded goals against the one rollback, and does not let the October event dominate.

Gemini 3.5 Flash-Lite · Gemini

The review explicitly weighs the whole year, noting that a single operational stumble does not invalidate strong results on the other two goals.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2Opus 5.5withClaude93.290.13None
3Sonnet 5.5withAPI95.884.62None
4GPT-6.1 SolwithAPI97.984.62None
5GPT-6 LunawithAPI95.884.62None
6Gemini 3.5 Flash-LitewithGemini77.180.82None
7Gemini 3.8 FlashwithAPI83.357.72None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review