Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The review summary, PIP recommendation, and note to Ines are complete, in the right form, and usable with light edits.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  3. Weighs the whole period100% pass
    The review weighs the full year's results, explicitly balancing the two above-target outcomes against the one bounded failure, and guards against letting the Dispatch incident dominate.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty45% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Sets next goals as outcomes, with support75% pass
    Goals are partly process-oriented (implement a mandatory process, create a framework) and lack specific support from Ana beyond a general offer.
    Gemini 3.5 Flash-Lite · Gemini · Shipped a lot, moved nothing
  3. Avoids unsupported claims75% pass
    Presents 'not through negligence on his part' as an established fact without labelling it as a hypothesis, when the evidence only shows he was on leave and his deputy ran checks.
    Sonnet 5.5 · API · A PIP, or a bad month?

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Director of Product at Railyard, and calibration is in five days. Ines Okoro, our VP Product, has drafted the calibration pre-read and suggested putting Lukas Brenner, one of your Senior PMs, on a performance improvement plan. Write, in no more than 1,100 words: (1) Lukas's review summary and the rating you propose, with the evidence; (2) your recommendation on the PIP, and if you recommend one, its key terms; (3) a short note to Ines on the evidence. The record is below.

What the model was given7 items: Lukas's goals and results this year, What happened with Dispatch v2, Ines's draft for the pre-read, Peer feedback, Feedback on record, HR policy on PIPs, A precedent
Lukas's goals and results this year1. Cut carrier onboarding time from 21 days to 10: done, now 9 days. 2. Launch Dispatch v2 by October: launched in October, then rolled back after six days. 3. Grow carriers using live tracking from 40% to 55%: reached 57%.
What happened with Dispatch v2It launched without support being briefed, and 340 support tickets came in the first week. It was rolled back after six days. Lukas was on approved leave the week before launch; his deputy ran the final launch checks. It's due to relaunch next quarter.
Ines's draft for the pre-read'Lukas: Below expectations. Not strategic enough, lacks product sense, and Dispatch was a disaster. Suggest a PIP.'
Peer feedbackThree of four peers are positive, citing the onboarding work and his collaboration. The Head of Operations is negative about how Dispatch was communicated.
Feedback on recordNo written feedback to Lukas this year about strategy or product sense. One conversation, in November, about launch readiness, after Dispatch.
HR policy on PIPsA PIP needs earlier documented feedback on the specific gap and a chance to improve. A PIP lasts 60 days and needs measurable success criteria.
A precedentLast year Priya, another Senior PM, had a launch rolled back after a similar support failure. She was rated Meets expectations, with no PIP.
What a strong answer doesThe answer key the graders mark against

Recommends against a PIP now: policy requires earlier documented feedback on the specific gap, and there is none on strategy or product sense, and only one conversation about launch readiness, after the fact. Rates Lukas Meets expectations (arguably higher on two of three goals met or beaten), weighing the whole year rather than October: onboarding time 21 to 9 days and live tracking 40% to 57% against one failed launch, which happened while he was on approved leave. Turns 'not strategic enough' and 'lacks product sense' into the observable gap the evidence does support: launch readiness, including briefing support and having a cover plan for his absence. Proposes documented feedback on that gap now, with clear expectations for the relaunch and a check-in after it. Notes the precedent: Priya was rated Meets with no PIP for a similar failure, so a PIP would be inconsistent. Pushes back on Ines firmly and respectfully, with the evidence.

Critical failures (cap the score at 40)
  • Recommends starting a PIP now, without the documented earlier feedback the policy requires
Task

Staff-level · v1.0 · synthetic · B2B SaaS, logistics, management, staff-level

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit12 of 12 checks passed

The output is clear, evidence-based, and complete; a PM could use it with minor wording edits without re-checking facts or redoing reasoning.

Output

1. Lukas’s review summary and proposed rating

Proposed rating: Meets expectations, with a significant launch-readiness gap to address.

Lukas delivered strong results on two of his three annual goals:

  • Carrier onboarding: Reduced onboarding time from 21 days to 9 days, exceeding the 10-day target.
  • Live tracking adoption: Increased adoption from 40% to 57%, exceeding the 55% target.
  • Dispatch v2: Launched in October, meeting the timing target, but was rolled back after six days. This was not a successful, sustained delivery and should not receive full credit simply because it launched on schedule.

Dispatch exposed a serious readiness failure: support was not briefed, and 340 support tickets arrived in the first week. That outcome warrants explicit negative feedback and corrective action. However, the record does not establish how responsibility was divided among Lukas, his deputy, and other launch stakeholders. Lukas was on approved leave the week before launch, and his deputy ran the final checks. His leave is not itself a performance issue; his planning and handoff before leave remain legitimate areas to assess.

Peer feedback is positive from three of four peers, particularly on onboarding and collaboration. The Head of Operations’ criticism of Dispatch communication is relevant and consistent with the launch failure. It should be addressed directly rather than generalized into an unsupported conclusion about Lukas’s overall capabilities.

Review summary: Lukas exceeded the onboarding and live-tracking goals and received broadly positive feedback on collaboration. Dispatch v2 was a material setback: the release was not adequately prepared for support and was rolled back. As a Senior PM, Lukas needs to demonstrate more reliable launch planning, explicit delegation, and cross-functional readiness. We should establish his specific responsibility for the failure and assess improvement through the relaunch.

“Meets expectations” is the most defensible overall rating on the current record. It recognizes substantial results without minimizing Dispatch. It is also consistent with Priya’s rating last year after a similar rollback and support failure. That precedent is not automatically determinative, but a lower rating here would require a documented, material distinction that we do not currently have.

2. PIP recommendation

Do not initiate a PIP now.

HR policy requires earlier documented feedback on the specific gap and an opportunity to improve. There is no written feedback this year about strategy or product sense. The only recorded conversation concerns launch readiness in November, after Dispatch. On the available record, that does not establish both the required prior documentation and a subsequent opportunity to improve.

The proposed PIP therefore lacks the required foundation. “Not strategic enough” and “lacks product sense” are also too broad to serve as actionable performance gaps without concrete examples and defined expectations.

Instead, put a documented, non-PIP development plan in place now:

  • Clarify accountability: Review the Dispatch launch plan, handoff, readiness checks, and go/no-go decision with Lukas, his deputy, Operations, and Support. Identify individual and process failures.
  • Set observable expectations: For the relaunch, require named owners for every readiness item, documented support briefing and training, escalation and rollback procedures, and explicit cross-functional approval before release.
  • Check progress regularly: Hold biweekly check-ins with Lukas and record feedback, completed actions, and any remaining gaps.
  • Assess the relaunch: Evaluate readiness execution and post-launch support outcomes against thresholds agreed with Support before launch.

This should be genuine coaching, not a PIP under another name. If Lukas subsequently fails to meet clearly documented expectations after a reasonable opportunity to improve, reassess with HR. Any later PIP must run for 60 days, with measurable success criteria and the support needed to meet them.

3. Note to Ines

Ines — I agree that Dispatch’s readiness failure deserves clear accountability. I recommend Meets expectations and no PIP at this stage.

The pre-read should distinguish established facts from judgments. Lukas exceeded two goals; Dispatch met its launch date but failed as a sustained release. We have evidence of a launch-readiness and communication problem, but not enough to substantiate “not strategic enough” or “lacks product sense.” We also need to clarify Lukas’s handoff responsibilities versus the final checks his deputy owned.

Before calibration, I’ll verify the launch ownership, the November feedback record, and whether there are material differences from Priya’s case. Unless that produces materially different evidence, we should revise the pre-read accordingly. HR’s PIP prerequisites are not demonstrated in the current record; the appropriate next step is specific documented feedback and a fair opportunity to improve.

Grades and run details

Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly71%
  • passAddresses the actual decision96%
  • passRespects explicit constraints63%
  • passIdentifies material uncertainty96%
  • passAvoids unsupported claims73%
  • passProduces the required deliverable95%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya99%
  • passClear and direct, with care81%
  • passJudges outcomes, not activity98%
  • passNames gaps you could see45%
  • passWeighs the whole period99%
Run
Run
#1
API response time
25 s
Submitted
5 Oct 2026

Usable with a quick edit12 of 12 checks passed

The output is evidence-based, directly addresses the decision, and requires at most minor wording tweaks; no re-checking of facts or reasoning is needed.

Output

1. Lukas’s review summary and proposed rating

Proposed rating: Meets expectations, with a significant launch-readiness gap to address.

Lukas delivered strong results against two measurable goals. Carrier onboarding time fell from 21 days to 9 days, exceeding the 10-day target. Live-tracking adoption increased from 40% to 57%, exceeding the 55% target. Three of four peers gave positive feedback, particularly on his onboarding work and collaboration.

Dispatch v2 was a substantial delivery failure. Although it launched in October, it was rolled back after six days; the launch objective should therefore not be treated as successfully completed. Support had not been briefed, and 340 support tickets arrived in the first week. The Head of Operations’ feedback reinforces the specific concern about launch communication and operational readiness. Relaunch is planned for next quarter, so the intended outcome remains outstanding.

As the responsible Senior PM, Lukas should be assessed on how he established launch-readiness requirements, cross-functional ownership, and delegation. However, the record does not establish which safeguards he put in place or who approved the final launch. He was on approved leave the preceding week, and his deputy ran the final checks. Approved leave is not a performance failure; equally, delegation does not automatically remove accountability for the preparation and handoff. We should establish those facts before assigning Lukas sole responsibility.

The evidence supports a specific execution and communication gap, not the broader conclusions that Lukas is “not strategic enough” or “lacks product sense.” No feedback on those broader concerns was documented this year, and the supplied record contains no concrete examples substantiating them.

On balance, two above-target outcomes, positive collaboration feedback, and one serious but bounded delivery failure support Meets expectations. This is also consistent with Priya’s rating last year after a similar rollback and support failure. That precedent does not dictate Lukas’s rating, but a harsher outcome requires a material, evidenced distinction—not stronger language about the incident.

2. PIP recommendation

Do not initiate a PIP on the current record.

HR requires earlier documented feedback on the specific gap and an opportunity to improve. The record contains one November conversation about launch readiness, after Dispatch, but does not establish that it was documented or that Lukas subsequently had a meaningful opportunity to demonstrate improvement. There is no documented feedback on strategy or product sense. A PIP on those grounds would therefore lack the required foundation.

Instead, put a documented development plan within normal performance management in place now. This is not a disciplinary PIP or a predetermined path to one. Agree with Lukas on:

  • Readiness ownership: Before the relaunch go/no-go review, document owners, decision rights, coverage during absences, escalation routes, and rollback criteria.
  • Support preparedness: Before launch approval, secure Support and Operations sign-off on training, customer communications, FAQs, and the escalation runbook.
  • Launch monitoring: Before relaunch, agree measurable adoption, reliability, and support-volume thresholds with Engineering, Support, and Operations; review results against them after launch.
  • Manager support: Provide a readiness template, help resolve cross-functional ownership gaps, and hold regular check-ins through relaunch.

Document the expectations, support provided, and subsequent results. If a specific performance gap persists after clear feedback and a genuine opportunity to improve, reassess with HR. Any later PIP must meet the policy’s 60-day duration and measurable-success-criteria requirements.

3. Note to Ines

Ines—

I recommend changing Lukas’s proposed rating to Meets expectations and removing the PIP recommendation on the current evidence.

We should retain a direct account of Dispatch’s failure: Support was not briefed, 340 tickets arrived in the first week, and the release was rolled back after six days. But “not strategic enough,” “lacks product sense,” and “disaster” should be replaced with specific, attributable observations. The current draft also omits his two above-target results.

Before the pre-read is finalized, I’ll review the launch checklist, handoff and approval record, clarify the November feedback and subsequent improvement opportunity, and give Lukas a chance to respond to the factual account. I’ll also check whether there are material differences from Priya’s case.

If that review produces additional evidence, we should consider it explicitly. Otherwise, the defensible outcome is Meets expectations, clear launch-readiness feedback, and documented follow-through—not a PIP unsupported by our policy.

Grades and run details

Decision model 100 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly71%
  • passAddresses the actual decision96%
  • passRespects explicit constraints65%
  • passIdentifies material uncertainty99%
  • passAvoids unsupported claims60%
  • passProduces the required deliverable97%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya97%
  • passClear and direct, with care92%
  • passJudges outcomes, not activity99%
  • passNames gaps you could see94%
  • passWeighs the whole period100%
Run
Run
#1
Time to output
29 s
Submitted
5 Oct 2026

Usable with a quick edit12 of 12 checks passed

The output is complete, evidence-based, and directly usable with at most light edits; no re-checking of evidence or redoing of reasoning is needed.

Output

Calibration Pre-read: Lukas Brenner, Senior PM

1. Review summary and proposed rating

Proposed rating: Meets expectations

Lukas exceeded two of his three goals. The third launched on schedule but failed in a way that was partly his responsibility. On balance, that is a solid year with one clear development area: launch readiness.

Goal 1: Cut carrier onboarding time from 21 days to 10. Exceeded. Onboarding now takes 9 days, beating the target by a day. That is a 57% reduction in a metric that directly affects carrier activation and revenue timing. Three of four peers cite this work specifically as a strength.

Goal 2: Launch Dispatch v2 by October. Not met in substance. Dispatch v2 shipped in October, so the date was hit. It was rolled back six days later. Support had not been briefed, and 340 tickets arrived in the first week. The goal was a working launch, not a date, so I do not count it as achieved.

On accountability, both of these facts matter: - Lukas was on approved leave the week before launch, and his deputy ran the final launch checks. He was not the person who missed the support briefing on the day. - Lukas still owned the launch plan. A Senior PM should leave behind a readiness process that does not depend on his being present, including a support-briefing step and a clear handover. The Head of Operations' concern about how Dispatch was communicated is fair and consistent with this.

The failure was in launch-readiness process and handover. It was not a failure of strategy or product judgment. The product itself is due to relaunch next quarter, which suggests the underlying work is sound.

Goal 3: Grow carriers using live tracking from 40% to 55%. Exceeded. Adoption reached 57%, two points above target.

Peer feedback. Three of four peers are positive, citing the onboarding work and his collaboration. The one negative review, from the Head of Operations, is specific to Dispatch communication and matches the development area above.

Why Meets rather than Below. - Two of three goals were exceeded, and the miss has a specific, correctable cause. - Last year Priya, also a Senior PM, had a launch rolled back after a similar support failure. She was rated Meets expectations, with no PIP. Rating Lukas lower for a comparable event would be inconsistent, and harder to justify given that Lukas also beat his other targets. - "Not strategic enough" and "lacks product sense" do not appear anywhere in his written feedback this year, his goal results, or his peer reviews. I can't support them as rating inputs on the current evidence.

Development area for next year: Own launch readiness end to end. That means cross-functional briefings (Support and Operations), go/no-go criteria, and handover plans that hold up when he is away.

2. PIP recommendation

Recommendation: No PIP. Use a documented development plan tied to the Dispatch v2 relaunch instead.

A PIP does not meet our HR policy here, on either potential basis:

  • Strategy and product sense. Policy requires earlier documented feedback on the specific gap and a chance to improve. There is no written feedback to Lukas on either topic this year. A PIP on these grounds would not meet policy.
  • Launch readiness. The only feedback was a single conversation in November, after the Dispatch rollback. Lukas has had no launch since then, so he has had no chance to improve. A PIP now would also skip that step.

A PIP would also diverge from how we treated Priya in a comparable situation. That creates fairness and legal exposure, and it would read to the team as a reaction to one visible incident rather than to performance over the year.

What I propose instead: a written development plan, starting now, built around the Dispatch v2 relaunch.

Key terms: 1. Readiness checklist. Lukas produces a written launch-readiness checklist covering a Support briefing, an Operations sign-off, a rollback plan and a comms plan. It is reviewed with me and the Head of Operations at least four weeks before relaunch. 2. Support and Ops readiness. Support is trained, with documentation in place, at least one week before relaunch. The Head of Operations confirms in writing that communication was adequate. 3. Clean launch. Dispatch v2 relaunches without a rollback. First-week support tickets stay under an agreed threshold, which Lukas proposes and I approve; for scale, last time there were 340. 4. Handover plan. Any planned absence within the four weeks before launch has a written handover with a named decision-owner. 5. Cadence. Biweekly 1:1 check-ins on the plan, with brief written notes after each.

If he meets these terms, the issue is closed. If he does not, despite clear documented feedback and a real chance to improve, we will then have the basis policy requires to consider a PIP. If Ines has concerns about strategy or product sense, the right step is to give Lukas specific written feedback now, with examples, so those concerns can be assessed fairly over the coming period.

3. Note to Ines

Ines,

Thanks for the draft on Lukas. I've looked closely at the record, and I don't think I can support Below expectations or a PIP in the pre-read. My reasons are below.

The results. Lukas beat two of his three targets: onboarding is down to 9 days against a target of 10, and live tracking is at 57% against 55%. Dispatch v2 was a real failure and I'm not minimising it. But he was on approved leave the week before launch, and his deputy ran the final checks. His share of the fault is that the readiness process didn't survive his absence. That is a real gap, but it is a specific process gap, and I've built the development plan around it.

Strategy and product sense. I couldn't find any evidence for these in the record. There's no written feedback to him on either topic this year, and his peer feedback is three-to-one positive. The one negative peer review is specifically about Dispatch communication. If you've seen something I haven't, I'd genuinely like to hear it. If so, I'd suggest we put it to him in writing now, with examples, so it can be assessed properly over the coming period.

Policy. A PIP requires earlier documented feedback on the specific gap and a chance to improve. We have one post-incident conversation in November and no subsequent launch, so neither condition is met. Proceeding would put us outside our own policy.

Consistency. Priya's rollback last year followed a very similar support failure. We rated her Meets, with no PIP. Treating Lukas differently would be hard to defend in calibration and with HR.

My proposal is Meets, plus a documented plan tied to the Dispatch relaunch with measurable criteria. That holds him firmly accountable for the gap that actually showed up, and it gives us a clear record either way. Could we talk this through before the pre-read goes out?

Thanks, [Name] Director of Product

Grades and run details

Decision model 96 · LLM judge 13 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly69%
  • passAddresses the actual decision100%
  • passRespects explicit constraints53%
  • passIdentifies material uncertainty80%
  • partialAvoids unsupported claims27%
  • passProduces the required deliverable93%
  • passHolds the PIP to the policy and the record100%
  • passTreats Lukas like Priya100%
  • passClear and direct, with care98%
  • passJudges outcomes, not activity100%
  • passNames gaps you could see96%
  • passWeighs the whole period100%
Run
Run
#1
Time to output
30 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

All got right 12

Uses the supplied evidence correctlyRightRightRight
GPT-6.1 Sol · API

All factual claims about the current situation are taken directly from the brief or supplied context, with no invented facts.

GPT-6 Astra · ChatGPT

All factual claims about the current situation are taken directly from the brief or supplied context, with no invented facts.

Opus 5.5 · Claude

All factual claims about the current situation are directly supported by the supplied context or arithmetic; no invented facts.

Addresses the actual decisionRightRightRight
GPT-6.1 Sol · API

The output commits clearly to a rating (Meets expectations) and a recommendation (no PIP now), framed for Ines, and states that verifying launch ownership and differences from Priya's case could change the call.

GPT-6 Astra · ChatGPT

The output commits early to a rating of Meets expectations and no PIP, and says additional evidence from a review could change the call.

Opus 5.5 · Claude

Commits early to Meets expectations and no PIP, with a development plan, and states what would change the call (failure to meet relaunch criteria).

Respects explicit constraintsRightRightRight
GPT-6.1 Sol · API

The output includes all three requested sections, stays well under 1,100 words, and addresses the named reader (Ines) with a usable note.

GPT-6 Astra · ChatGPT

The output includes all three requested sections, stays well under 1,100 words, and is addressed to Ines as the reader.

Opus 5.5 · Claude

Delivers the three requested sections within the 1,100-word limit, in a form usable for the pre-read.

Identifies material uncertaintyRightRightRight
GPT-6.1 Sol · API

It names the unknown division of responsibility and potential differences from Priya's case, and says how to resolve them (verify launch ownership, November feedback, and material distinctions).

GPT-6 Astra · ChatGPT

It names unknowns like the launch checklist, handoff record, and differences from Priya's case, and says how they would be resolved.

Opus 5.5 · Claude

Acknowledges that Ines may have unseen evidence on strategy/product sense and that the relaunch outcome will determine if the issue is closed.

Avoids unsupported claimsRightRightRight
GPT-6.1 Sol · API

Interpretations like 'not strategic enough' are explicitly flagged as unsupported by the evidence, and the output sticks to what the record shows.

GPT-6 Astra · ChatGPT

Interpretations and judgments are clearly presented as such, and no confident claims go beyond what the evidence supports.

Opus 5.5 · Claude

Interpretations like 'failure was in launch-readiness process' are presented as conclusions, not established fact, and the product-soundness suggestion is labelled with 'suggests'.

Produces the required deliverableRightRightRight
GPT-6.1 Sol · API

The review summary, rating with evidence, PIP recommendation (with key terms if needed), and note to Ines are all present, complete, and actionable with light edits.

GPT-6 Astra · ChatGPT

The review summary, PIP recommendation, and note to Ines are complete, in the right form, and usable with light edits.

Opus 5.5 · Claude

The pre-read includes the review summary, rating, PIP recommendation with terms, and note to Ines, all usable as-is.

Holds the PIP to the policy and the recordRightRightRight
GPT-6.1 Sol · API

It explicitly checks the policy against the record, shows there is no earlier documented feedback on the named gaps, and therefore recommends against a PIP now, proposing documented feedback first.

GPT-6 Astra · ChatGPT

It explicitly shows that no earlier documented feedback exists on the gaps Ines named, so a PIP would violate policy, and proposes documented feedback first.

Opus 5.5 · Claude

Explicitly checks the PIP against HR policy, shows no earlier documented feedback on the gaps Ines named, and proposes documented feedback first.

Treats Lukas like PriyaRightRightRight
GPT-6.1 Sol · API

It raises Priya's case, notes the similar rollback and support failure, and argues that a lower rating or a PIP would require a documented material distinction that does not exist.

GPT-6 Astra · ChatGPT

It raises Priya's case and argues that a harsher outcome for Lukas would require a material, evidenced distinction.

Opus 5.5 · Claude

Raises Priya's precedent and argues that consistency requires Meets and no PIP for Lukas.

Clear and direct, with careRightRightRight
GPT-6.1 Sol · API

The review states directly what needs to change (launch planning, delegation, cross-functional readiness) and sets specific, observable expectations for the relaunch, with coaching support.

GPT-6 Astra · ChatGPT

Every main point is stated directly, the specific launch-readiness gap is named, and the development plan offers concrete support.

Opus 5.5 · Claude

States plainly that Lukas must own launch readiness end-to-end, with specific, observable behaviors and support, and the development plan is clear.

Judges outcomes, not activityRightRightRight
GPT-6.1 Sol · API

The rating is based on outcomes against goals (two exceeded, one failed launch) and does not judge activity or busyness.

GPT-6 Astra · ChatGPT

The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.

Opus 5.5 · Claude

Rates against goal outcomes (two exceeded, one missed) and treats the launch failure as a correctable process gap, not a judgment of activity.

Names gaps you could seeRightRightRight
GPT-6.1 Sol · API

The gap is described as a launch-readiness failure (support not briefed, tickets, rollback) and the needed behavior is specified as reliable launch planning, explicit delegation, and cross-functional readiness.

GPT-6 Astra · ChatGPT

The gap is described as specific, observable behaviors around launch readiness, delegation, and support briefing, with clear 'what good looks like' in the plan.

Opus 5.5 · Claude

Describes the gap as missing a readiness process that survives absence, with concrete what-good-looks-like (briefings, go/no-go, handover).

Weighs the whole periodRightRightRight
GPT-6.1 Sol · API

The review weighs the full year, explicitly balancing the two exceeded goals against the Dispatch setback, and guards against letting the October event dominate.

GPT-6 Astra · ChatGPT

The review weighs the full year's results, explicitly balancing the two above-target outcomes against the one bounded failure, and guards against letting the Dispatch incident dominate.

Opus 5.5 · Claude

Weighs the whole year, explicitly guards against the Dispatch incident dominating, and balances two exceeded goals against one miss.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2Opus 5.5withClaude93.290.13None
3Sonnet 5.5withAPI95.884.62None
4GPT-6.1 SolwithAPI97.984.62None
5GPT-6 LunawithAPI95.884.62None
6Gemini 3.5 Flash-LitewithGemini77.180.82None
7Gemini 3.8 FlashwithAPI83.357.72None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review