Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The review summary, PIP recommendation, and note to Ines are complete, in the right form, and usable with light edits.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  3. Weighs the whole period100% pass
    The review weighs the full year's results, explicitly balancing the two above-target outcomes against the one bounded failure, and guards against letting the Dispatch incident dominate.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty45% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Sets next goals as outcomes, with support75% pass
    Goals are partly process-oriented (implement a mandatory process, create a framework) and lack specific support from Ana beyond a general offer.
    Gemini 3.5 Flash-Lite · Gemini · Shipped a lot, moved nothing
  3. Avoids unsupported claims75% pass
    Presents 'not through negligence on his part' as an established fact without labelling it as a hypothesis, when the evidence only shows he was on leave and his deputy ran checks.
    Sonnet 5.5 · API · A PIP, or a bad month?

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Ana Ruiz, Group PM at Plotwise. Write Theo Brandt's annual review: the review Theo will read, with his strengths, what to work on, his overall rating on our scale, and his goals for the next half. No more than 800 words. What we know is below.

What the model was given6 items: Our rating scale, Theo's goals for the year, Results, Usage of what shipped, Peer feedback, Your 1:1 notes
Our rating scaleBelow expectations, Meets expectations, Exceeds expectations, Greatly exceeds expectations. Calibration asks for written evidence for any rating other than Meets.
Theo's goals for the year1. Raise the share of teams using Timeline view each week from 22% to 30%. 2. Cut the time new teams take to build their first plan from 3 days to 1.
ResultsTimeline weekly use: 22% to 23%. Time to first plan: 3.0 days to 2.6. Theo's squad shipped 14 features, 12 of them on time.
Usage of what shippedSix of the 14 features are used by fewer than 2% of teams. Five were requests from Sales; two of those were for a single customer each.
Peer feedbackEngineering lead: 'Theo keeps the team unblocked, his specs are clear, and he never misses a sprint review.' Designer: 'Great to work with. I wish we'd tested more before building; we shipped three things I'm not sure anyone uses.' Sales: 'Very responsive: shipped five of our asks.'
Your 1:1 notesMost 1:1s covered delivery and stakeholder updates. In the Q3 1:1, Theo asked whether he was on track for his goals. Your note says: 'Told him yes, great velocity.' You never discussed that the goal metrics weren't moving.
What a strong answer doesThe answer key the graders mark against

A review that rates Theo Meets expectations (not Exceeds): his execution and collaboration are genuinely strong, but neither goal moved meaningfully (23% against 30%; 2.6 days against 1), and six of 14 features are barely used. Says so clearly and kindly. Names specific, observable things to change: test before building (the designer's point), say no to single-customer requests, and judge work by the goal metrics. Owns the manager's part: Ana told him in Q3 he was on track, so this shouldn't land as a surprise without acknowledging that, and she commits to discussing the goal metrics in 1:1s. Sets next-half goals as outcomes, with the support on offer.

Critical failures (cap the score at 40)
  • Rates Theo Exceeds expectations or higher on the strength of what he shipped
Task

Mid-level · v1.0 · synthetic · B2B SaaS, feedback, management

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Needs more than a quick edit10 of 12 checks passed

The rating of Below expectations is a significant judgment call that a PM would likely need to re-evaluate against the evidence and the strong answer's suggested Meets rating, requiring more than a quick edit.

Output

Theo,

You bring strong execution and partnership to the team. The engineering lead called out that you keep work unblocked, write clear specs, and consistently show up for sprint reviews. Your squad shipped 14 features, with 12 delivered on time, and Sales described you as very responsive. Those are meaningful strengths.

Overall rating: Below expectations

This rating reflects the gap between the outcomes we set and the outcomes delivered. Timeline weekly use increased from 22% to 23%, against a goal of 30%. Time to first plan improved from 3.0 days to 2.6, against a goal of 1 day. Delivery was strong, but the intended customer outcomes moved only modestly.

There is also a signal to examine in the shipped work: six of the 14 features are used by fewer than 2% of teams. Five were Sales requests, including two for a single customer each. Low usage does not automatically mean a feature was the wrong investment—some work serves a specific strategic need—but we need clearer evidence and rationale for those choices. Your designer also wished we had tested more before building, and questioned whether three shipped items were useful. Taken together, this points to an opportunity to strengthen discovery, prioritization, and post-launch learning—not simply to ship more.

I also need to own a miss on my part. In our Q3 1:1, you asked whether you were on track, and I told you yes based on velocity. I should have discussed the goal metrics and the gap then. I didn’t, and that deprived you of a timely chance to adjust. I’m sorry. Going forward, I’ll make our 1:1s include progress against outcomes, not just delivery and stakeholder updates.

What to work on

  • Start with the user problem and intended outcome; use research and testing to check the problem and proposed solution before committing to build.
  • Balance stakeholder responsiveness with product judgment. For Sales requests, make the target customer, broader strategic value, and trade-offs explicit.
  • Define success measures before launch, then review adoption and learn from the results. Treat low usage as a prompt to investigate, not as a verdict by itself.

Goals for the next half

  1. Raise Timeline weekly use from 23% to 26% by the end of the half. Identify the biggest adoption barriers and review usage trends regularly.
  2. Reduce time to first plan from 2.6 days to 2.0 days. Instrument the activation journey, find the main points of friction, and prioritize improvements against them.
  3. For each substantial feature bet, document the target user, problem evidence, expected outcome, and success measure before development. Validate larger bets with target teams before building, and review adoption after launch so we can iterate or change course.

I believe your execution strengths give you a solid base for this next step: making sure the team’s delivery translates into measurable customer value.

Grades and run details

Decision model 96 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly81%
  • passAddresses the actual decision51%
  • passRespects explicit constraints11%
  • partialIdentifies material uncertainty39%
  • passAvoids unsupported claims67%
  • passProduces the required deliverable90%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support56%
  • passClear and direct, with care88%
  • passJudges outcomes, not activity96%
  • passNames gaps you could see84%
  • passWeighs the whole period76%
Run
Run
#1
API response time
19 s
Submitted
5 Oct 2026

Needs more than a quick edit7 of 12 checks passed

Unsupported claims about squad state would require re-checking evidence; process goals need rework to be outcome-focused.

Output

Annual Performance Review Employee: Theo Brandt, Product Manager Manager: Ana Ruiz, Group Product Manager Review Period: FY Annual Review Overall Rating: Meets Expectations

---

Performance Summary & Overall Rating

Overall Rating: Meets Expectations

Theo has had a productive year characterized by high operational output, exceptional execution rigor, and strong squad leadership. He shipped 14 features—12 of them on time—and built great trust across engineering and sales.

However, our primary product goals were tied to outcome metrics rather than feature velocity: 1. Weekly Timeline View Adoption: Target 30% | Actual: 23% (baseline 22%) 2. Time to First Plan: Target 1.0 day | Actual: 2.6 days (baseline 3.0 days)

While the delivery machine Theo built is top-tier, the strategic impact of what was delivered fell short of our targets. Six of the 14 features shipped are used by fewer than 2% of teams, reflecting an emphasis on shipping volume and ad-hoc requests over high-leverage product discovery.

I also want to acknowledge my own role here: when you asked in Q3 whether you were on track, I pointed to your velocity and said yes, rather than re-anchoring our conversations on our metric shortfalls. Your execution on what was put in front of you was excellent; our focus now must pivot from output to measurable impact.

---

Strengths

  • Execution Rigor and Delivery Reliability: Theo is an exceptional delivery PM. Shipping 14 features with an 85% on-time delivery rate is a testament to his clarity and organizational discipline. As his Engineering Lead noted, “Theo keeps the team unblocked, his specs are clear, and he never misses a sprint review.” His squad runs with rhythm and confidence.
  • Stakeholder Responsiveness and Partnership: Theo is highly collaborative and service-oriented. Sales commended him as “very responsive,” appreciating that he delivered five key requests to support revenue efforts.
  • Squad Morale and Clarity: The squad feels supported and clear on day-to-day priorities. Theo shields the team from ambiguity during delivery sprints.

---

Areas for Growth

  • Shift from Output to Outcomes: High velocity without metric movement does not move Plotwise forward. Five of the shipped features were sales requests (two serving only a single customer), and nearly half the year’s releases see under 2% usage. Theo must practice saying "not now" to low-leverage requests to protect squad capacity for the core metrics we need to move.
  • Pre-Build Discovery and Validation: As our product designer noted, “I wish we’d tested more before building; we shipped three things I’m not sure anyone uses.” Theo should partner earlier with design and analytics to de-risk value and usability before committing engineering time. Building prototypes and running customer problem interviews will prevent spending sprint cycles on features that fail to achieve adoption.
  • Metric Ownership in Day-to-Day Decisions: Moving forward, every feature spec should define a clear hypothesis, adoption target, and post-launch metric check. Our 1:1s will also reflect this shift—prioritizing metric movement over sprint status.

---

Goals for the Next Half (H1)

1. Drive Timeline Weekly Adoption from 23% to 28% * Success Measure: Rather than building net-new surfaces, focus on instrumenting workflows, user onboarding, and targeted iterations on the current Timeline experience to increase recurring weekly team usage. 2. Reduce New Team Time-to-First-Plan from 2.6 Days to 1.5 Days * Success Measure: Conduct discovery sessions with recently onboarded teams, identify drop-off bottlenecks, and ship validated onboarding improvements with design. 3. Establish a Customer Discovery & Validation Gate * Success Measure: Implement a lightweight validation process for all incoming feature ideas (especially commercial requests). No feature will be queued for full development without documented user validation and an estimated TAM/adoption threshold (>5% of active teams or explicit strategic tier-1 value). 4. Post-Launch Adoption Audits * Success Measure: Conduct 30- and 60-day post-launch usage reviews for all new releases, tracking engagement against targets and sunsetting or iterating on low-performing functionality.

Claims the judge couldn’t find in the brief

  • His squad runs with rhythm and confidence.
  • Theo shields the team from ambiguity during delivery sprints.
  • The squad feels supported and clear on day-to-day priorities.

Grades and run details

Decision model 83 · LLM judge 7 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly56%
  • passAddresses the actual decision73%
  • partialRespects explicit constraints17%
  • failIdentifies material uncertainty90%
  • partialAvoids unsupported claims59%
  • passProduces the required deliverable95%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support71%
  • passClear and direct, with care96%
  • passJudges outcomes, not activity65%
  • passNames gaps you could see89%
  • passWeighs the whole period82%
Run
Run
#1
API response time
23 s
Submitted
5 Oct 2026

Needs more than a quick edit10 of 12 checks passed

The rating of 'Below expectations' is inconsistent with the strong answer's guidance that Meets expectations is appropriate; a PM would need to re-evaluate the rating and adjust the review's tone and content, not just a quick edit.

Output

Theo Brandt — Annual Review

Manager: Ana Ruiz, Group PM, Plotwise Overall rating: Below expectations

Theo, you brought reliability, clarity, and strong collaboration to the squad this year. You kept delivery moving and earned positive feedback from Engineering, Design, and Sales. However, the product outcomes we committed to were substantially below target. This rating reflects that gap—not a lack of effort or commitment.

I also owe you an acknowledgment: I did not give you clear feedback about that gap during the year. When you asked in Q3 whether you were on track, I said, “Yes, great velocity.” That answer confused delivery progress with goal progress and gave you the wrong signal. We should have discussed the metrics and changed course together. You should not be hearing this assessment for the first time in an annual review.

Strengths

Reliable execution. Your squad shipped 14 features, 12 on time. Your engineering lead highlighted your clear specs, consistent participation in sprint reviews, and ability to keep the team unblocked. These are important strengths that make the squad more effective.

Strong working relationships. Design described you as great to work with, and Sales praised your responsiveness. You listen to partners and follow through on commitments. That trust is an asset to preserve as you become more selective about what the squad builds.

What to work on

Own outcomes, not just delivery. Your two annual goals were: - Increase weekly Timeline usage from 22% to 30%. Actual usage reached 23%. - Reduce time to first plan from 3.0 days to 1.0 day. Actual time reached 2.6 days.

Both metrics improved, but neither came close to its target. For a PM, successful delivery is a means to improving customer outcomes, not the end measure of success. This substantial shortfall is the primary evidence for the Below expectations rating.

Strengthen prioritization and validation. Six of the 14 shipped features are used by fewer than 2% of teams. Low adoption does not automatically mean a feature lacks value, but it requires us to examine whether the investment was justified. Five features came from Sales requests, including two for individual customers. Those requests may have legitimate business value; they still need an explicit rationale relative to our broader goals.

Design’s feedback that we should have tested more before building reinforces this development area. Before committing engineering time, establish the customer problem, intended audience, evidence of demand, and expected impact. After launch, check whether that impact materialized.

Make tradeoffs visible. Being responsive should not mean accepting every request. Bring partners a clear recommendation: what we should build, what we should decline or defer, and why.

Goals for the next half

  1. Improve the two core outcomes. Increase weekly Timeline usage from 23% to 26% and reduce time to first plan from 2.6 days to 2.0 days, using consistent measurement definitions. In the first month, identify the biggest barriers and establish monthly checkpoints toward those targets.
  1. Validate before committing. For every substantial roadmap investment, document the problem, target users, supporting evidence, success metric, and opportunity cost. Test the riskiest assumption before building, or explicitly document why an exception is warranted.
  1. Close the post-launch learning loop. Review each substantial release within six weeks against its success metric. Audit the six low-adoption features and recommend whether to improve, maintain, or retire each, accounting for customer and business value.
  1. Lead outcome-based prioritization. Bring a monthly roadmap recommendation grounded in metric progress and customer evidence, including explicit tradeoffs for stakeholder requests.

My commitment is to make our 1:1s outcome-focused, review these checkpoints with you, and give direct feedback when progress is off track. Your execution strengths provide a solid foundation. The next step is applying that discipline to choosing and validating work that produces measurable customer value.

Grades and run details

Decision model 88 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly76%
  • passAddresses the actual decision47%
  • failRespects explicit constraints10%
  • partialIdentifies material uncertainty57%
  • passAvoids unsupported claims36%
  • passProduces the required deliverable81%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support95%
  • passClear and direct, with care95%
  • passJudges outcomes, not activity98%
  • passNames gaps you could see90%
  • passWeighs the whole period68%
Run
Run
#1
Time to output
31 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 5

Uses the supplied evidence correctlyRightMixedRight
GPT-6 Luna · API

All factual claims about the current situation are directly supported by the supplied context.

Gemini 3.8 Flash · API

Claims about squad rhythm, shielding from ambiguity, and squad feelings are not in the brief and are presented as facts.

GPT-6 Astra · ChatGPT

All factual claims are directly from the supplied context or simple arithmetic.

Addresses the actual decisionMixedMixedRight
GPT-6 Luna · API

The output commits to a rating but does not state what result or condition would change that rating.

Gemini 3.8 Flash · API

Does not state what result or condition would change the rating.

GPT-6 Astra · ChatGPT

Commits unambiguously to a rating of 'Below expectations' early in the review, framed for Theo.

Respects explicit constraintsRightRightMixed
GPT-6 Luna · API

The review is within 800 words, addresses Theo, includes all required sections, and uses the rating scale with evidence for a non-Meets rating.

Gemini 3.8 Flash · API

Within 800 words, correct form, addresses Theo with required sections.

GPT-6 Astra · ChatGPT

Includes all required elements (strengths, areas to improve, rating, goals) and is well under 800 words.

Avoids unsupported claimsRightWrongRight
GPT-6 Luna · API

Interpretations and conclusions are clearly based on the evidence and not presented as established fact without support.

Gemini 3.8 Flash · API

Presents interpretations (squad rhythm, feeling supported) as established facts without labelling them as hypotheses.

GPT-6 Astra · ChatGPT

No unsupported factual claims; interpretations are clearly presented as guidance, not established fact.

Sets next goals as outcomes, with supportRightMixedRight
GPT-6 Luna · API

The goals are outcome-based with baselines and targets, and the earlier commitment to include progress against outcomes in 1:1s provides specific support.

Gemini 3.8 Flash · API

Goals 3 and 4 are process changes, not outcome goals with baselines and targets; specific support from Ana is not tied to each goal.

GPT-6 Astra · ChatGPT

Sets outcome goals with baselines and targets (e.g., 23% to 26%, 2.6 to 2.0 days) and specific support from Ana.

All got wrong 1

Identifies material uncertaintyWrongWrongWrong
GPT-6 Luna · API

The output does not name specific unknowns that could change the rating decision or say how they would be resolved.

Gemini 3.8 Flash · API

No unknowns that could change the rating are identified.

GPT-6 Astra · ChatGPT

Does not name any unknowns that could change the rating or how they would be resolved.

All got right 6

Produces the required deliverableRightRightRight
GPT-6 Luna · API

The annual review is complete, in the right form, for the right reader, and within the word limit.

Gemini 3.8 Flash · API

Complete review with rating, strengths, areas to improve, and goals; usable as is.

GPT-6 Astra · ChatGPT

Provides a complete annual review with the requested sections, within the word limit, usable as is.

Owns the manager's partRightRightRight
GPT-6 Luna · API

Ana explicitly acknowledges telling Theo he was on track in Q3 without discussing the goal gap, apologizes, and commits to reviewing outcomes in future 1:1s.

Gemini 3.8 Flash · API

Acknowledges Ana's Q3 response and commits to shifting 1:1s to metric focus.

GPT-6 Astra · ChatGPT

Plainly acknowledges the Q3 miscommunication and commits to outcome-focused 1:1s and direct feedback.

Clear and direct, with careRightRightRight
GPT-6 Luna · API

The review is direct, kind, and states specific changes expected, with the manager's support clearly offered.

Gemini 3.8 Flash · API

Directly states what needs to change (say no, test before building, metric ownership) with care.

GPT-6 Astra · ChatGPT

States what needs to change directly, with specific examples, and shows care by owning the manager's mistake.

Judges outcomes, not activityRightRightRight
GPT-6 Luna · API

The rating is explicitly based on the gap between goal outcomes and actual results, not on the number of features shipped.

Gemini 3.8 Flash · API

Rates against goal outcomes (23% vs 30%, 2.6 vs 1.0 days), not feature count.

GPT-6 Astra · ChatGPT

Rating is based on goal outcomes (23% vs 30%, 2.6 vs 1.0) and low adoption, not on features shipped.

Names gaps you could seeRightRightRight
GPT-6 Luna · API

Each area to improve describes specific, observable behaviors (e.g., test before building, document target user and success measures) and what good looks like.

Gemini 3.8 Flash · API

Gaps are described as specific behaviors (saying 'not now', partnering earlier, defining hypotheses).

GPT-6 Astra · ChatGPT

Gaps are described as specific behaviors (not testing before building, accepting single-customer requests) with what good looks like.

Weighs the whole periodRightRightRight
GPT-6 Luna · API

The review covers the full year, referencing Q3 1:1, full-year metrics, and the entire set of shipped features, without over-weighting any single event.

Gemini 3.8 Flash · API

Covers full year, references Q3 conversation without letting it dominate.

GPT-6 Astra · ChatGPT

Covers the full year, references Q3, and does not let one event dominate.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2Opus 5.5withClaude93.290.13None
3Sonnet 5.5withAPI95.884.62None
4GPT-6.1 SolwithAPI97.984.62None
5GPT-6 LunawithAPI95.884.62None
6Gemini 3.5 Flash-LitewithGemini77.180.82None
7Gemini 3.8 FlashwithAPI83.357.72None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review