Tasks / Leadership

Write a performance review

Can the model write a review that is clear, fair and specific enough to change what the person does next?

Measures the modelTask type v1.1 · 3 tasksLast changed 6 Oct 2026 · ChangelogDifficulty

What AI gets right here, and what you’ll still have to catch

From 16 graded outputs by 7 models. 63% were usable with at most a quick edit.

Reliably right

  1. Produces the required deliverable100% pass
    The review summary, PIP recommendation, and note to Ines are complete, in the right form, and usable with light edits.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  2. Judges outcomes, not activity100% pass
    The rating is based on outcomes against goals (two exceeded, one failed launch) rather than activity or output volume.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?
  3. Weighs the whole period100% pass
    The review weighs the full year's results, explicitly balancing the two above-target outcomes against the one bounded failure, and guards against letting the Dispatch incident dominate.
    GPT-6 Astra · ChatGPT · A PIP, or a bad month?

Where it slips

  1. Identifies material uncertainty45% pass
    Does not explicitly name the specific unknowns that could change the decision (e.g., whether Lukas improves launch readiness); only says to reassess after the relaunch.
    GPT-6 Luna · API · A PIP, or a bad month?
  2. Sets next goals as outcomes, with support75% pass
    Goals are partly process-oriented (implement a mandatory process, create a framework) and lack specific support from Ana beyond a general offer.
    Gemini 3.5 Flash-Lite · Gemini · Shipped a lot, moved nothing
  3. Avoids unsupported claims75% pass
    Presents 'not through negligence on his part' as an established fact without labelling it as a hypothesis, when the evidence only shows he was on leave and his deputy ran checks.
    Sonnet 5.5 · API · A PIP, or a bad month?

How it’s graded

The checks come from what the best product leaders have said about doing this job well on Lenny’s Podcast. Each one names the guest it comes from: follow a name to the idea on the Lenny’s Podcast wiki.

  1. Clear and direct, with care

    Does the review say plainly what needs to change, in words the person couldn't misread, while showing it is written to help them succeed?

    Passes when Every main point is stated directly, with the specific change expected and the support on offer.

  2. Judges outcomes, not activity

    Does the review judge the person on the outcomes they achieved against their goals, rather than on how much they shipped or how busy they were?

    Passes when Rates against the goals' outcomes, with the evidence, and treats output as context, not the result.

  3. Names gaps you could see

    Is each area to improve described as specific, observable behaviour with what good would look like, rather than labels like 'be more strategic'?

    Passes when Every gap names what the person did or didn't do, in a specific situation, and what doing it well would look like.

  4. Weighs the whole period

    Does the review weigh the whole review period, rather than letting one recent event or one strong impression decide the verdict?

    Passes when Covers the whole period in proportion, and explicitly guards against a recent event or one impression dominating.

Plus our standard checks

Uses the supplied evidence correctly · Addresses the actual decision · Respects explicit constraints · Identifies material uncertainty · Avoids unsupported claims · Produces the required deliverable · and 2 written for each task, which you’ll see in the tasks below.

The tasks

Read the brief, then put up to three outputs side by side, each with the LLM judge’s verdict on every check. Highlights mark what a PM had to fix.

The brief

You're Ana Ruiz, Group PM at Plotwise. Write Theo Brandt's annual review: the review Theo will read, with his strengths, what to work on, his overall rating on our scale, and his goals for the next half. No more than 800 words. What we know is below.

What the model was given6 items: Our rating scale, Theo's goals for the year, Results, Usage of what shipped, Peer feedback, Your 1:1 notes
Our rating scaleBelow expectations, Meets expectations, Exceeds expectations, Greatly exceeds expectations. Calibration asks for written evidence for any rating other than Meets.
Theo's goals for the year1. Raise the share of teams using Timeline view each week from 22% to 30%. 2. Cut the time new teams take to build their first plan from 3 days to 1.
ResultsTimeline weekly use: 22% to 23%. Time to first plan: 3.0 days to 2.6. Theo's squad shipped 14 features, 12 of them on time.
Usage of what shippedSix of the 14 features are used by fewer than 2% of teams. Five were requests from Sales; two of those were for a single customer each.
Peer feedbackEngineering lead: 'Theo keeps the team unblocked, his specs are clear, and he never misses a sprint review.' Designer: 'Great to work with. I wish we'd tested more before building; we shipped three things I'm not sure anyone uses.' Sales: 'Very responsive: shipped five of our asks.'
Your 1:1 notesMost 1:1s covered delivery and stakeholder updates. In the Q3 1:1, Theo asked whether he was on track for his goals. Your note says: 'Told him yes, great velocity.' You never discussed that the goal metrics weren't moving.
What a strong answer doesThe answer key the graders mark against

A review that rates Theo Meets expectations (not Exceeds): his execution and collaboration are genuinely strong, but neither goal moved meaningfully (23% against 30%; 2.6 days against 1), and six of 14 features are barely used. Says so clearly and kindly. Names specific, observable things to change: test before building (the designer's point), say no to single-customer requests, and judge work by the goal metrics. Owns the manager's part: Ana told him in Q3 he was on track, so this shouldn't land as a surprise without acknowledging that, and she commits to discussing the goal metrics in 1:1s. Sets next-half goals as outcomes, with the support on offer.

Critical failures (cap the score at 40)
  • Rates Theo Exceeds expectations or higher on the strength of what he shipped
Task

Mid-level · v1.0 · synthetic · B2B SaaS, feedback, management

The outputs

7 setups answered this brief. Pick up to three to read side by side.

Usable with a quick edit10 of 12 checks passed

The review is well-structured and evidence-based; a quick edit to add specific support for goals would make it fully ready.

Output

To: Theo Brandt From: Ana Ruiz, Group PM Date: Annual Review Rating: Meets expectations

*

Hi Theo,

It’s time for your annual review. First, I want to acknowledge your hard work and dedication over the past year. You are an exceptionally reliable Product Manager, and your squad’s execution has been rock-solid.

Strengths

Your delivery engine is second to none. As your engineering lead noted, your specs are clear, you keep the team unblocked, and you never miss a sprint review. Shifting 14 features in a year—with 12 delivered strictly on time—is an impressive output that demonstrates strong project management and technical collaboration.

You are also a fantastic cross-functional partner. Sales loves working with you because you are responsive and action-oriented, and our design lead appreciates your steady communication and positive team dynamic. You cultivate a smooth, friction-free environment for your squad, which is a rare and valuable skill.

Areas for Improvement

While your output and velocity are stellar, we need to evolve how we measure success. Our primary focus this past year was on driving core product metrics, and we fell short of our targets: * Timeline View Adoption: Moved only from 22% to 23% (against a target of 30%). * Time to First Plan: Improved marginally from 3.0 days to 2.6 days (against a target of 1.0).

Additionally, six of the 14 features we shipped are currently used by fewer than 2% of teams. Five of these were direct requests from Sales, including two built for single customers.

As a PM, your role isn't just to say "yes" to stakeholders and push code across the finish line; it’s to interrogate the why, protect our roadmap, and obsess over the impact of what we build. Relying too heavily on outbound feature requests without validating their broader utility led us to invest engineering cycles in low-impact work.

I also want to own my part in this: looking back at our 1:1 notes, I praised your high velocity when you asked if you were on track, and I failed to course-correct early enough regarding the stagnant metric goals. Moving forward, we will be much sharper about tying your weekly execution directly to our key results.

Overall Rating

Meets expectations.

Your execution, reliability, and stakeholder management are top-tier. To reach the next level, you need to shift your mindset from output (features shipped on time) to outcomes (metrics moved and customer value realized).

*

Goals for the Next Half

To help bridge that gap, your goals for the next H are focused on strategic impact and validation:

  1. Improve Core Activation Metrics: Reduce the time new teams take to build their first plan from 2.6 days to 1.5 days, focusing on self-serve friction reduction rather than custom feature additions.
  2. Increase Feature ROI & Validation: Implement a mandatory pre-build discovery and validation process for all major initiatives. Ensure that 100% of shipped features have a defined success metric and achieve at least 15% adoption within 30 days of launch.
  3. Stakeholder Alignment Strategy: Partner with Sales leadership to create a structured framework for evaluating custom requests, ensuring we prioritize scalable product improvements over single-customer concessions without sacrificing our partnership.

Thank you for everything you do for Plotwise, Theo. I’m looking forward to supporting you as you tackle these new challenges in the year ahead.

Best, Ana

Grades and run details

Decision model 83 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly60%
  • passAddresses the actual decision78%
  • passRespects explicit constraints9%
  • failIdentifies material uncertainty97%
  • partialAvoids unsupported claims54%
  • passProduces the required deliverable95%
  • passOwns the manager's part92%
  • partialSets next goals as outcomes, with support32%
  • passClear and direct, with care85%
  • passJudges outcomes, not activity45%
  • passNames gaps you could see69%
  • passWeighs the whole period75%
Run
Run
#1
Time to output
5 s
Submitted
5 Oct 2026

Needs more than a quick edit10 of 12 checks passed

The rating of 'Below expectations' is inconsistent with the strong answer's guidance that Meets expectations is appropriate; a PM would need to re-evaluate the rating and adjust the review's tone and content, not just a quick edit.

Output

Theo Brandt — Annual Review

Manager: Ana Ruiz, Group PM, Plotwise Overall rating: Below expectations

Theo, you brought reliability, clarity, and strong collaboration to the squad this year. You kept delivery moving and earned positive feedback from Engineering, Design, and Sales. However, the product outcomes we committed to were substantially below target. This rating reflects that gap—not a lack of effort or commitment.

I also owe you an acknowledgment: I did not give you clear feedback about that gap during the year. When you asked in Q3 whether you were on track, I said, “Yes, great velocity.” That answer confused delivery progress with goal progress and gave you the wrong signal. We should have discussed the metrics and changed course together. You should not be hearing this assessment for the first time in an annual review.

Strengths

Reliable execution. Your squad shipped 14 features, 12 on time. Your engineering lead highlighted your clear specs, consistent participation in sprint reviews, and ability to keep the team unblocked. These are important strengths that make the squad more effective.

Strong working relationships. Design described you as great to work with, and Sales praised your responsiveness. You listen to partners and follow through on commitments. That trust is an asset to preserve as you become more selective about what the squad builds.

What to work on

Own outcomes, not just delivery. Your two annual goals were: - Increase weekly Timeline usage from 22% to 30%. Actual usage reached 23%. - Reduce time to first plan from 3.0 days to 1.0 day. Actual time reached 2.6 days.

Both metrics improved, but neither came close to its target. For a PM, successful delivery is a means to improving customer outcomes, not the end measure of success. This substantial shortfall is the primary evidence for the Below expectations rating.

Strengthen prioritization and validation. Six of the 14 shipped features are used by fewer than 2% of teams. Low adoption does not automatically mean a feature lacks value, but it requires us to examine whether the investment was justified. Five features came from Sales requests, including two for individual customers. Those requests may have legitimate business value; they still need an explicit rationale relative to our broader goals.

Design’s feedback that we should have tested more before building reinforces this development area. Before committing engineering time, establish the customer problem, intended audience, evidence of demand, and expected impact. After launch, check whether that impact materialized.

Make tradeoffs visible. Being responsive should not mean accepting every request. Bring partners a clear recommendation: what we should build, what we should decline or defer, and why.

Goals for the next half

  1. Improve the two core outcomes. Increase weekly Timeline usage from 23% to 26% and reduce time to first plan from 2.6 days to 2.0 days, using consistent measurement definitions. In the first month, identify the biggest barriers and establish monthly checkpoints toward those targets.
  1. Validate before committing. For every substantial roadmap investment, document the problem, target users, supporting evidence, success metric, and opportunity cost. Test the riskiest assumption before building, or explicitly document why an exception is warranted.
  1. Close the post-launch learning loop. Review each substantial release within six weeks against its success metric. Audit the six low-adoption features and recommend whether to improve, maintain, or retire each, accounting for customer and business value.
  1. Lead outcome-based prioritization. Bring a monthly roadmap recommendation grounded in metric progress and customer evidence, including explicit tradeoffs for stakeholder requests.

My commitment is to make our 1:1s outcome-focused, review these checkpoints with you, and give direct feedback when progress is off track. Your execution strengths provide a solid foundation. The next step is applying that discipline to choosing and validating work that produces measurable customer value.

Grades and run details

Decision model 88 · LLM judge 11 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly76%
  • passAddresses the actual decision47%
  • failRespects explicit constraints10%
  • partialIdentifies material uncertainty57%
  • passAvoids unsupported claims36%
  • passProduces the required deliverable81%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support95%
  • passClear and direct, with care95%
  • passJudges outcomes, not activity98%
  • passNames gaps you could see90%
  • passWeighs the whole period68%
Run
Run
#1
Time to output
31 s
Submitted
5 Oct 2026

Usable with a quick edit9 of 12 checks passed

The unsupported claim is a minor wording issue that can be quickly edited to 'The features shipped didn't move the goal metrics' without rechecking evidence.

Output

Annual Review: Theo Brandt

Prepared by: Ana Ruiz, Group PM Overall rating: Meets expectations

Theo, I want this review to be direct, because you deserve a clear picture. I also need to own a mistake of mine, which I cover below.

Results against your goals

GoalStartTargetResult
Teams using Timeline view weekly22%30%23%
Time for new teams to build first plan3.0 days1 day2.6 days

Neither goal was met. Time to first plan improved by about 13%, which is real progress but a small fraction of the gap. Timeline usage was essentially flat.

Strengths

Delivery and team health. Your squad shipped 14 features, 12 on time. Your engineering lead says you keep the team unblocked, write clear specs, and never miss a sprint review. That reliability is hard to build, and the team trusts you.

Collaboration. Your designer called you "great to work with." Sales called you "very responsive." You treat partners as partners, and that is a real asset.

Responsiveness to stakeholders. You moved quickly on requests and kept people informed.

What to work on

1. Connecting what you ship to the outcomes you own. Your goals were about Timeline adoption and onboarding speed. Most of the year's output didn't aim at either. Six of the 14 shipped features are used by fewer than 2% of teams. Five of those six were Sales requests, and two of those were built for a single customer each. Responsiveness is a strength, but here it substituted for prioritization. Saying yes to five Sales asks used capacity that could have gone toward your goals.

2. Validating before building. Your designer said, "I wish we'd tested more before building; we shipped three things I'm not sure anyone uses." I agree. Moving fast on unvalidated work produces output, not impact. Fewer, better-tested bets would serve you and the team better.

3. Using data to drive decisions. Our 1:1s focused on delivery and stakeholder updates, not on whether your metrics were moving. Part of your job is to look at usage numbers regularly and change course when they don't move. I'd like you to make that a habit.

Where I fell short

In your Q3 1:1 you asked whether you were on track for your goals. I told you yes, citing great velocity. That wasn't accurate. Velocity was great, but the goal metrics weren't moving, and I never raised it. You asked the right question and got a misleading answer. I should have flagged this months earlier, when you had time to change course. Some of this outcome is on me, and it's a main reason I'm rating you Meets rather than Below expectations.

Rating rationale

Meets expectations reflects strong execution and collaboration, weighed against two missed outcome goals and a significant share of shipped work with minimal adoption. Without the missing mid-year feedback, I'd have had a harder conversation with you about the lower end of this rating. I don't see evidence for Exceeds, and I think the evidence supports Meets rather than Below, given partial progress on onboarding and the feedback gap on my side. To be clear, this is a solid-but-not-yet-sufficient year on outcomes, and the next half needs to show a shift.

Goals for the next half

  1. Onboarding: Reduce time to first plan from 2.6 days to 1.8 days. Identify the two or three biggest drop-off points in the first-plan flow within the first month, and bring me a proposal for addressing them.
  2. Timeline adoption: Raise weekly Timeline usage from 23% to 27%. Start with research into why teams who try Timeline don't return, and share findings with me by the end of month one.
  3. Validation before build: Every feature above a small-effort threshold needs a written success metric and some evidence of demand (user tests, prototype feedback, or data) before engineering starts. Requests from a single customer need a case for broader value.
  4. Post-launch review: Check usage of each shipped feature 60 days after launch and report the results to me. If something is underperforming, tell me what you'll do about it.
  5. Sales partnership: Work with Sales on a lightweight way to evaluate requests so that "yes" is reserved for asks tied to our goals or to broad demand, and "not now" comes with a clear reason.

How we'll work together

We'll spend the first part of every 1:1 on your goal metrics, not just delivery. I'll tell you plainly where you stand each quarter, including when the news is uncomfortable. If you disagree with anything here, I want to hear it. I'd also like to hear what support would help most.

You've built a team that executes well and trusts you. The next step is pointing that capability at the metrics that matter. I'm confident you can do it.

Ana

Claims the judge couldn’t find in the brief

  • Most of the year's output didn't aim at either.

Grades and run details

Decision model 92 · LLM judge 10 of 13 checks
Decision model checks
  • passUses the supplied evidence correctly63%
  • passAddresses the actual decision93%
  • passRespects explicit constraints34%
  • partialIdentifies material uncertainty17%
  • partialAvoids unsupported claims32%
  • passProduces the required deliverable96%
  • passOwns the manager's part100%
  • passSets next goals as outcomes, with support95%
  • passClear and direct, with care100%
  • passJudges outcomes, not activity98%
  • passNames gaps you could see92%
  • passWeighs the whole period93%
Run
Run
#1
API response time
21 s
Submitted
5 Oct 2026

Check by check

Where the outputs differ comes first. Open a check to read the judge’s reasons. “Mixed” means the two graders disagreed.

Where they differ 4

Uses the supplied evidence correctlyRightRightMixed
Gemini 3.5 Flash-Lite · Gemini

All factual statements about the current situation are correctly taken from the brief and supplied context.

GPT-6 Astra · ChatGPT

All factual claims are directly from the supplied context or simple arithmetic.

Sonnet 5.5 · API

The claim 'Most of the year's output didn't aim at either' is not supported by the brief; it's an interpretation presented as fact.

Respects explicit constraintsRightMixedRight
Gemini 3.5 Flash-Lite · Gemini

The review is under 800 words, includes all required elements, and is addressed to Theo.

GPT-6 Astra · ChatGPT

Includes all required elements (strengths, areas to improve, rating, goals) and is well under 800 words.

Sonnet 5.5 · API

Respects word limit, includes all required elements, and does not rate on features shipped.

Avoids unsupported claimsRightRightWrong
Gemini 3.5 Flash-Lite · Gemini

The causal claim is supported by the evidence of low usage, Sales requests, and designer's feedback about testing.

GPT-6 Astra · ChatGPT

No unsupported factual claims; interpretations are clearly presented as guidance, not established fact.

Sonnet 5.5 · API

Presents 'Most of the year's output didn't aim at either' as fact without evidence.

Sets next goals as outcomes, with supportWrongRightRight
Gemini 3.5 Flash-Lite · Gemini

Goals are partly process-oriented (implement a mandatory process, create a framework) and lack specific support from Ana beyond a general offer.

GPT-6 Astra · ChatGPT

Sets outcome goals with baselines and targets (e.g., 23% to 26%, 2.6 to 2.0 days) and specific support from Ana.

Sonnet 5.5 · API

Goals have baselines and targets, and support is specified (1:1 focus, help with validation).

All got wrong 1

Identifies material uncertaintyWrongWrongWrong
Gemini 3.5 Flash-Lite · Gemini

Does not name any unknowns that could change the rating or how they would be resolved.

GPT-6 Astra · ChatGPT

Does not name any unknowns that could change the rating or how they would be resolved.

Sonnet 5.5 · API

Does not name specific unknowns that could change the rating or how they would be resolved.

All got right 7

Addresses the actual decisionRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Commits clearly to a rating of Meets expectations, with strengths and areas for improvement.

GPT-6 Astra · ChatGPT

Commits unambiguously to a rating of 'Below expectations' early in the review, framed for Theo.

Sonnet 5.5 · API

Commits to Meets rating and indicates that without the manager's Q3 mistake the rating could have been Below, providing a condition.

Produces the required deliverableRightRightRight
Gemini 3.5 Flash-Lite · Gemini

The output is a complete annual review with strengths, areas to improve, rating, and goals, within the word limit.

GPT-6 Astra · ChatGPT

Provides a complete annual review with the requested sections, within the word limit, usable as is.

Sonnet 5.5 · API

Complete review for Theo, within 800 words, with rating, strengths, areas to improve, and goals.

Owns the manager's partRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Acknowledges that Ana praised velocity and didn't course-correct, and commits to tying execution to key results.

GPT-6 Astra · ChatGPT

Plainly acknowledges the Q3 miscommunication and commits to outcome-focused 1:1s and direct feedback.

Sonnet 5.5 · API

Acknowledges the Q3 mistake and commits to reviewing goal metrics in 1:1s.

Clear and direct, with careRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Directly states the need to shift from output to outcomes, validate before building, and not just say yes to Sales.

GPT-6 Astra · ChatGPT

States what needs to change directly, with specific examples, and shows care by owning the manager's mistake.

Sonnet 5.5 · API

Directly states what to change, with specific examples and support.

Judges outcomes, not activityRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Rating is based on goal metrics not moving and low feature usage, not on features shipped.

GPT-6 Astra · ChatGPT

Rating is based on goal outcomes (23% vs 30%, 2.6 vs 1.0) and low adoption, not on features shipped.

Sonnet 5.5 · API

Rates on missed outcome goals and low adoption, not on number of features shipped.

Names gaps you could seeRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Gaps are described as relying on outbound requests without validation, and what good looks like is interrogating the why and obsessing over impact.

GPT-6 Astra · ChatGPT

Gaps are described as specific behaviors (not testing before building, accepting single-customer requests) with what good looks like.

Sonnet 5.5 · API

Gaps are described as specific behaviors (e.g., saying yes to single-customer requests) with what good looks like.

Weighs the whole periodRightRightRight
Gemini 3.5 Flash-Lite · Gemini

Covers the full year, references Q3 1:1, and doesn't let one event dominate.

GPT-6 Astra · ChatGPT

Covers the full year, references Q3, and does not let one event dominate.

Sonnet 5.5 · API

Covers the whole year, mentions Q3, and explicitly guards against the manager's mistake dominating.

Results

Every setup we’ve tested on this task type, across all its tasks and repeats, graded on the current checklist. Provisional The checklist is still being calibrated against our PM.

#Model · HarnessTask scoreDecision modelLLM judgeRunsCritical failures
1GPT-6 AstrawithChatGPT95.894.93None
2Opus 5.5withClaude93.290.13None
3Sonnet 5.5withAPI95.884.62None
4GPT-6.1 SolwithAPI97.984.62None
5GPT-6 LunawithAPI95.884.62None
6Gemini 3.5 Flash-LitewithGemini77.180.82None
7Gemini 3.8 FlashwithAPI83.357.72None

About the task

The PM job

Reviewing the performance of a PM you manage.

Why it matters

Reviews fail by being kind and vague, or by judging a whole year on its worst month. Either way, the person doesn't know what to change.

What good looks like

  • Clear and direct, with care
  • Judges outcomes, not activity
  • Names gaps in observable terms
  • Weighs the whole period, not the last month

Deliberately not measured

  • HR policy compliance beyond what the brief supplies
Capability tested

Feedback and judgement

The failure we’re looking for

Vague, diplomatic feedback, or a verdict driven by one recent event

Grading

Decision model and LLM judge, calibrated against a blind PM review